Shih-Fu Chang

dblp:c/ShihFuChang · DBLP profile ↗
← Back
388ranked-venue papers
32as first author
48since 2021 · last 2026
0000-0003-1444-1205ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 308 · 26 first-author · 31 since 2021Artificial intelligence and machine learning · 153 · 42 since 2021Databases, data management, data science and information retrieval · 23 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-authorComputer networks · 9 · 2 first-authorSystems, architecture and hardware · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3Security and privacy · 1
YearPublicationVenuePosition
2026 Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting
Guangxing Han, Long Chen 0016, Jiawei Ma, Shiyuan Huang 0001, Rama Chellappa, Shih-Fu Chang
Int. J. Comput. Vis.6
2025 From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
abstract
Data visualization in the form of charts plays a pivotal role in data analysis, offering critical insights and aiding in informed decision-making. Automatic chart understanding has witnessed significant advancements with the rise of large foundation models in recent years. Foundation models, such as large language models, have revolutionized various natural language processing tasks and are increasingly being applied to chart understanding tasks. This survey paper provides a comprehensive overview of the recent developments, challenges, and future directions in chart understanding within the context of these foundation models. We review fundamental building blocks crucial for studying chart understanding tasks. Additionally, we explore various tasks and their evaluation metrics and sources of both charts and textual inputs. Various modeling strategies are then examined, encompassing both classification-based and generation-based approaches, along with tool augmentation techniques that enhance chart understanding performance. Furthermore, we discuss the state-of-the-art performance of each task and discuss how we can improve the performance. Challenges and future directions are addressed, highlighting the importance of several topics, such as domain-specific charts, lack of efforts in developing evaluation metrics, and agent-oriented settings. This survey paper aims to provide valuable insights and directions for future research in chart understanding leveraging large foundation models.
Kung-Hsiang Huang, Hou Pong Chan, May Fung, Haoyi Qiu, Shafiq R. Joty, Shih-Fu Chang, Heng Ji 0001
IEEE Trans. Knowl. Data Eng.7
2024 Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities
abstract
Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical (via grounding) and thus, on the same semantic level. However, grounding fails to capture the intricate cross-event relations that exist due to the same events being referred to on many semantic levels. For example, the abstract event of "war'' manifests at a lower semantic level through subevents "tanks firing'' (in video) and airplane "shot'' (in text), leading to a hierarchical, multimodal relationship between the events. In this paper, we propose the task of extracting event hierarchies from multimodal (video and text) data to capture how the same event manifests itself in different modalities at different semantic levels. This reveals the structure of events and is critical to understanding them. To support research on this task, we introduce the Multimodal Hierarchical Events (MultiHiEve) dataset. Unlike prior video-language datasets, MultiHiEve is composed of news video-article pairs, which makes it rich in event hierarchies. We densely annotate a part of the dataset to construct the test benchmark. We show the limitations of state-of-the-art unimodal and multimodal baselines on this task. Further, we address these limitations via a new weakly supervised model, leveraging only unannotated video-article pairs from MultiHiEve. We perform a thorough evaluation of our proposed method which demonstrates improved performance on this task and highlight opportunities for future research. Data: https://github.com/hayyubi/multihieve
Hammad A. Ayyubi, Christopher Thomas 0004, Lovish Chum, Rahul Lokesh, Long Chen 0016, Yulei Niu, Xudong Lin 0003, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang
AAAI11
2024 What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
abstract
Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box supervision. This work addresses this task from a multimodal supervision perspective, proposing a framework for spatio-temporal action grounding trained on loose video and subtitle supervision only, without human annotation. To this end, we combine local representation learning, which focuses on leveraging fine-grained spatial information, with a global representation encoding that captures higher-level representations and incorporates both in a joint approach. To evaluate this challenging task in a real-life setting, a new benchmark dataset is proposed, providing dense spatio-temporal grounding annotations in long, untrimmed, multi-action instructional videos for over 5K events. We evaluate the proposed approach and other methods on the proposed and standard downstream tasks, showing that our method improves over current baselines in various settings, including spatial, temporal, and untrimmed multi-action spatio-temporal grounding.
Brian Chen 0001, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas 0001, Shih-Fu Chang, Rogério Feris, James R. Glass, Hilde Kuehne
CVPR6
2024 MoDE: CLIP Data Experts via Clustering
abstract
The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web- crawled data. We present Mixture of Data Experts (MoDE) and learn a system of CLIP data experts via clustering. Each data expert is trained on one data cluster, being less sensitive to false negative noises in other clusters. At inference time, we ensemble their outputs by applying weights determined through the correlation between task metadata and cluster conditions. To estimate the correlation pre-cisely, the samples in one cluster should be semantically similar, but the number of data experts should still be rea-sonable for training and inference. As such, we consider the ontology in human language and propose to use fine- grained cluster centers to represent each data expert at a coarse-grained level. Experimental studies show that four CLIP data experts on ViT-B/16 outperform the ViT-L/14 by OpenAI CLIP and OpenCLIP on zero-shot image classification but with less (<35%) training cost. Meanwhile, MoDE can train all data expert asynchronously and can flexibly include new data experts. The code is available here.
Jiawei Ma, Po-Yao Huang 0001, Saining Xie, Shang-Wen Li 0001, Luke Zettlemoyer, Shih-Fu Chang, Scott Yih, Hu Xu 0001
CVPR6
2024 RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
Ali Zare, Yulei Niu, Hammad A. Ayyubi, Shih-Fu Chang
ECCV (77)4
2024 Training-free Deep Concept Injection Enables Language Models for Video Question Answering
abstract
Recently, enabling pretrained language models (PLMs) to perform zero-shot crossmodal tasks such as video question answering has been extensively studied.A popular approach is to learn a projection network that projects visual features into the input text embedding space of a PLM, as well as feed-forward adaptation layers, with the weights of the PLM frozen.However, is it really necessary to learn such additional layers?In this paper, we make the first attempt to demonstrate that the PLM is able to perform zero-shot crossmodal tasks without any crossmodal pretraining, when the observed visual concepts are injected as both additional input text tokens and augmentation in the intermediate features within each feed-forward network for the PLM.Specifically, inputting observed visual concepts as text tokens helps to inject them through the self-attention layers in the PLM; to augment the intermediate features in a way that is compatible with the PLM, we propose to construct adaptation layers based on the intermediate representation of concepts (obtained by solely inputting them to the PLM).These two complementary injection mechanisms form the proposed Deep Concept Injection, which comprehensively enables the PLM to perceive instantly without crossmodal pretraining.Extensive empirical analysis on zero-shot video question answering, as well as visual question answering, shows Deep Concept Injection achieves competitive or even better results in both zero-shot and fine-tuning settings, compared to state-of-the-art methods that require crossmodal pretraining. Input Video 𝑣𝑣 What is the woman wearing? [mask]ℱ 𝑉𝑉
Xudong Lin 0003, Manling Li, Richard S. Zemel, Heng Ji 0001, Shih-Fu Chang
EMNLP5
2024 VIEWS: Entity-Aware News Video Captioning
abstract
Hammad Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, Shih-Fu Chang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Hammad A. Ayyubi, Tianqi Liu 0002, Arsha Nagrani, Xudong Lin 0003, Anurag Arnab, Yukun Zhu, Xuande Feng, Shih-Fu Chang
EMNLP12
2024 SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
abstract
We study the problem of procedure planning in instructional videos, which aims to make a goal-oriented sequence of action steps given partial visual state observations. The motivation of this problem is to learn a structured and plannable state and action space. Recent works succeeded in sequence modeling of steps with only sequence-level annotations accessible during training, which overlooked the roles of states in the procedures. In this work, we point out that State CHangEs MAtter (SCHEMA) for procedure planning in instructional videos. We aim to establish a more structured state space by investigating the causal relations between steps and states in procedures. Specifically, we explicitly represent each step as state changes and track the state changes in procedures. For step representation, we leveraged the commonsense knowledge in large language models (LLMs) to describe the state changes of steps via our designed chain-of-thought prompting. For state changes tracking, we align visual state observations with language state descriptions via cross-modal contrastive learning, and explicitly model the intermediate states of the procedure using LLM-generated state descriptions. Experiments on CrossTask, COIN, and NIV benchmark datasets demonstrate that our proposed SCHEMA model achieves state-of-the-art performance and obtains explainable visualizations.
Yulei Niu, Wenliang Guo, Long Chen 0016, Xudong Lin 0003, Shih-Fu Chang
ICLR5
2024 Ferret: Refer and Ground Anything Anywhere at Any Granularity
abstract
We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with an additional 130K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination.
Haoxuan You, Haotian Zhang 0005, Zhe Gan, Xianzhi Du, Bowen Zhang 0002, Liangliang Cao, Shih-Fu Chang, Yinfei Yang
ICLR8
2024 Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
Junzhang Liu, Zhecan Wang, Hammad A. Ayyubi, Haoxuan You, Christopher Thomas 0004, Rui Sun 0011, Shih-Fu Chang, Kai-Wei Chang 0001
ACM Multimedia7
2024 JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
abstract
Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts.As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors.Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.
Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun 0011, Md. Atabuzzaman, Hammad A. Ayyubi, Haoxuan You, Alvi Md. Ishmam, Kai-Wei Chang 0001, Shih-Fu Chang, Christopher Thomas 0004
NeurIPS13
2023 Video Event Extraction via Tracking Visual States of Arguments
abstract
Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the overall visual scene of each frame, ignoring fine-grained argument-level information. Inspired by the definition of events as changes of states, we propose a novel framework to detect video events by tracking the changes in the visual states of all involved arguments, which are expected to provide the most informative evidence for the extraction of video events. In order to capture the visual state changes of arguments, we decompose them into changes in pixels within objects, displacements of objects, and interactions among multiple arguments. We further propose Object State Embedding, Object Motion-aware Embedding and Argument Interaction Embedding to encode and track these changes respectively. Experiments on various video event extraction tasks demonstrate significant improvements compared to state-of-the-art models. In particular, on verb classification, we achieve 3.49% absolute gains (19.53% relative gains) in F1@5 on Video Situation Recognition. Our Code is publicly available at https://github.com/Shinetism/VStates for research purposes.
Guang Yang 0046, Manling Li, Xudong Lin 0003, Heng Ji 0001, Shih-Fu Chang
AAAI6
2023 Non-Sequential Graph Script Induction via Multimedia Grounding
abstract
Online resources such as wikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. 1 However, the scripts are always presented in a linear manner, which does not reflect the flexibility displayed by people executing tasks in real life.For example, in the CrossTask Dataset, 64.5% of consecutive step pairs are also observed in the reverse order, suggesting their ordering is not fixed.In addition, each step has an average of 2.56 frequent 2 next steps, demonstrating "branching".In this paper, we propose a new challenging task of non-sequential graph script induction, aiming to capture optional and interchangeable steps in procedural planning.To automate the induction of such graph scripts for given tasks, we propose to take advantage of loosely aligned videos of people performing the tasks.In particular, we design a multimodal framework to ground procedural videos to wikiHow textual steps and thus transform each video into an observed step path on the latent ground truth graph script.This key transformation enables us to train a script knowledge model capable of both generating explicit graph scripts for learnt tasks and predicting future steps given a partial step sequence.Our best model outperforms the strongest pure text/vision baselines by 17.52% absolute gains on F 1 @3 for next step prediction and 13.8% absolute gains on Acc@1 for partial sequence completion.Human evaluation shows our model outperforming the wikiHow linear baseline by 48.76% absolute gains in capturing sequential and non-sequential step relations.
Yu Zhou 0030, Manling Li, Xudong Lin 0003, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001
ACL (1)5
2023 Towards Fast Adaptation of Pretrained Contrastive Models for Multi-channel Video-Language Retrieval
abstract
Multi-channel video-language retrieval require models to understand information from different channels (e.g. video+question, video+speech) to correctly link a video with a textual response or query. Fortunately, contrastive multimodal models are shown to be highly effective at aligning entities in images/videos and text, e.g., CLIP [20]; text contrastive models are extensively studied recently for their strong ability of producing discriminative sentence embeddings, e.g., SimCSE [5]. However, there is not a clear way to quickly adapt these two lines to multi-channel video-language retrieval with limited data and resources. In this paper, we identify a principled model design space with two axes: how to represent videos and how to fuse video and text information. Based on categorization of recent methods, we investigate the options of representing videos using continuous feature vectors or discrete text tokens; for the fusion method, we explore the use of a multimodal transformer or a pretrained contrastive text model. We extensively evaluate the four combinations on five video-language datasets. We surprisingly find that discrete text tokens coupled with a pretrained contrastive text model yields the best performance, which can even outperform state-of-the-art on the iVQA and How2QA datasets without additional training on millions of video-text data. Further analysis shows that this is because representing videos as text tokens captures the key visual information and text tokens are naturally aligned with text models that are strong retrievers after the contrastive pretraining process. All the empirical analysis establishes a solid foundation for future research on affordable and upgradable multimodal intelligence.
Xudong Lin 0003, Simran Tiwari, Shiyuan Huang 0001, Manling Li, Zheng Shou 0001, Heng Ji 0001, Shih-Fu Chang
CVPR7
2023 Supervised Masked Knowledge Distillation for Few-Shot Transformers
abstract
Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However, under few-shot learning (FSL) settings on small datasets with only a few labeled data, ViT tends to overfit and suffers from severe performance degradation due to its absence of CNN-alike inductive bias. Previous works in FSL avoid such problem either through the help of self-supervised auxiliary losses, or through the dextile uses of label information under supervised settings. But the gap between self-supervised and supervised few-shot Transformers is still unfilled. Inspired by recent advances in self-supervised knowledge distillation and masked image modeling (MIM), we propose a novel Supervised Masked Knowledge Distillation model (SMKD) for few-shot Transformers which incorporates label information into self-distillation frameworks. Compared with previous self-supervised methods, we allow intra-class knowledge distillation on both class and patch tokens, and introduce the challenging task of masked patch tokens reconstruction across intra-class images. Experimental results on four few-shot classification benchmark datasets show that our method with simple design outperforms previous methods by a large margin and achieves a new start-of-the-art. Detailed ablation studies confirm the effectiveness of each component of our model. Code for this paper is available here: https://github.com/HL-hanlin/SMKD.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Xudong Lin 0003, Shih-Fu Chang
CVPR6
2023 DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection
abstract
Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approaches enhance few-shot generalization with the sacrifice of base-class performance, or maintain high precision in base-class detection with limited improvement in novel-class adaptation. In this paper, we point out the reason is insufficient Discriminative feature learning for all of the classes. As such, we propose a new training framework, DiGeo, to learn Geometry-aware features of interclass separation and intra-class compactness. To guide the separation of feature clusters, we derive an offline simplex equiangular tight frame (ETF) classifier whose weights serve as class centers and are maximally and equally separated. To tighten the cluster for each class, we include adaptive class-specific margins into the classification loss and encourage the features close to the class centers. Experimental studies on two few-shot benchmark datasets (VOC, COCO) and one long-tail dataset (LVIS) demonstrate that, with a single model, our method can effectively improve generalization on novel classes without hurting the detection of base classes. Our code can be found here.
Jiawei Ma, Yulei Niu, Jincheng Xu, Shiyuan Huang 0001, Guangxing Han, Shih-Fu Chang
CVPR6
2023 TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Xudong Lin 0003, Guangxing Han, Shih-Fu Chang
ICLR7
2023 PreViTS: Contrastive Pretraining with Video Tracking Supervision
abstract
Videos are a rich source for self-supervised learning (SSL) of visual representations due to the presence of natural temporal transformations of objects. However, current methods typically randomly sample video clips for learning, which results in an imperfect supervisory signal. In this work, we propose PreViTS, an SSL framework that utilizes an unsupervised tracking signal for selecting clips containing the same object, which helps better utilize temporal transformations of objects. PreViTS further uses the tracking signal to spatially constrain the frame regions to learn from and trains the model to locate meaningful objects by providing supervision on Grad-CAM attention maps. To evaluate our approach, we train a momentum contrastive (MoCo) encoder on VGG-Sound and Kinetics-400 datasets with PreViTS. Training with PreViTS outperforms representations learnt by contrastive strategy alone on video downstream tasks, obtaining state-of-the-art performance on action classification. PreViTS helps learn feature representations that are more robust to changes in background and context, as seen by experiments on datasets with background changes. Our experiment also demonstrates various visual transformation invariance captured by our model. Learning from large-scale videos with PreViTS could lead to more accurate and robust visual feature representations.
Brian Chen 0001, Ramprasaath R. Selvaraju, Shih-Fu Chang, Juan Carlos Niebles
WACV3
2022 Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment
abstract
Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern object detectors. However, the quality of proposals generated for few-shot classes using existing methods is far worse than that of many-shot classes, e.g., missing boxes for few-shot classes due to misclassification or inaccurate spatial locations with respect to true objects. To address the noisy proposal problem, we propose a novel meta-learning based FSOD model by jointly optimizing the few-shot proposal generation and fine-grained few-shot proposal classification. To improve proposal generation for few-shot classes, we propose to learn a lightweight metric-learning based prototype matching network, instead of the conventional simple linear object/nonobject classifier, e.g., used in RPN. Our non-linear classifier with the feature fusion network could improve the discriminative prototype matching and the proposal recall for few-shot classes. To improve the fine-grained few-shot proposal classification, we propose a novel attentive feature alignment method to address the spatial misalignment between the noisy proposals and few-shot classes, thus improving the performance of few-shot object detection. Meanwhile we learn a separate Faster R-CNN detection head for many-shot base classes and show strong performance of maintaining base-classes knowledge. Our model achieves state-of-the-art performance on multiple FSOD benchmarks over most of the shots and metrics.
Guangxing Han, Shiyuan Huang 0001, Jiawei Ma, Yicheng He, Shih-Fu Chang
AAAI5
2022 MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and Grounding
abstract
Recently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images is often limited to just picking the answer from a pre-defined set of options. In addition, images in the real world, especially in news, have objects that are co-referential to the text, with complementary information from both modalities. In this paper, we present a new QA evaluation benchmark with 1,384 questions over news articles that require cross-media grounding of objects in images onto text. Specifically, the task involves multi-hop questions that require reasoning over image-caption pairs to identify the grounded visual object being referred to and then predicting a span from the news body text to answer the question. In addition, we introduce a novel multimedia data augmentation framework, based on cross-media knowledge extraction and synthetic question-answer generation, to automatically augment data that can provide weak supervision for this task. We evaluate both pipeline-based and end-to-end pretraining-based multimedia QA models on our benchmark, and show that they achieve promising performance, while considerably lagging behind human performance hence leaving large room for future work on this challenging new task.
Revanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin 0003, Haoyang Wen, Jaemin Cho 0001, Lifu Huang, Mohit Bansal, Avirup Sil, Shih-Fu Chang, Alexander G. Schwing, Heng Ji 0001
AAAI10
2022 SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense Reasoning
abstract
Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as a strong reasoning ability. Recently, multimodal Transformers have made a great progress in the task of Visual Commonsense Reasoning (VCR), by jointly understanding visual objects and text tokens through layers of cross-modality attention. However, these approaches do not utilize the rich structure of the scene and the interactions between objects which are essential in answering complex commonsense questions. We propose a Scene Graph Enhanced Image-Text Learning (SGEITL) framework to incorporate visual scene graph in commonsense reasoning. In order to exploit the scene graph structure, at the model structure level, we propose a multihop graph transformer for regularizing attention interaction among hops. As for pre-training, a scene-graph-aware pre-training method is proposed to leverage structure knowledge extracted in visual scene graph. Moreover, we introduce a method to train and generate domain relevant visual scene graph using textual annotations in a weakly-supervised manner. Extensive experiments on VCR and other tasks show significant performance boost compared with the state-of-the-art methods, and prove the efficacy of each proposed component.
Zhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian, Suji Park, Yiqing Liang, Kai-Wei Chang 0001, Shih-Fu Chang
AAAI8
2022 Learning To Recognize Procedural Activities with Distant Supervision
abstract
In this paper we consider the problem of classifying fine-grained, multi-step activities (e.g., cooking different recipes, making disparate home improvements, creating various forms of arts and crafts) from long videos spanning up to several minutes. Accurately categorizing these activities requires not only recognizing the individual steps that compose the task but also capturing their temporal dependencies. This problem is dramatically different from traditional action classification, where models are typically optimized on videos that span only a few seconds and that are manually trimmed to contain simple atomic actions. While step annotations could enable the training of models to recognize the individual steps of procedural activities, existing large-scale datasets in this area do not include such segment labels due to the prohibitive cost of manually annotating temporal boundaries in long videos. To address this issue, we propose to automatically identify steps in instructional videos by leveraging the distant supervision of a textual knowledge base (wikiHow) that includes detailed descriptions of the steps needed for the execution of a wide variety of complex activities. Our method uses a language model to match noisy, automatically-transcribed speech from the video to step descriptions in the knowledge base. We demonstrate that video models trained to recognize these automatically-labeled steps (without manual supervision) yield a representation that achieves superior generalization performance on four downstream tasks: recognition of procedural activities, step classification, step forecasting and egocentric video classification.
Xudong Lin 0003, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, Lorenzo Torresani
CVPR5
2022 Few-Shot Object Detection with Fully Cross-Transformer
abstract
Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and calculate the similarity between image regions and few-shot examples for detection. However, in previous works, the interaction between the two branches is only restricted in the detection head, while leaving the remaining hundreds of layers for separate feature extraction. Inspired by the recent work on vision transformers and vision-language transformers, we propose a novel Fully Cross-Transformer based model (FCT) for FSOD by incorporating cross-transformer into both the feature backbone and detection head. The asymmetric-batched cross-attention is proposed to aggregate the key information from the two branches with different batch sizes. Our model can improve the few-shot similarity learning between the two branches by introducing the multi-level interactions. Comprehensive experiments on both PASCAL VOC and MSCOCO FSOD benchmarks demonstrate the effectiveness of our model.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Shih-Fu Chang
CVPR5
2022 Task-Adaptive Negative Envision for Few-Shot Open-Set Recognition
abstract
We study the problem of few-shot open-set recognition (FSOR), which learns a recognition system capable of both fast adaptation to new classes with limited labeled exam-ples and rejection of unknown negative samples. Traditional large-scale open-set methods have been shown in-effective for FSOR problem due to data limitation. Current FSOR methods typically calibrate few-shot closed-set clas-sifiers to be sensitive to negative samples so that they can be rejected via thresholding. However, threshold tuning is a challenging process as different FSOR tasks may require different rejection powers. In this paper, we instead propose task-adaptive negative class envision for FSOR to integrate threshold tuning into the learning process. Specifically, we augment the few-shot closed-set classifier with additional negative prototypes generated from few-shot examples. By incorporating few-shot class correlations in the negative generation process, we are able to learn dynamic rejection boundaries for FSOR tasks. Besides, we extend our method to generalized few-shot open-set recognition (GF-SOR), which requires classification on both many-shot and few-shot classes as well as rejection of negative samples. Extensive experiments on public benchmarks validate our methods on both problems.11Code available at https://github.com/shiyuanh/TANE
Shiyuan Huang 0001, Jiawei Ma, Guangxing Han, Shih-Fu Chang
CVPR4
2022 CLIP-Event: Connecting Text and Images with Event Structures
abstract
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing vision-language pretraining models primarily focus on understanding objects in images or entities in text, they often ignore the alignment at the level of events and their argument structures. In this work, we propose a contrastive learning framework to enforce vision-language pretraining models to comprehend events and associated argument (participant) roles. To achieve this, we take advantage of text information extraction technologies to obtain event structural knowledge, and utilize multiple prompt functions to contrast difficult negative descriptions by manipulating event structures. We also design an event graph alignment loss based on optimal transport to capture event argument structures. In addition, we collect a large event-rich dataset (106,875 images) for pretraining, which provides a more challenging image retrieval benchmark to assess the understanding of complicated lengthy sentences11The data and code are publicly available for research purpose in https://github.com/limanling/clip-event.. Experiments show that our zero-shot CLIP-Event outperforms the state-of-the-art supervised model in argument extraction on Multimedia Event Extraction, achieving more than 5% absolute F-score gain in event extraction, as well as significant improvements on a variety of downstream tasks under zero-shot settings.
Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin 0003, Chenguang Zhu 0001, Michael Zeng 0001, Heng Ji 0001, Shih-Fu Chang
CVPR9
2022 Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding Across Heads
Jiawei Ma, Guangxing Han, Shiyuan Huang 0001, Yuncong Yang, Shih-Fu Chang
ECCV (26)5
2022 Fine-Grained Visual Entailment
Christopher Thomas 0004, Shih-Fu Chang
ECCV (36)3
2022 Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao 0004, Noel Codella, Yu Cheng 0001, Ruochen Xu, Shih-Fu Chang, Lu Yuan 0001
ECCV (27)7
2022 Weakly-Supervised Temporal Article Grounding
abstract
Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Long Chen 0016, Yulei Niu, Brian Chen 0001, Xudong Lin 0003, Guangxing Han, Christopher Thomas 0004, Hammad A. Ayyubi, Heng Ji 0001, Shih-Fu Chang
EMNLP9
2022 Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense
abstract
Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also crossreference in-between to fully integrate and achieve comprehension of the visual scene described.Recently, various approaches have been developed and have achieved high performance on visual commonsense benchmarks.However, it is unclear whether the models really understand the visual scene and underlying commonsense knowledge due to limited evaluation data resources.To provide an indepth analysis, we present a Multimodal Evaluation (ME) pipeline to automatically generate question-answer pairs to test models' understanding of the visual scene, text, and related knowledge.We then take a step further to show that training with the ME data boosts model's performance in standard VCR evaluation.Lastly, our in-depth analysis and comparison reveal interesting findings: (1) semantically low-level information can assist learning of high-level information but not the opposite;(2) visual information is generally under utilization compared with text.
Zhecan Wang, Haoxuan You, Yicheng He, Kai-Wei Chang 0001, Shih-Fu Chang
EMNLP6
2022 Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers
abstract
Multimodal active speaker detection (ASD) methods assign a speaking/not-speaking label per individual in a video clip. ASD is critical for applications such as natural human-computer interaction, speaker diarization, and video reframing. Recent work has shown the success of transformers in multimodal settings, thus we propose a novel framework that leverages modern transformer and concatenation mechanisms to efficiently capture the interaction between audio and video modalities for ASD. We achieve mAP similar to state-of-the-art (93.0% vs 93.5%) on the AVA-ActiveSpeaker dataset. Further, our model has ~3× smaller size (15.23MB vs 49.82MB), reduced FLOPs count (11.8 vs 14.3), and lower training time (15h vs 38h). To verify our model is making predictions from the right visual cues, we computed saliency maps over input images. We found that in addition to mouth regions, the nose, cheek, and area under the eye were helpful in identifying active speakers. Our ablation study reveals that the mouth region alone achieved lower mAP (91.9% vs 93.0%) compared to full face region, supporting our hypothesis that facial expressions in addition to mouth region are useful for ASD.
Gourav Datta, Tyler Etchart, Vivek Yadav, Varsha Hedau, Pradeep Natarajan, Shih-Fu Chang
ICASSP6
2022 Few-Shot Gaze Estimation with Model Offset Predictors
abstract
Due to the variance of optical properties across different people, the performance of a person-agnostic gaze estimation model may not generalize well on a specific person. Though one may achieve better performance by training a person-specific model, it typically requires a large number of samples which is not available in real-life scenarios. Hence, few-shot gaze estimation method is preferred for the small number of samples from a target person. However, the key question is how to close the performance gap between a "few-shot" model and the "many-shot" model. In this paper, we propose to learn a person-specific offset predictor which outputs the difference between the person-agnostic model and the many-shot person-specific model with as few as one training sample. We adapt the knowledge to a new person by using the average of meta-learned offset predictors parameters as the initialization of the new offset predictor. Experiments show that the proposed few-shot person-specific model is not only closer to the corresponding many-shot person-specific model but also has better accuracy than the SOTA few-shot gaze estimation methods in multiple gaze datasets.
Jiawei Ma, Xu Zhang 0022, Yue Wu 0001, Varsha Hedau, Shih-Fu Chang
ICASSP5
2022 Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
abstract
The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets.Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei 0003, Xudong Lin 0003, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001
NeurIPS11
2022 Augmentation Invariant and Instance Spreading Feature for Softmax Embedding
abstract
Deep embedding learning plays a key role in learning discriminative feature representations, where the visually similar samples are pulled closer and dissimilar samples are pushed away in the low-dimensional embedding space. This paper studies the unsupervised embedding learning problem by learning such a representation without using any category labels. This task faces two primary challenges: mining reliable positive supervision from highly similar fine-grained classes, and generalizing to unseen testing categories. To approximate the positive concentration and negative separation properties in category-wise supervised learning, we introduce a data augmentation invariant and instance spreading feature using the instance-wise supervision. We also design two novel domain-agnostic augmentation strategies to further extend the supervision in feature space, which simulates the large batch training using a small batch size and the augmented features. To learn such a representation, we propose a novel instance-wise softmax embedding, which directly perform the optimization over the augmented instance features with the binary discrmination softmax encoding. It significantly accelerates the learning speed with much higher accuracy than existing methods, under both seen and unseen testing categories. The unsupervised embedding performs well even without pre-trained network over samples from fine-grained categories. We also develop a variant using category-wise supervision, namely category-wise softmax embedding, which achieves competitive performance over the state-of-of-the-arts, without using any auxiliary information or restrict sample mining.
Mang Ye, Jianbing Shen, Xu Zhang 0022, Pong C. Yuen, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Beyond Triplet Loss: Meta Prototypical N-Tuple Loss for Person Re-identification
abstract
Person Re-identification (ReID) aims at matching a person of interest across images. In convolutional neural network (CNN) based approaches, loss design plays a vital role in pulling closer features of the same identity and pushing far apart features of different identities. In recent years, triplet loss achieves superior performance and is predominant in ReID. However, triplet loss considers only three instances of two classes in per-query optimization (with an anchor sample as query) and it is actually equivalent to a two-class classification. There is a lack of loss design which enables the joint optimization of multiple instances (of multiple classes) within per-query optimization for person ReID. In this paper, we introduce a multi-class classification loss,i.e., N-tuple loss, to jointly consider multiple ($N$) instances for per-query optimization. This in fact aligns better with the ReID test/inference process, which conducts the ranking/comparisons among multiple instances. Furthermore, for more efficient multi-class classification, we propose a new meta prototypical N-tuple loss. With the multi-class classification incorporated, our model achieves the state-of-the-art performance on the benchmark person ReID datasets
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang
IEEE Trans. Multim.5
2021 Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
abstract
The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to one of the proposals. Existing two-stage solutions mostly focus on the grounding step, which aims to align the expressions with the proposals. In this paper, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., expression-agnostic), hoping that the proposals contain all right instances in the expression (i.e., expression-aware). Due to this mismatch, current two-stage methods suffer from a severe performance drop between detected and ground-truth proposals. To this end, we propose Ref-NMS, which is the first method to yield expression-aware proposals at the first stage. Ref-NMS regards all nouns in the expression as critical objects, and introduces a lightweight module to predict a score for aligning each box with a critical object. These scores can guide the NMS operation to filter out the boxes irrelevant to the expression, increasing the recall of critical objects, resulting in a significantly improved grounding performance. Since Ref- NMS is agnostic to the grounding step, it can be easily integrated into any state-of-the-art two-stage method. Extensive ablation studies on several backbones, benchmarks, and tasks consistently demonstrate the superiority of Ref-NMS. Codes are available at: https://github.com/ChopinSharp/ref-nms.
Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shih-Fu Chang
AAAI5
2021 InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection
abstract
Yi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, Avi Sil. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yi R. Fung 0001, Christopher Thomas 0004, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji 0001, Shih-Fu Chang, Kathy McKeown, Mohit Bansal, Avirup Sil
ACL/IJCNLP (1)6
2021 Vx2Text: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
abstract
We present VX2TEXT, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each modality is first converted into a set of language embeddings by a learnable tokenizer. This allows our approach to perform multimodal fusion in the language space, thus eliminating the need for ad-hoc cross-modal fusion modules. To address the non-differentiability of tokenization on continuous inputs (e.g., video or audio), we utilize a relaxation scheme that enables end-to-end training. Furthermore, unlike prior encoder-only models, our network includes an autoregressive decoder to generate open-ended text from the multimodal embeddings fused by the language encoder. This renders our approach fully generative and makes it directly applicable to different "video+x to text" problems without the need to design specialized network heads for each task. The proposed framework is not only conceptually simple but also remarkably effective: experiments demonstrate that our approach based on a single architecture outperforms the state-of-the-art on three videobased text-generation tasks—captioning, question answering and audio-visual scene-aware dialog.
Xudong Lin 0003, Gedas Bertasius, Jue Wang 0001, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani
CVPR4
2021 Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in Videos
abstract
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking, proposal-based matching), we tackle the problem from a novel perspective, co-grounding, with an elegant one-stage framework. We enhance the single-frame grounding accuracy by semantic attention learning and improve the cross-frame grounding consistency with co-grounding feature learning. Semantic attention learning explicitly parses referring cues in different attributes to reduce the ambiguity in the complex expression. Co-grounding feature learning boosts visual feature representations by integrating temporal correlation to reduce the ambiguity caused by scene dynamics. Experiment results demonstrate the superiority of our framework on the video grounding datasets VID and LiOTB in generating accurate and stable results across frames. Our model is also applicable to referring expression comprehension in images, illustrated by the improved performance on the RefCOCO dataset. Our project is available at https://sijiesong.github.io/co-grounding.
Sijie Song, Xudong Lin 0003, Jiaying Liu 0001, Zongming Guo, Shih-Fu Chang
CVPR5
2021 Open-Vocabulary Object Detection Using Captions
abstract
Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements. Particularly, learning more object categories typically requires proportionally more bounding box annotations. Weakly supervised and zero-shot learning techniques have been explored to scale object detectors to more categories with less supervision, but they have not been as successful and widely adopted as supervised models. In this paper, we put forth a novel formulation of the object detection problem, namely open-vocabulary object detection, which is more general, more practical, and more effective than weakly supervised and zero-shot approaches. We propose a new method to train object detectors using bounding box annotations for a limited set of object categories, as well as image-caption pairs that cover a larger variety of objects at a significantly lower cost. We show that the proposed method can detect and localize objects for which no bounding box annotation is provided during training, at a significantly higher accuracy than zero-shot approaches. Meanwhile, objects with bounding box annotation can be detected almost as accurately as supervised methods, which is significantly better than weakly supervised baselines. Accordingly, we establish a new state of the art for scalable object detection.
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, Shih-Fu Chang
CVPR4
2021 Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos
abstract
Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone, learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets.
Brian Chen 0001, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas 0001, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang
ICCV13
2021 Query Adaptive Few-Shot Object Detection with Heterogeneous Graph Convolutional Networks
abstract
Few-shot object detection (FSOD) aims to detect never-seen objects using few examples. This field sees recent improvement owing to the meta-learning techniques by learning how to match between the query image and few-shot class examples, such that the learned model can generalize to few-shot novel classes. However, currently, most of the meta-learning-based methods perform parwise matching between query image regions (usually proposals) and novel classes separately, therefore failing to take into account multiple relationships among them. In this paper, we propose a novel FSOD model using heterogeneous graph convolutional networks. Through efficient message passing among all the proposal and class nodes with three different types of edges, we could obtain context-aware proposal features and query-adaptive, multiclass-enhanced prototype representations for each class, which could help promote the pairwise matching and improve final FSOD accuracy. Extensive experimental results show that our proposed model, denoted as QA-FewDet, outperforms the current state-of-the-art approaches on the PASCAL VOC and MSCOCO FSOD benchmarks under different shots and evaluation metrics.
Guangxing Han, Yicheng He, Shiyuan Huang 0001, Jiawei Ma, Shih-Fu Chang
ICCV5
2021 Partner-Assisted Learning for Few-Shot Image Classification
abstract
Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. In this paper, we focus on the design of training strategy to obtain an elemental representation such that the prototype of each novel class can be estimated from a few labeled samples. We propose a two-stage training scheme, Partner-Assisted Learning (PAL), which first trains a Partner Encoder to model pair-wise similarities and extract features serving as soft-anchors, and then trains a Main Encoder by aligning its outputs with soft-anchors while attempting to maximize classification performance. Two alignment constraints from logit-level and feature-level are designed individually. For each few-shot task, we perform prototype classification. Our method consistently outperforms the state-of-the-art methods on four benchmarks. Detailed ablation studies of PAL are provided to justify the selection of each component involved in training.
Jiawei Ma, Hanchen Xie, Guangxing Han, Shih-Fu Chang, Aram Galstyan, Wael Abd-Almageed
ICCV4
2021 Uncertainty-Aware Few-Shot Image Classification
abstract
Few-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network has different uncertainties on its calculated similarities of different pairs. Understanding and modeling the uncertainty on the similarity could promote the exploitation of limited samples in few-shot optimization. In this work, we propose Uncertainty-Aware Few-Shot framework for image classification by modeling uncertainty of the similarities of query-support pairs and performing uncertainty-aware optimization. Particularly, we exploit such uncertainty by converting observed similarities to probabilistic representations and incorporate them to the loss for more effective optimization. In order to jointly consider the similarities between a query and the prototypes in a support set, a graph-based model is utilized to estimate the uncertainty of the pairs. Extensive experiments show our proposed method brings significant improvements on top of a strong baseline and achieves the state-of-the-art performance.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang
IJCAI5
2021 Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
abstract
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, Kai-Wei Chang 0001
NAACL-HLT5
2021 VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
abstract
We present a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures. Specifically, our Video-Audio-Text Transformer (VATT) takes raw signals as inputs and extracts multimodal representations that are rich enough to benefit a variety of downstream tasks. We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and text-to-video retrieval. Furthermore, we study a modality-agnostic single-backbone Transformer by sharing weights among the three modalities. We show that the convolution-free VATT outperforms state-of-the-art ConvNet-based architectures in the downstream tasks. Especially, VATT's vision Transformer achieves the top-1 accuracy of 82.1% on Kinetics-400, 83.6% on Kinetics-600, 72.7% on Kinetics-700, and 41.1% on Moments in Time, new records while avoiding supervised pre-training. Transferring to image classification leads to 78.7% top-1 accuracy on ImageNet compared to 64.7% by training the same Transformer from scratch, showing the generalizability of our model despite the domain gap between videos and images. VATT's audio Transformer also sets a new record on waveform-based audio event recognition by achieving the mAP of 39.4% on AudioSet without any supervised pre-training.
Hassan Akbari, Liangzhe Yuan, Rui Qian 0003, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, Boqing Gong
NeurIPS5
2021 Variational Context: Exploiting Visual and Textual Context for Grounding Referring Expressions
abstract
We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., "largest elephant standing behind baby elephant". This is a general yet challenging vision-language task since it does not only require the localization of objects, but also the multimodal comprehension of context - visual attributes (e.g., "largest", "baby") and relationships (e.g., "behind") that help to distinguish the referent from other objects, especially those of the same category. Due to the exponential complexity involved in modeling the context associated with multiple image regions, existing work oversimplifies this task to pairwise region modeling by multiple instance learning. In this paper, we propose a variational Bayesian method, called Variational Context, to solve the problem of complex context modeling in referring expression grounding. Specifically, our framework exploits the reciprocal relation between the referent and context, i.e., either of them influences estimation of the posterior distribution of the other, and thereby the search space of context can be greatly reduced. In addition to reciprocity, our framework considers the semantic information of context, i.e., the referring expression can be reproduced based on the estimated context. We also extend the model to unsupervised setting where no annotation for the referent is available. Extensive experiments on various benchmarks show consistent improvement over state-of-the-art methods in both supervised and unsupervised settings.
Yulei Niu, Hanwang Zhang, Zhiwu Lu 0001, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 General Partial Label Learning via Dual Bipartite Graph Autoencoder
abstract
We formulate a practical yet challenging problem: General Partial Label Learning (GPLL). Compared to the traditional Partial Label Learning (PLL) problem, GPLL relaxes the supervision assumption from instance-level — a label set partially labels an instance — to group-level: 1) a label set partially labels a group of instances, where the within-group instance-label link annotations are missing, and 2) cross-group links are allowed — instances in a group may be partially linked to the label set from another group. Such ambiguous group-level supervision is more practical in real-world scenarios as additional annotation on the instance-level is no longer required, e.g., face-naming in videos where the group consists of faces in a frame, labeled by a name set in the corresponding caption. In this paper, we propose a novel graph convolutional network (GCN) called Dual Bipartite Graph Autoencoder (DB-GAE) to tackle the label ambiguity challenge of GPLL. First, we exploit the cross-group correlations to represent the instance groups as dual bipartite graphs: within-group and cross-group, which reciprocally complements each other to resolve the linking ambiguities. Second, we design a GCN autoencoder to encode and decode them, where the decodings are considered as the refined results. It is worth noting that DB-GAE is self-supervised and transductive, as it only uses the group-level supervision without a separate offline training stage. Extensive experiments on two real-world datasets demonstrate that DB-GAE significantly outperforms the best baseline over absolute 0.159 F1-score and 24.8% accuracy. We further offer analysis on various levels of label ambiguities.
Brian Chen 0001, Bo Wu 0018, Alireza Zareian, Hanwang Zhang, Shih-Fu Chang
AAAI5
2020 Cross-media Structured Common Space for Multimedia Event Extraction
abstract
We introduce a new task, MultiMedia Event Extraction (M 2 E 2 ), which aims to extract events and their arguments from multimedia documents.We develop the first benchmark and collect a dataset of 245 multimedia news articles with extensively annotated events and arguments.1 We propose a novel method, Weakly Aligned Structured Embedding (WASE), that encodes structured representations of semantic information from textual and visual data into a common embedding space.The structures are aligned across modalities by employing a weakly supervised training strategy, which enables exploiting available resources without explicit cross-media annotation.Compared to unimodal state-of-the-art methods, our approach achieves 4.0% and 9.8% absolute F-score gains on text event argument role labeling and visual event extraction.Compared to stateof-the-art multimedia unstructured representations, we achieve 8.3% and 5.0% absolute Fscore gains on multimedia event extraction and argument role labeling, respectively.By utilizing images, we extract 21.4% more event mentions than traditional text-only methods.
Manling Li, Alireza Zareian, Qi Zeng 0001, Spencer Whitehead, Di Lu 0003, Heng Ji 0001, Shih-Fu Chang
ACL7
2020 Weakly Supervised Visual Semantic Parsing
abstract
Scene Graph Generation (SGG) aims to extract entities, predicates and their semantic structure from images, enabling deep understanding of visual content, with many applications such as visual reasoning and image retrieval. Nevertheless, existing SGG methods require millions of manually annotated bounding boxes for training, and are computationally inefficient, as they exhaustively process all pairs of object proposals to detect predicates. In this paper, we address those two limitations by first proposing a generalized formulation of SGG, namely Visual Semantic Parsing, which disentangles entity and predicate recognition, and enables sub-quadratic performance. Then we propose the Visual Semantic Parsing Network, VSPNet, based on a dynamic, attention-based, bipartite message passing framework that jointly infers graph nodes and edges through an iterative process. Additionally, we propose the first graph-based weakly supervised learning framework, based on a novel graph alignment algorithm, which enables training without bounding box annotations. Through extensive experiments, we show that VSPNet outperforms weakly supervised baselines significantly and approaches fully supervised performance, while being several times faster. We publicly release the source code of our method.
Alireza Zareian, Svebor Karaman, Shih-Fu Chang
CVPR3
2020 Context-Gated Convolution
Xudong Lin 0003, Lin Ma 0002, Wei Liu 0005, Shih-Fu Chang
ECCV (18)4
2020 Learning to Learn Words from Visual Scenes
Didac Suris, Dave Epstein, Heng Ji 0001, Shih-Fu Chang, Carl Vondrick
ECCV (29)4
2020 Bridging Knowledge Graphs to Generate Scene Graphs
Alireza Zareian, Svebor Karaman, Shih-Fu Chang
ECCV (23)3
2020 Learning Visual Commonsense for Robust Scene Graph Generation
Alireza Zareian, Zhecan Wang, Haoxuan You, Shih-Fu Chang
ECCV (23)4
2020 Cross-lingual Structure Transfer for Zero-resource Event Extraction
abstract
Most of the current cross-lingual transfer learning methods for Information Extraction (IE) have been only applied to name tagging. To tackle more complex tasks such as event extraction we need to transfer graph structures (event trigger linked to multiple arguments with various roles) across languages. We develop a novel share-and-transfer framework to reach this goal with three steps: (1) Convert each sentence in any language to language-universal graph structures; in this paper we explore two approaches based on universal dependency parses and complete graphs, respectively. (2) Represent each node in the graph structure with a cross-lingual word embedding so that all sentences in multiple languages can be represented with one shared semantic space. (3) Using this common semantic space, train event extractors from English training data and apply them to languages that do not have any event annotations. Experimental results on three languages (Spanish, Russian and Ukrainian) without any annotations show this framework achieves comparable performance to a state-of-the-art supervised model trained from more than 1,500 manually annotated event mentions.
Di Lu 0003, Ananya Subburathinam, Heng Ji 0001, Jonathan May, Shih-Fu Chang, Avirup Sil, Clare R. Voss
LREC5
2020 FATE/MM 20: 2nd International Workshop on Fairness, Accountability, Transparency and Ethics in MultiMedia
abstract
The series of FAT/FAccT events aim at bringing together researchers and practitioners interested in fairness, accountability, transparency and ethics of computational methods. The FATE/MM workshop focuses on addressing these issues in the Multimedia field. Multimedia computing technologies operate today at an unprecedented scale, with a growing community of scientists interested in multimedia models, tools and applications. Such continued growth has great implications not only for the scientific community, but also for the society as a whole. Typical risks of large-scale computational models include model bias and algorithmic discrimination. These risks become particularly prominent in the multimedia field, which historically has been focusing on user-centered technologies. To ensure a healthy and constructive development of the best multimedia technologies, this workshop offers a space to discuss how to develop ethical, fair, unbiased, representative, and transparent multimedia models, bringing together researchers from different areas to present computational solutions to these issues.
Xavier Alameda-Pineda, Miriam Redi, Jahna Otterbacher, Nicu Sebe, Shih-Fu Chang
ACM Multimedia5
2019 Multi-Level Multimodal Common Semantic Space for Image-Phrase Grounding
abstract
We address the problem of phrase grounding by learning a multi-level common semantic space shared by the textual and visual modalities. We exploit multiple levels of feature maps of a Deep Convolutional Neural Network, as well as contextualized word and sentence embeddings extracted from a character-based language model. Following dedicated non-linear mappings for visual features at each level, word, and sentence embeddings, we obtain multiple instantiations of our common semantic space in which comparisons between any target text and the visual content is performed with cosine similarity. We guide the model by a multi-level multimodal attention mechanism which outputs attended visual features at each level. The best level is chosen to be compared with text content for maximizing the pertinence scores of image-sentence pairs of the ground truth. Experiments conducted on three publicly available datasets show significant performance gains (20%-60% relative) over the state-of-the-art in phrase localization and set a new performance record on those datasets. We provide a detailed ablation study to show the contribution of each element of our approach and release our code on GitHub.
Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen 0001, Carl Vondrick, Shih-Fu Chang
CVPR6
2019 Multi-Granularity Generator for Temporal Action Proposal
abstract
Temporal action proposal generation is an important task, aiming to localize the video segments containing human actions in an untrimmed video. In this paper, we propose a multi-granularity generator (MGG) to perform the temporal action proposal from different granularity perspectives, relying on the video visual features equipped with the position embedding information. First, we propose to use a bilinear matching model to exploit the rich local information within the video sequence. Afterwards, two components, namely segment proposal producer (SPP) and frame actionness producer (FAP), are combined to perform the task of temporal action proposal at two distinct granularities. SPP considers the whole video in the form of feature pyramid and generates segment proposals from one coarse perspective, while FAP carries out a finer actionness evaluation for each video frame. Our proposed MGG can be trained in an end-to-end fashion. Through temporally adjusting the segment proposals with fine-grained information based on frame actionness, MGG achieves the superior performance over state-of-the-art methods on the public THUMOS-14 and ActivityNet-1.3 datasets. Moreover, we employ existing action classifiers to perform the classification of the proposals generated by MGG, leading to significant improvements compared against the competing methods for the video detection task.
Lin Ma 0002, Yifeng Zhang 0001, Wei Liu 0005, Shih-Fu Chang
CVPR5
2019 DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
abstract
Motion has shown to be useful for video understanding, where motion is typically represented by optical flow. However, computing flow from video frames is very timeconsuming. Recent works directly leverage the motion vectors and residuals readily available in the compressed video to represent motion at no cost. While this avoids flow computation, it also hurts accuracy since the motion vector is noisy and has substantially reduced resolution, which makes it a less discriminative motion representation. To remedy these issues, we propose a lightweight generator network, which reduces noises in motion vectors and captures fine motion details, achieving a more Discriminative Motion Cue (DMC) representation. Since optical flow is a more accurate motion representation, we train the DMC generator to approximate flow using a reconstruction loss and a generative adversarial loss, jointly with the downstream action classification task. Extensive evaluations on three action recognition benchmarks (HMDB-51, UCF-101, and a subset of Kinetics) confirm the effectiveness of our method. Our full system, consisting of the generator and the classifier, is coined as DMC-Net which obtains high accuracy close to that of using flow and runs two orders of magnitude faster than using optical flow at inference time.
Zheng Shou 0001, Xudong Lin 0003, Yannis Kalantidis, Laura Sevilla-Lara, Marcus Rohrbach, Shih-Fu Chang, Zhicheng Yan 0001
CVPR6
2019 Unsupervised Embedding Learning via Invariant and Spreading Instance Feature
abstract
This paper studies the unsupervised embedding learning problem, which requires an effective similarity measurement between samples in low-dimensional embedding space. Motivated by the positive concentrated and negative separated properties observed from category-wise supervised learning, we propose to utilize the instance-wise supervision to approximate these properties, which aims at learning data augmentation invariant and instance spread-out features. To achieve this goal, we propose a novel instance based softmax embedding method, which directly optimizes the `real' instance features on top of the softmax function. It achieves significantly faster learning speed and higher accuracy than all existing methods. The proposed method performs well for both seen and unseen testing categories with cosine similarity. It also achieves competitive performance even without pre-trained network over samples from fine-grained categories.
Mang Ye, Xu Zhang 0022, Pong C. Yuen, Shih-Fu Chang
CVPR4
2019 Cross-lingual Structure Transfer for Relation and Event Extraction
abstract
Ananya Subburathinam, Di Lu, Heng Ji, Jonathan May, Shih-Fu Chang, Avirup Sil, Clare Voss. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ananya Subburathinam, Di Lu 0003, Heng Ji 0001, Jonathan May, Shih-Fu Chang, Avirup Sil, Clare R. Voss
EMNLP/IJCNLP (1)5
2019 Counterfactual Critic Multi-Agent Training for Scene Graph Generation
abstract
Scene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects. For example, ``person'' on ``bike'' can help to determine the relationship ``ride'', which in turn contributes to the confidence of the two objects. However, we argue that the visual context is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes should not be penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach. CMAT is a multi-agent policy gradient method that frames objects into cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the predictions of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art performance by significant gains under various settings and metrics.
Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Xiangnan He 0001, Shiliang Pu, Shih-Fu Chang
ICCV6
2019 Multimodal Social Media Analysis for Gang Violence Prevention
Philipp Blandfort, Desmond Upton Patton, William R. Frey, Svebor Karaman, Surabhi Bhargava, Fei-Tzin Lee, Siddharth Varia, Chris Kedzie, Michael B. Gaskell, Rossano Schifanella, Kathy McKeown, Shih-Fu Chang
ICWSM12
2019 Unsupervised Rank-Preserving Hashing for Large-Scale Image Retrieval
abstract
We propose an unsupervised hashing method, exploiting a shallow neural network, that aims to produce binary codes that preserve the ranking induced by an original real-valued representation. This is motivated by the emergence of small-world graph-based approximate search methods that rely on local neighborhood ranking. We formalize the training process in an intuitive way by considering each training sample as a query and aiming to obtain a ranking of a random subset of the training set using the hash codes that is the same as the ranking using the original features. We also explore the use of a decoder to obtain an approximated reconstruction of the original features. At test time, we retrieve the most promising database samples using only the hash codes and perform re-ranking using the reconstructed features, thus allowing the complete elimination of the original real-valued features and the associated high memory cost. Experiments conducted on publicly available large-scale datasets show that our method consistently outperforms all compared state-of-the-art unsupervised hashing methods and that the reconstruction procedure can effectively boost the search accuracy with a minimal constant additional cost.
Svebor Karaman, Xudong Lin 0003, Xuefeng Hu, Shih-Fu Chang
ICMR4
2019 FAT/MM'19: 1st International Workshop on Fairness, Accountability, and Transparency in MultiMedia
abstract
The series of FAT* events aim at bringing together researchers and practitioners interested in fairness, accountability, and transparency of computational methods. The FAT/MM workshop focuses on addressing these issues in the Multimedia field. Multimedia computing technologies operate today at an unprecedented scale, with a growing community of scientists interested in multimedia models, tools and applications. Such continued growth has great implications not only for the scientific community, but also for the society as a whole. Typical risks of large-scale computational models include model bias and algorithmic discrimination. These risks become particularly prominent in the multimedia field, which historically has been focusing on user-centered technologies. To ensure a healthy and constructive development of the best multimedia technologies, this workshop offers a space to discuss how to develop fair, unbiased, representative, and transparent multimedia models, bringing together researchers from different areas to present computational solutions to these issues.
Xavier Alameda-Pineda, Miriam Redi, L. Elisa Celis, Nicu Sebe, Shih-Fu Chang
ACM Multimedia5
2019 PANEL: Challenges for Multimedia/Multimodal Research in the Next Decade
abstract
The multimedia and multi-modal community is witnessing an explosive transformation in the recent years with major societal impact. With the unprecedented deployment of multimedia devices and systems, multimedia research is critical to our abilities and prospects in advancing state-of-the-art technologies and solving real-world challenges facing the society and the nation. To respond to these challenges and further advance the frontiers of the field of multimedia, this panel will discuss the challenges and visions that may guide future research in the next ten years.
Shih-Fu Chang, Louis-Philippe Morency, Alex Hauptmann 0001, Alberto Del Bimbo, Cathal Gurrin, Hayley Hung, Heng Ji 0001, Alan F. Smeaton
ACM Multimedia1
2019 Multi-Modal Multi-Scale Deep Learning for Large-Scale Image Annotation
abstract
Image annotation aims to annotate a given image with a variable number of class labels corresponding to diverse visual concepts. In this paper, we address two main issues in large-scale image annotation: 1) how to learn a rich feature representation suitable for predicting a diverse set of visual concepts ranging from object, scene to abstract concept and 2) how to annotate an image with the optimal number of class labels. To address the first issue, we propose a novel multi-scale deep model for extracting rich and discriminative features capable of representing a wide range of visual concepts. Specifically, a novel two-branch deep neural network architecture is proposed, which comprises a very deep main network branch and a companion feature fusion network branch designed for fusing the multi-scale features computed from the main branch. The deep model is also made multi-modal by taking noisy user-provided tags as model input to complement the image input. For tackling the second issue, we introduce a label quantity prediction auxiliary task to the main label prediction task to explicitly estimate the optimal label number for a given image. Extensive experiments are carried out on two large-scale image annotation benchmark datasets, and the results show that our method significantly outperforms the state of the art.
Yulei Niu, Zhiwu Lu 0001, Ji-Rong Wen, Tao Xiang 0002, Shih-Fu Chang
IEEE Trans. Image Process.5
2019 Special Section on Multimodal Understanding of Social, Affective, and Subjective Attributes
abstract
Multimedia scientists have largely focused their research on the recognition of tangible properties of data such as objects and scenes. Recently, the field has started evolving toward the modeling of more complex properties. For example, the understanding of social, affective, and subjective attributes of visual data has attracted the attention of many research teams at the crossroads of computer vision, multimedia, and social sciences. These intangible attributes include, for example, visual beauty, video popularity, or user behavior. Multiple, diverse challenges arise when modeling such properties from multimedia data. The sections concern technical aspects such as reliable groundtruth collection, the effective learning of subjective properties, or the impact of context in subjective perception; see Refs. [2] and [3].
Xavier Alameda-Pineda, Miriam Redi, Mohammad Soleymani 0001, Nicu Sebe, Shih-Fu Chang, Samuel D. Gosling
ACM Trans. Multim. Comput. Commun. Appl.5
2018 Zero-Shot Visual Recognition Using Semantics-Preserving Adversarial Embedding Networks
abstract
We propose a novel framework called Semantics-Preserving Adversarial Embedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test images and their classes are both unseen during training. SP-AEN aims to tackle the inherent problem - semantic loss - in the prevailing family of embedding-based ZSL, where some semantics would be discarded during training if they are non-discriminative for training classes, but could become critical for recognizing test classes. Specifically, SP-AEN prevents the semantic loss by introducing an independent visual-to-semantic space embedder which disentangles the semantic space into two subspaces for the two arguably conflicting objectives: classification and reconstruction. Through adversarial learning of the two subspaces, SP-AEN can transfer the semantics from the reconstructive subspace to the discriminative one, accomplishing the improved zero-shot recognition of unseen classes. Comparing with prior works, SP-AEN can not only improve classification but also generate photo-realistic images, demonstrating the effectiveness of semantic preservation. On four popular benchmarks: CUB, AWA, SUN and aPY, SP-AEN considerably outperforms other state-of-the-art methods by an absolute performance difference of 12.2%, 9.3%, 4.0% and 3.6% in terms of harmonic mean values [62].
Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Wei Liu 0005, Shih-Fu Chang
CVPR5
2018 Grounding Referring Expressions in Images by Variational Context
abstract
We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., "largest elephant standing behind baby elephant". This is a general yet challenging vision-language task since it does not only require the localization of objects, but also the multimodal comprehension of context - visual attributes (e.g., "largest", "baby") and relationships (e.g., "behind") that help to distinguish the referent from other objects, especially those of the same category. Due to the exponential complexity involved in modeling the context associated with multiple image regions, existing work oversimplifies this task to pairwise region modeling by multiple instance learning. In this paper, we propose a variational Bayesian method, called Variational Context, to solve the problem of complex context modeling in referring expression grounding. Our model exploits the reciprocal relation between the referent and context, i.e., either of them influences estimation of the posterior distribution of the other, and thereby the search space of context can be greatly reduced. We also extend the model to unsupervised setting where no annotation for the referent is available. Extensive experiments on various benchmarks show consistent improvement over state-of-the-art methods in both supervised and unsupervised settings. The code is available at https://github.com/yuleiniu/vc/.
Hanwang Zhang, Yulei Niu, Shih-Fu Chang
CVPR3
2018 AutoLoc: Weakly-Supervised Temporal Action Localization in Untrimmed Videos
Zheng Shou 0001, Lei Zhang 0001, Kazuyuki Miyazawa, Shih-Fu Chang
ECCV (16)5
2018 Online Detection of Action Start in Untrimmed, Streaming Videos
Zheng Shou 0001, Junting Pan, Kazuyuki Miyazawa, Hassan Mansour, Anthony Vetro, Xavier Giró-i-Nieto, Shih-Fu Chang
ECCV (3)8
2018 Entity-aware Image Caption Generation
abstract
Current image captioning approaches generate descriptions which lack specific information, such as named entities that are involved in the images.In this paper we propose a new task which aims to generate informative image captions, given images and hashtags as input.We propose a simple but effective approach to tackle this problem.We first train a convolutional neural networks -long short term memory networks (CNN-LSTM) model to generate a template caption based on the input image.Then we use a knowledge graph based collective inference algorithm to fill in the template with specific named entities retrieved via the hashtags.Experiments on a new benchmark dataset collected from Flickr show that our model generates news-style image descriptions with much richer information.Our model outperforms unimodal baselines significantly with various evaluation metrics. 1
Di Lu 0003, Spencer Whitehead, Lifu Huang, Heng Ji 0001, Shih-Fu Chang
EMNLP5
2018 Incorporating Background Knowledge into Video Description Generation
abstract
Most previous efforts toward video captioning focus on generating generic descriptions, such as, "A man is talking."We collect a news video dataset to generate enriched descriptions that include important background knowledge, such as named entities and related events, which allows the user to fully understand the video content.We develop an approach that uses video meta-data to retrieve topically related news documents for a video and extracts the events and named entities from these documents.Then, given the video as well as the extracted events and entities, we generate a description using a Knowledgeaware Video Description network.The model learns to incorporate entities found in the topically related documents into the description via an entity pointer network and the generation procedure is guided by the event and entity types from the topically related documents through a knowledge gate, which is a gating mechanism added to the model's decoder that takes a one-hot vector of these types.We evaluate our approach on the new dataset of news videos we have collected, establishing the first benchmark for this dataset as well as proposing a new metric to evaluate these descriptions.
Spencer Whitehead, Heng Ji 0001, Mohit Bansal, Shih-Fu Chang, Clare R. Voss
EMNLP4
2018 Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
Victor Campos 0001, Brendan Jou, Xavier Giró-i-Nieto, Jordi Torres, Shih-Fu Chang
ICLR (Poster)5
2018 PatternNet: Visual Pattern Mining with Deep Neural Network
abstract
Visual patterns represent the discernible regularity in the visual world. They capture the essential nature of visual objects or scenes. Understanding and modeling visual patterns is a fundamental problem in visual recognition that has wide ranging applications. In this paper, we study the problem of visual pattern mining and propose a novel deep neural network architecture called PatternNet for discovering these patterns that are both discriminative and representative. The proposed PatternNet leverages the filters in the last convolution layer of a convolutional neural network to find locally consistent visual patches, and by combining these filters we can effectively discover unique visual patterns. In addition, PatternNet can discover visual patterns efficiently without performing expensive image patch sampling, and this advantage provides an order of magnitude speedup compared to most other approaches. We evaluate the proposed PatternNet subjectively by showing randomly selected visual patterns which are discovered by our method and quantitatively by performing image classification with the identified visual patterns and comparing our performance with the current state-of-the-art. We also directly evaluate the quality of the discovered visual patterns by leveraging the identified patterns as proposed objects in an image and compare with other relevant methods. Our proposed network and procedure, PatterNet, is able to outperform competing methods for the tasks described.
Hongzhi Li 0001, Joseph G. Ellis, Lei Zhang 0001, Shih-Fu Chang
ICMR4
2018 EE-USAD: ACM MM 2018Workshop on UnderstandingSubjective Attributes of Data focus on Evoked Emotions
abstract
The series of events devoted to the computational Understanding of Subjective Attributes (e.g. beauty, sentiment) of Data (USAD)provide a complementary perspective to the analysis of tangible properties (objects, scenes), which overwhelmingly covered the spectra of applications in multimedia. Partly fostered by the wide-spread usage of social media, the analysis of subjective attributes has attracted lots of attention in the recent years, and many research teams at the crossroads of multimedia, computer vision and social sciences, devoted time and effort to this topic. Among the subjective attributes there are those assessed by individuals (e.g. safety,interestingness, evoked emotions [2], memorability [3]) as well as aggregated emergent properties (such as popularity or virality [1]).This edition of the workshop (see below for the workshop's history)is devoted to the multimodal recognition of evoked emotions (EE).
Xavier Alameda-Pineda, Miriam Redi, Nicu Sebe, Shih-Fu Chang, Jiebo Luo 0001
ACM Multimedia4
2018 Low-shot Learning via Covariance-Preserving Adversarial Augmentation Networks
abstract
Deep neural networks suffer from over-fitting and catastrophic forgetting when trained with small data. One natural remedy for this problem is data augmentation, which has been recently shown to be effective. However, previous works either assume that intra-class variances can always be generalized to new classes, or employ naive generation methods to hallucinate finite examples without modeling their latent distributions. In this work, we propose Covariance-Preserving Adversarial Augmentation Networks to overcome existing limits of low-shot learning. Specifically, a novel Generative Adversarial Network is designed to model the latent distribution of each novel class given its related base counterparts. Since direct estimation on novel classes can be inductively biased, we explicitly preserve covariance information as the ``variability'' of base examples during the generation process. Empirical results show that our model can generate realistic yet diverse examples, leading to substantial improvements on the ImageNet benchmark over the state of the art.
Zheng Shou 0001, Alireza Zareian, Hanwang Zhang, Shih-Fu Chang
NeurIPS5
2018 Guest Editorial
Lamberto Ballan, Shih-Fu Chang, Gang Hua 0001, Thomas Mensink, Greg Mori, Rahul Sukthankar
Comput. Vis. Image Underst.2
2018 Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
abstract
In this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Although extensive efforts have been devoted in recent years, most existing works combined multiple video features using simple fusion strategies and neglected the utilization of inter-class semantic relationships. This paper proposes a novel unified framework that jointly exploits the feature relationships and the class relationships for improved categorization performance. Specifically, these two types of relationships are estimated and utilized by imposing regularizations in the learning process of a deep neural network (DNN). Through arming the DNN with better capability of harnessing both the feature and the class relationships, the proposed regularized DNN (rDNN) is more suitable for modeling video semantics. We show that rDNN produces better performance over several state-of-the-art approaches. Competitive results are reported on the well-known Hollywood2 and Columbia Consumer Video benchmarks. In addition, to stimulate future research on large scale video categorization, we collect and release a new benchmark dataset, called FCVID, which contains 91,223 Internet videos and 239 manually annotated categories.
Yu-Gang Jiang 0001, Zuxuan Wu, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Model-Driven Feedforward Prediction for Manipulation of Deformable Objects
abstract
Robotic manipulation of deformable objects is a difficult problem especially because of the complexity of the many different ways an object can deform. Searching such a high-dimensional state space makes it difficult to recognize, track, and manipulate deformable objects. In this paper, we introduce a predictive, model-driven approach to address this challenge, using a precomputed, simulated database of deformable object models. Mesh models of common deformable garments are simulated with the garments picked up in multiple different poses under gravity, and stored in a database for fast and efficient retrieval. To validate this approach, we developed a comprehensive pipeline for manipulating clothing as in a typical laundry task. First, the database is used for category and the pose estimation is used for a garment in an arbitrary position. A fully featured 3-D model of the garment is constructed in real time, and volumetric features are then used to obtain the most similar model in the database to predict the object category and pose. Second, the database can significantly benefit the manipulation of deformable objects via nonrigid registration, providing accurate correspondences between the reconstructed object model and the database models. Third, the accurate model simulation can also be used to optimize the trajectories for the manipulation of deformable objects, such as the folding of garments. Extensive experimental results are shown for the above tasks using a variety of different clothings. Note to Practitioners-This paper provides an open source, extensible, 3-D database for dissemination to the robotics and graphics communities. Model-driven methods are proliferating, and they need to be applied, tested, and validated in real environments. A key idea we have exploited is to have an innovative and novel use of simulation. This database will serve as infrastructure for developing advanced robotic machine learning algorithms. We want to address this machine learning idea ourselves, but we expect the dissemination of the database to other researchers with different agendas and task applications, which will bring wide progress in this area. Our proposed methods, as mentioned earlier, can be easily applied to interrelated areas. One example is that the 3-D shape-based matching algorithm can be used for other objects, such as bottles, papers, and food. After integrating with other robotic systems, the use of the robot can be easily extended to other tasks, such as making food, cleaning room, and fetching objects, to assist our daily life.
Yinxiao Li, Yan Wang 0059, Yonghao Yue, Danfei Xu, Michael Case, Shih-Fu Chang, Eitan Grinspun, Peter K. Allen
IEEE Trans Autom. Sci. Eng.6
2018 Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
abstract
Videos are inherently multimodal. This paper studies the problem of exploiting the abundant multimodal clues for improved video classification performance. We introduce a novel hybrid deep learning framework that integrates useful clues from multiple modalities, including static spatial appearance information, motion patterns within a short time window, audio information, as well as long-range temporal dynamics. More specifically, we utilize three Convolutional Neural Networks (CNNs) operating on appearance, motion, and audio signals to extract their corresponding features. We then employ a feature fusion network to derive a unified representation with an aim to capture the relationships among features. Furthermore, to exploit the long-range temporal dynamics in videos, we apply two long short-term memory (LSTM) networks with extracted appearance and motion features as inputs. Finally, we also propose refining the prediction scores by leveraging contextual relationships among video semantics. The hybrid deep learning framework is able to exploit a comprehensive set of multimodal features for video classification. Through an extensive set of experiments, we demonstrate that: 1) LSTM networks that model sequences in an explicitly recurrent manner are highly complementary to the CNN models; 2) the feature fusion network that produces a fused representation through modeling feature relationships outperforms a large set of alternative fusion strategies; and 3) the semantic context of video classes can help further refine the predictions for improved performance. Experimental results on two challenging benchmarks-the UCF-101 and the Columbia Consumer Videos (CCV)-provide strong quantitative evidence that our framework can produce promising results: 93.1% on the UCF-101 and 84.5% on the CCV, outperforming several competing methods with clear margins.
Yu-Gang Jiang 0001, Zuxuan Wu, Jinhui Tang 0001, Zechao Li, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Multim.6
2018 Editorial IEEE Transactions on Multimedia Special Section on Video Analytics: Challenges, Algorithms, and Applications
abstract
The papers in this special section focus on the topic of video analytics. Also known as video content analysis, video analytics refer to the capability of automatically analyzing video to extract knowledge/information and detect and determine temporal and spatial events. The algorithms designed for these analytics can be implemented as software on general-purpose machines, or as hardware in specialized video processing units. Video analytics is still an emerging technology with techniques that are continuously being developed to help make widespread implementation feasible in the years ahead. Such analytics has been typically used in semantic categorization and retrieval of video databases. A goal of this special issue is to focus on video analytics beyond categorization and retrieval. With increasing hardware capability and advances in algorithms used, real-time video analytics is now being used in a wide range of domains including entertainment, health-care, retail, automotive, transport, home automation, emotion analysis, aesthetics, inappropriate content detection, safety and security. For instance, video analytics are increasingly being deployed for real-time alerts in situation monitoring systems such as traffic surveillance (vehicle counting), counting people in lines (some hospitals are using this to get more nurses from a less busy department to serve thewaiting patients) and in manufacturing (for monitoring and counting). From the sensing aspect, 3-D cameras such as RGB-D and LiDAR (Light Detection and Ranging) cameras are becoming more and more affordable, enabling additional areas of research and applications, such as self-driving cars employing video analytics on LiDAR captured data for path planning as well as obstacle detection.
B. Prabhakaran 0001, Yu-Gang Jiang 0001, Hari Kalva, Shih-Fu Chang
IEEE Trans. Multim.4
2017 Localizing Actions from Video Labels and Pseudo-Annotations
Pascal Mettes, Cees Snoek, Shih-Fu Chang
BMVC3
2017 CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
abstract
Temporal action localization is an important yet challenging problem. Given a long, untrimmed video consisting of multiple action instances and complex background contents, we need not only to recognize their action categories, but also to localize the start time and end time of each instance. Many state-of-the-art systems use segment-level classifiers to select and rank proposal segments of pre-determined boundaries. However, a desirable model should move beyond segment-level and make dense predictions at a fine granularity in time to determine precise temporal boundaries. To this end, we design a novel Convolutional-De-Convolutional (CDC) network that places CDC filters on top of 3D ConvNets, which have been shown to be effective for abstracting action semantics but reduce the temporal length of the input data. The proposed CDC filter performs the required temporal upsampling and spatial downsampling operations simultaneously to predict actions at the frame-level granularity. It is unique in jointly modeling action semantics in space-time and fine-grained temporal dynamics. We train the CDC network in an end-to-end manner efficiently. Our model not only achieves superior performance in detecting actions in every frame, but also significantly boosts the precision of localizing temporal boundaries. Finally, the CDC network demonstrates a very high efficiency with the ability to process 500 frames per second on a single GPU server. Source code and trained models are available online at https://bitbucket.org/columbiadvmm/cdc.
Zheng Shou 0001, Alireza Zareian, Kazuyuki Miyazawa, Shih-Fu Chang
CVPR5
2017 Visual Translation Embedding Network for Visual Relation Detection
abstract
Visual relations, such as person ride bike and bike next to car, offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language. However, due to the challenging combinatorial complexity of modeling subject-predicate-object relation triplets, very little work has been done to localize and predict visual relations. Inspired by the recent advances in relational representation learning of knowledge bases and convolutional object detection networks, we propose a Visual Translation Embedding network (VTransE) for visual relation detection. VTransE places objects in a low-dimensional relation space where a relation can be modeled as a simple vector translation, i.e., subject + predicate ≈ object. We propose a novel feature extraction layer that enables object-relation knowledge transfer in a fully-convolutional fashion that supports training and inference in a single forward/backward pass. To the best of our knowledge, VTransE is the first end-toend relation detection network. We demonstrate the effectiveness of VTransE over other state-of-the-art methods on two large-scale datasets: Visual Relationship and Visual Genome. Note that even though VTransE is a purely visual model, it is still competitive to the Lu's multi-modal model with language priors [27].
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, Tat-Seng Chua
CVPR3
2017 Learning Discriminative and Transformation Covariant Local Feature Detectors
abstract
Robust covariant local feature detectors are important for detecting local features that are (1) discriminative of the image content and (2) can be repeatably detected at consistent locations when the image undergoes diverse transformations. Such detectors are critical for applications such as image search and scene reconstruction. Many learning-based local feature detectors address one of these two problems while overlooking the other. In this work, we propose a novel learning-based method to simultaneously address both issues. Specifically, we extend the covariant constraint proposed by Lenc and Vedaldi [8] by defining the concepts of standard patch and canonical feature and leverage these to train a novel robust covariant detector. We show that the introduction of these concepts greatly simplifies the learning stage of the covariant detector, and also makes the detector much more robust. Extensive experiments show that our method outperforms previous hand-crafted and learning-based detectors by large margins in terms of repeatability.
Xu Zhang 0022, Felix X. Yu, Svebor Karaman, Shih-Fu Chang
CVPR4
2017 PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang
ICCV4
2017 Learning Spread-Out Local Feature Descriptors
abstract
We propose a simple, yet powerful regularization technique that can be used to significantly improve both the pairwise and triplet losses in learning local feature descriptors. The idea is that in order to fully utilize the expressive power of the descriptor space, good local feature descriptors should be sufficiently “spread-out” over the space. In this work, we propose a regularization term to maximize the spread in feature descriptor inspired by the property of uniform distribution. We show that the proposed regularization with triplet loss outperforms existing Euclidean distance based descriptor learning techniques by a large margin. As an extension, the proposed regularization technique can also be used to improve image-level deep feature embedding.
Xu Zhang 0022, Felix X. Yu, Sanjiv Kumar, Shih-Fu Chang
ICCV4
2017 MUSA2: First ACM Workshop on Multimodal Understanding of Social, Affective and Subjective Attributes
abstract
Multimedia scientists have largely focused their research on the recognition of tangible properties of data, such as objects and scenes. Recently, the field has started evolving towards the modeling of more complex properties. For example, the understanding of social, affective and subjective attributes of data has attracted the attention of many research teams at the crossroads of computer vision, multimedia, and social sciences. These intangible attributes include, for example, visual beauty, video popularity, or user behavior. Multiple, diverse challenges arise when modeling such properties from multimedia data. Issues concern technical aspects such as reliable groundtruth collection, the effective learning of subjective properties, or the impact of context in subjective perception. The first edition of the ACM MM'17 MUSA2 workshop has gathered together high-quality research works focusing on the computational understanding of intangible properties from multimodal data, including visual emotions, user intent, human relationships, and personality.
Xavier Alameda-Pineda, Miriam Redi, Mohammad Soleymani 0001, Nicu Sebe, Shih-Fu Chang, Samuel D. Gosling
ACM Multimedia5
2017 LSVC2017: Large-Scale Video Classification Challenge
abstract
Recognizing visual contents in unconstrained videos has become a very important problem for many applications, such as Web video search and recommendation, smart advertising, robotics, etc. This workshop and challenge aims at exploring new challenges and approaches for large-scale video classification with large number of classes from open source videos in a realistic setting, based upon an extension of Fudan-Columbia Video Dataset (FCVID). This newly collected dataset contains over 8000 hours of video data from YouTube and Flicker, annotated into 500 categories. We hope this dataset can stimulate innovative research on this challenging and important problem.
Zuxuan Wu, Yu-Gang Jiang 0001, Larry Davis 0001, Shih-Fu Chang
ACM Multimedia4
2017 Improving Event Extraction via Multimodal Integration
abstract
In this paper, we focus on improving Event Extraction (EE) by incorporating visual knowledge with words and phrases from text documents. We first discover visual patterns from large-scale text-image pairs in a weakly-supervised manner and then propose a multimodal event extraction algorithm where the event extractor is jointly trained with textual features and visual patterns. Extensive experimental results on benchmark data sets demonstrate that the proposed multimodal EE method can achieve significantly better performance on event extraction: absolute 7.1% F-score gain on event trigger labeling and 8.5% F-score gain on event argument labeling.
Tongtao Zhang, Spencer Whitehead, Hanwang Zhang, Hongzhi Li 0001, Joseph G. Ellis, Lifu Huang, Wei Liu 0005, Heng Ji 0001, Shih-Fu Chang
ACM Multimedia9
2017 Deep Image Set Hashing
abstract
In applications involving matching of image sets, the information from multiple images must be effectively exploited to represent each set. State-of-the-art methods use probabilistic distribution or subspace to model a set and use specific distance measure to compare two sets. These methods are slow to compute and not compact to use in a large scale scenario. Learning-based hashing is often used in large scale image retrieval as they provide a compact representation of each sample and the Hamming distance can be used to efficiently compare two samples. However, most hashing methods encode each image separately and discard knowledge that multiple images in the same set represent the same object or person. We investigate the set hashing problem by combining both set representation and hashing in a single deep neural network. An image set is first passed to a CNN module to extract image features, then these features are aggregated using two types of set feature to capture both set specific and database-wide distribution information. The computed set feature is then fed into a multilayer perceptron to learn a compact binary embedding trained with triplet loss. We extensively evaluate our approach on multiple image datasets and show highly competitive performance compared to state-of-the-art methods.
Svebor Karaman, Shih-Fu Chang
WACV3
2017 A survey of multimodal sentiment analysis
Mohammad Soleymani 0001, David García 0001, Brendan Jou, Björn W. Schuller, Shih-Fu Chang, Maja Pantic
Image Vis. Comput.5
2017 Guest editorial: Multimodal sentiment analysis and mining in the wild
Mohammad Soleymani 0001, Björn W. Schuller, Shih-Fu Chang
Image Vis. Comput.3
2017 On Binary Embedding using Circulant Matrices
Felix X. Yu, Aditya Bhaskara, Sanjiv Kumar, Yunchao Gong, Shih-Fu Chang
J. Mach. Learn. Res.5
2017 Hash Bit Selection for Nearest Neighbor Search
abstract
To overcome the barrier of storage and computation when dealing with gigantic-scale data sets, compact hashing has been studied extensively to approximate the nearest neighbor search. Despite the recent advances, critical design issues remain open in how to select the right features, hashing algorithms, and/or parameter settings. In this paper, we address these by posing an optimal hash bit selection problem, in which an optimal subset of hash bits are selected from a pool of candidate bits generated by different features, algorithms, or parameters. Inspired by the optimization criteria used in existing hashing algorithms, we adopt the bit reliability and their complementarity as the selection criteria that can be carefully tailored for hashing performance in different tasks. Then, the bit selection solution is discovered by finding the best tradeoff between search accuracy and time using a modified dynamic programming method. To further reduce the computational complexity, we employ the pairwise relationship among hash bits to approximate the high-order independence property, and formulate it as an efficient quadratic programming method that is theoretically equivalent to the normalized dominant set problem in a vertex- and edge-weighted graph. Extensive large-scale experiments have been conducted under several important application scenarios of hash techniques, where our bit selection framework can achieve superior performance over both the naive selection methods and the state-of-the-art hashing algorithms, with significant accuracy gains ranging from 10% to 50%, relatively.
Xianglong Liu 0001, Junfeng He, Shih-Fu Chang
IEEE Trans. Image Process.3
2016 A Multi-media Approach to Cross-lingual Entity Knowledge Transfer
abstract
When a large-scale incident or disaster occurs, there is often a great demand for rapidly developing a system to extract detailed and new information from lowresource languages (LLs).We propose a novel approach to discover comparable documents in high-resource languages (HLs), and project Entity Discovery and Linking results from HLs documents back to LLs.We leverage a wide variety of language-independent forms from multiple data modalities, including image processing (image-to-image retrieval, visual similarity and face recognition) and sound matching.We also propose novel methods to learn entity priors from a large-scale HL corpus and knowledge base.Using Hausa and Chinese as the LLs and English as the HL, experiments show that our approach achieves 36.1% higher Hausa name tagging F-score over a costly supervised model, and 9.4% higher Chineseto-English Entity Linking accuracy over state-of-the-art.
Di Lu 0003, Xiaoman Pan, Nima Pourdamghani, Shih-Fu Chang, Heng Ji 0001, Kevin Knight
ACL (1)4
2016 Interactive Segmentation on RGBD Images via Cue Selection
abstract
Interactive image segmentation is an important problem in computer vision with many applications including image editing, object recognition and image retrieval. Most existing interactive segmentation methods only operate on color images. Until recently, very few works have been proposed to leverage depth information from low-cost sensors to improve interactive segmentation. While these methods achieve better results than color-based methods, they are still limited in either using depth as an additional color channel or simply combining depth with color in a linear way. We propose a novel interactive segmentation algorithm which can incorporate multiple feature cues like color, depth, and normals in an unified graph cut framework to leverage these cues more effectively. A key contribution of our method is that it automatically selects a single cue to be used at each pixel, based on the intuition that only one cue is necessary to determine the segmentation label locally. This is achieved by optimizing over both segmentation labels and cue labels, using terms designed to decide where both the segmentation and label cues should change. Our algorithm thus produces not only the segmentation mask but also a cue label map that indicates where each cue contributes to the final result. Extensive experiments on five large scale RGBD datasets show that our proposed algorithm performs significantly better than both other color-based and RGBD based algorithms in reducing the amount of user inputs as well as increasing segmentation accuracy.
Brian L. Price, Scott Cohen, Shih-Fu Chang
CVPR4
2016 Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
abstract
We address temporal action localization in untrimmed long videos. This is important because videos in real applications are usually unconstrained and contain multiple action instances plus video content of background scenes or other activities. To address this challenging issue, we exploit the effectiveness of deep networks in temporal action localization via three segment-based 3D ConvNets: (1) a proposal network identifies candidate segments in a long video that may contain actions, (2) a classification network learns one-vs-all action classification model to serve as initialization for the localization network, and (3) a localization network fine-tunes the learned classification network to localize each action instance. We propose a novel loss function for the localization network to explicitly consider temporal overlap and achieve high temporal localization accuracy. In the end, only the proposal network and the localization network are used during prediction. On two largescale benchmarks, our approach achieves significantly superior performances compared with other state-of-the-art systems: mAP increases from 1.7% to 7.4% on MEXaction2 and increases from 15.0% to 19.0% on THUMOS 2014.
Zheng Shou 0001, Dongang Wang, Shih-Fu Chang
CVPR3
2016 PanoSwarm: Collaborative and Synchronized Multi-Device Panoramic Photography
abstract
Taking a picture has been traditionally a one-person task. In this paper we present a novel system that allows multiple mobile devices to work collaboratively in a synchronized fashion to capture a panorama of a highly dynamic scene, creating an entirely new photography experience that encourages social interactions and teamwork. Our system contains two components: a client app that runs on all participating devices, and a server program that monitors and communicates with each device. In a capturing session, the server collects in realtime the viewfinder images of all devices and stitches them on-the-fly to create a panorama preview, which is then streamed to all devices as visual guidance. The system also allows one camera to be the host and send direct visual instructions to others to guide camera adjustment. When ready, all devices take pictures at the same time for panorama stitching. Our preliminary study suggests that the proposed system can help users capture high quality panoramas with an enjoyable teamwork experience.
Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang
IUI4
2016 New Frontiers of Large Scale Multimedia Information Retrieval
abstract
Multimedia information retrieval aims to automatically extract useful information from large collection of images, videos, and combinations with other data like text and speech. As reported in recent news, it's now possible to search information over millions or more of products with just an example image on the mobile phone. Intelligent apps are being deployed by major companies to automatically generate keywords or even captions of an image at a sophistication level that could not be imagined before. In this talk, I will review core technologies involved and discuss challenges and opportunities ahead. First, to address the complexity bottleneck when scaling up the data size, I will present extremely compact hash codes and deep learning image classification models that can reduce complexity by orders of magnitude while preserving approximate accuracy. Second, to support easy extension of recognition systems to new domains, instead of relying on fixed image categories, we introduce a new paradigm to automatically discover unique multimodal concepts and structures using large amounts of multimedia data available. Last, to support emerging applications beyond basic image categorization, I will discuss on-going efforts in understanding how images are used in expressing sentiments and emotions in online social media and how languages/cultures may influence such online multimedia communication.
Shih-Fu Chang
ICMR1
2016 SentiCart: Cartography and Geo-contextualization for Multilingual Visual Sentiment
abstract
Where in the world are pictures of cute animals or ancient architecture most shared from? And are they equally sentimentally perceived across different languages? We demonstrate a series of visualization tools, that we collectively call SentiCart, for answering such questions and navigating the landscape of how sentiment-biased images are shared around the world in multiple languages. We present visualizations using a large-scale, self-gathered geodata corpus of >1.54M geo-references coming from over 235 countries mined from >15K visual concepts over 12 languages. We also highlight several compelling data-driven findings about multilingual visual sentiment in geo-social interactions.
Brendan Jou, Margaret Yuying Qian, Shih-Fu Chang
ICMR3
2016 Complura: Exploring and Leveraging a Large-scale Multilingual Visual Sentiment Ontology
abstract
What would someone from another culture think of this photograph I just took? Would they think my picture of this "wilted flower" was also sentimentally positive or would they perceive it negatively instead? Or what if I wanted to find other photographs that are semantically related to my image as well as sentimentally sensitive, but from other cultures? In fact, this cultural and sentimental relevancy are features that we would expect of any recommender system and query expansion engine, respectively. Motivated by this, we present an online demonstration of a system called Complura. Our system implements three major functions: an interactive multilingual ontology browser, a cross-lingual image-based sentiment analyzer, and a culturally-coherent, sentiment-aware image query expansion engine. We ground our system on a multilingual visual sentiment ontology, containing over 10k sentiment-polarized visual concepts over 12 languages and over 7.3M images.
Hongyi Liu 0004, Brendan Jou, Tao Chen 0015, Mercan Topkara, Nikolaos Pappas 0002, Miriam Redi, Shih-Fu Chang
ICMR7
2016 Multilingual Visual Sentiment Concept Matching
abstract
The impact of culture in visual emotion perception has recently captured the attention of multimedia research. In this study, we provide powerful computational linguistics tools to explore, retrieve and browse a dataset of 16K multilingual affective visual concepts and 7.3M Flickr images. First, we design an effective crowdsourcing experiment to collect human judgements of sentiment connected to the visual concepts. We then use word embeddings to represent these concepts in a low dimensional vector space, allowing us to expand the meaning around concepts, and thus enabling insight about commonalities and differences among different languages. We compare a variety of concept representations through a novel evaluation task based on the notion of visual semantic relatedness. Based on these representations, we design clustering schemes to group multilingual visual concepts, and evaluate them with novel metrics based on the crowdsourced sentiment annotations as well as visual semantic relatedness. The proposed clustering framework enables us to analyze the full multilingual dataset in-depth and also show an application on a facial data subset, exploring cultural insights of portrait-related affective visual concepts.
Nikolaos Pappas 0002, Miriam Redi, Mercan Topkara, Brendan Jou, Hongyi Liu 0004, Tao Chen 0015, Shih-Fu Chang
ICMR7
2016 Watching What and How Politicians Discuss Various Topics: A Large-Scale Video Analytics UI
abstract
Accurately gauging the political atmosphere is especially difficult in this day and age, as individuals have access to a constantly growing collection of written and audiovisual news sources. This is especially true with regards to the U.S. presidential election, as there are numerous candidates, countless stories, and opinion articles discussing the merits of each particular candidate. It is therefore challenging for people to make an accurate assessment of what each candidate represents and how they would act if they were elected into office. To address this problem, we present a large-scale dataset comprised of videos of politicians speaking organized by the topics they are speaking about, and a user interface for exploring this interesting dataset. Our interface links people and events to relevant pieces of audiovisual media, and presents the desired information in a meaningful and intuitive manner. Our approach is unique by direct linking to actual speaking by politicians about specific topics, rather than links to textual quotes only. We describe the larger underlying infrastructure, a novel automated system that crawls thousands of internet news sources and 100 television news channels daily, and automatically discovers entities and indexes the content into events and topics. We examine how our user interface provides helpful and unique insights to its users, and give an example of the type of large scale trend analysis that can be performed with our system. Our online demo can be accessed at: http://www.ee.columbia.edu/dvmm/PoliticialSpeakerDemo
Emily Song, Joseph G. Ellis, Hongzhi Li 0001, Shih-Fu Chang
ICMR4
2016 Placing Broadcast News Videos in their Social Media Context Using Hashtags
abstract
With the growth of social media platforms in recent years, social media is now a major source of information and news for many people around the world. In particular the rise of hashtags have helped to build communities of discussion around particular news, topics, opinions, and ideologies. However, television news programs still provide value and are used by a vast majority of the population to obtain their news, but these videos are not easily linked to broader discussion on social media. We have built a novel pipeline that allows television news to be placed in its relevant social media context, by leveraging hashtags. In this paper, we present a method for automatically collecting television news and social media content (Twitter) and discovering the hashtags that are relevant for a TV news video. Our algorithms incorporate both the visual and text information within social media and television content, and we show that by leveraging both modalities we can improve performance over single modality approaches.
Joseph G. Ellis, Svebor Karaman, Hongzhi Li 0001, Hong Bin Shim, Shih-Fu Chang
ACM Multimedia5
2016 Tamp: A Library for Compact Deep Neural Networks with Structured Matrices
abstract
We introduce Tamp, an open source C++ library for reducing the space and time costs of deep neural network models. In particular, Tamp implements several recent works which use structured matrices to replace unstructured matrices which are often bottlenecks in neural networks.
Bingchen Gong, Brendan Jou, Felix X. Yu, Shih-Fu Chang
ACM Multimedia4
2016 Deep Cross Residual Learning for Multitask Visual Recognition
abstract
Residual learning has recently surfaced as an effective means of constructing very deep neural networks for object recognition. However, current incarnations of residual networks do not allow for the modeling and integration of complex relations between closely coupled recognition tasks or across domains. Such problems are often encountered in multimedia applications involving large-scale content recognition. We propose a novel extension of residual learning for deep networks that enables intuitive learning across multiple related tasks using cross-connections called cross-residuals. These cross-residuals connections can be viewed as a form of in-network regularization and enables greater network generalization. We show how cross-residual learning (CRL) can be integrated in multitask networks to jointly train and detect visual concepts across several tasks. We present a single multitask cross-residual network with >40% less parameters that is able to achieve competitive, or even better, detection performance on a visual sentiment concept detection problem normally requiring multiple specialized single-task networks. The resulting multitask cross-residual network also achieves better detection performance by about 10.4% over a standard multitask residual network without cross-residuals with even a small amount of cross-task weighting.
Brendan Jou, Shih-Fu Chang
ACM Multimedia2
2016 Event Specific Multimodal Pattern Mining for Knowledge Base Construction
abstract
Knowledge bases, which consist of a collection of entities, attributes, and the relations between them are widely used and important for many information retrieval tasks. Knowledge base schemas are often constructed manually using experts with specific domain knowledge for the field of interest. Once the knowledge base is generated then many tasks such as automatic content extraction and knowledge base population can be performed, which have so far been robustly studied by the Natural Language Processing community. However, the current approaches ignore visual information that could be used to build or populate these structured ontologies. Preliminary work on visual knowledge base construction only explores limited basic objects and scene relations. In this paper, we propose a novel multimodal pattern mining approach towards constructing a high-level "event" schema semi-automatically, which has the capability to extend text only methods for schema construction. We utilize a large unconstrained corpus of weakly-supervised image-caption pairs related to high-level events such as "attack" and "demonstration" to both discover visual aspects of an event, and name these visual components automatically. We compare our method with several state-of-the-art visual pattern mining approaches and demonstrate that our proposed method can achieve dramatic improvements in terms of the number of concepts discovered (33% gain), semantic consistence of visual patterns (52% gain), and correctness of pattern naming (150% gain).
Hongzhi Li 0001, Joseph G. Ellis, Heng Ji 0001, Shih-Fu Chang
ACM Multimedia4
2016 3D shape retrieval using a single depth image from low-cost sensors
abstract
Content-based 3D shape retrieval is an important problem in computer vision. Traditional retrieval interfaces require a 2D sketch or a manually designed 3D model as the query, which is difficult to specify and thus not practical in real applications. With the recent advance in low-cost 3D sensors such as Microsoft Kinect and Intel Realsense, capturing depth images that carry 3D information is fairly simple, making shape retrieval more practical and user-friendly. In this paper, we study the problem of cross-domain 3D shape retrieval using a single depth image from low-cost sensors as the query to search for similar human designed CAD models. We propose a novel method using an ensemble of autoencoders in which each autoencoder is trained to learn a compressed representation of depth views synthesize d from each database object. By viewing each autoencoder as a probabilistic model, a likelihood score can be derived as a similarity measure. A domain adaptation layer is built on top of autoencoder outputs to explicitly address the cross-domain issue (between noisy sensory data and clean 3D models) by incorporating training data of sensor depth images and their category labels in a weakly supennsed learning formulation. Experiments using real-world depth images and a large-scale CAD dataset demonstrate the effectiveness of our approach, which offers significant improvements over state-of-the-art 3D shape retrieval methods.
Yan Wang 0059, Shih-Fu Chang
WACV3
2016 Learning to Hash for Indexing Big Data - A Survey
abstract
The explosive growth in Big Data has attracted much attention in designing efficient indexing and search methods recently. In many critical applications such as large-scale search and pattern matching, finding the nearest neighbors to a query is a fundamental research problem. However, the straightforward solution using exhaustive comparison is infeasible due to the prohibitive computational complexity and memory requirement. In response, approximate nearest neighbor (ANN) search based on hashing techniques has become popular due to its promising performance in both efficiency and accuracy. Prior randomized hashing methods, e.g., locality-sensitive hashing (LSH), explore data-independent hash functions with random projections or permutations. Although having elegant theoretic guarantees on the search quality in certain metric spaces, performance of randomized hashing has been shown insufficient in many real-world applications. As a remedy, new approaches incorporating data-driven learning methods in development of advanced hash functions have emerged. Such learning-to-hash methods exploit information such as data distributions or class labels when optimizing the hash codes or functions. Importantly, the learned hash codes are able to preserve the proximity of neighboring data in the original feature spaces in the hash code spaces. The goal of this paper is to provide readers with systematic understanding of insights, pros, and cons of the emerging techniques. We provide a comprehensive survey of the learning-to-hash framework and representative techniques of various types, including unsupervised, semisupervised, and supervised. In addition, we also summarize recent hashing approaches utilizing the deep learning models. Finally, we discuss the future direction and trends of research in this area.
Jun Wang 0006, Wei Liu 0005, Sanjiv Kumar, Shih-Fu Chang
Proc. IEEE4
2015 Low-Rank Similarity Metric Learning in High Dimensions
abstract
Metric learning has become a widespreadly used tool in machine learning. To reduce expensive costs brought in by increasing dimensionality, low-rank metric learning arises as it can be more economical in storage and computation. However, existing low-rank metric learning algorithms usually adopt nonconvex objectives, and are hence sensitive to the choice of a heuristic low-rank basis. In this paper, we propose a novel low-rank metric learning algorithm to yield bilinear similarity functions. This algorithm scales linearly with input dimensionality in both space and time, therefore applicable to high-dimensional data domains. A convex objective free of heuristics is formulated by leveraging trace norm regularization to promote low-rankness. Crucially, we prove that all globally optimal metric solutions must retain a certain low-rank structure, which enables our algorithm to decompose the high-dimensional learning task into two steps: an SVD-based projection and a metric learning problem with reduced dimensionality. The latter step can be tackled efficiently through employing a linearized Alternating Direction Method of Multipliers. The efficacy of the proposed algorithm is demonstrated through experiments performed on four benchmark datasets with tens of thousands of dimensions.
Wei Liu 0005, Cun Mu, Rongrong Ji, Shiqian Ma, John R. Smith, Shih-Fu Chang
AAAI6
2015 An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinement
abstract
We describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR.
Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai
AVSS31
2015 Attributes and categories for generic instance search from one example
abstract
This paper aims for generic instance search from one example where the instance can be an arbitrary 3D object like shoes, not just near-planar and one-sided instances like buildings and logos. Firstly, we evaluate state-of-the-art instance search methods on this problem. We observe that what works for buildings loses its generality on shoes. Secondly, we propose to use automatically learned category-specific attributes to address the large appearance variations present in generic instance search. On the problem of searching among instances from the same category as the query, the category-specific attributes outperform existing approaches by a large margin. On a shoe dataset containing 6624 shoe images recorded from all viewing angles, we improve the performance from 36.73 to 56.56 using category-specific attributes. Thirdly, we extend our methods to search objects without restricting to the specifically known category. We show the combination of category-level information and the category-specific attributes is superior to combining category-level information with low-level features such as Fisher vector.
Ran Tao 0004, Arnold W. M. Smeulders, Shih-Fu Chang
CVPR3
2015 New insights into Laplacian similarity search
abstract
Graph-based computer vision applications rely critically on similarity metrics which compute the pairwise similarity between any pair of vertices on graphs. This paper investigates the fundamental design of commonly used similarity metrics, and provides new insights to guide their use in practice. In particular, we introduce a family of similarity metrics in the form of (L + αΛ)-1, where L is the graph Laplacian, Λ is a positive diagonal matrix acting as a regularizer, and α is a positive balancing factor. Such metrics respect graph topology when a is small, and reproduce well-known metrics such as hitting times and the pseudo-inverse of graph Laplacian with different regularizer Λ. This paper is the first to analyze the important impact of selecting Λ in retrieving the local cluster from a seed. We find that different Λ can lead to surprisingly complementary behaviors: Λ = D (degree matrix) can reliably extract the cluster of a query if it is sparser than surrounding clusters, while Λ = I (identity matrix) is preferred if it is denser than surrounding clusters. Since in practice there is no reliable way to determine the local density in order to select the right model, we propose a new design of Λ that automatically adapts to the local density. Experiments on image retrieval verify our theoretical arguments and confirm the benefit of the proposed metric. We expect the insights of our theory to provide guidelines for more applications in computer vision and other domains.
Xiao-Ming Wu 0003, Zhenguo Li, Shih-Fu Chang
CVPR3
2015 Cross-document Event Coreference Resolution based on Cross-media Features
abstract
In this paper we focus on a new problem of event coreference resolution across television news videos.Based on the observation that the contents from multiple data modalities are complementary, we develop a novel approach to jointly encode effective features from both closed captions and video key frames.Experiment results demonstrate that visual features provided 7.2% absolute F-score gain on stateof-the-art text based event extraction and coreference resolution.
Tongtao Zhang, Hongzhi Li 0001, Heng Ji 0001, Shih-Fu Chang
EMNLP4
2015 An Exploration of Parameter Redundancy in Deep Networks with Circulant Projections
abstract
We explore the redundancy of parameters in deep neural networks by replacing the conventional linear projection in fully-connected layers with the circulant projection. The circulant structure substantially reduces memory footprint and enables the use of the Fast Fourier Transform to speed up the computation. Considering a fully-connected neural network layer with d input nodes, and d output nodes, this method improves the time complexity from O(d2) to O(dlogd) and space complexity from O(d2) to O(d). The space savings are particularly important for modern deep convolutional neural network architectures, where fully-connected layers typically contain more than 90% of the network parameters. We further show that the gradient computation and optimization of the circulant projections can be performed very efficiently. Our experiments on three standard datasets show that the proposed approach achieves this significant gain in storage and efficiency with minimal increase in error rate compared to neural networks with unstructured projections.
Yu Cheng 0001, Felix X. Yu, Rogério Feris, Sanjiv Kumar, Alok N. Choudhary, Shih-Fu Chang
ICCV6
2015 Fast Orthogonal Projection Based on Kronecker Product
abstract
We propose a family of structured matrices to speed up orthogonal projections for high-dimensional data commonly seen in computer vision applications. In this, a structured matrix is formed by the Kronecker product of a series of smaller orthogonal matrices. This achieves O(dlogd) computational complexity and O(logd) space complexity for d-dimensional data, a drastic improvement over the standard unstructured projections whose computational and space complexities are both O(d^2). The proposed structured matrices are applicable to a number of application domains, and are faster and more compact than other structured matrices used in the past. We also introduce an efficient learning procedure for optimizing such matrices in a data dependent fashion. We demonstrate the significant advantages of the proposed approach in solving the approximate nearest neighbor (ANN) image search problem with both binary embedding and quantization. We find that the orthogonality plays a very important role in solving ANN problem, since the random orthogonal Kronecker projection has already provided promising performance. Comprehensive experiments show that the proposed approach can achieve similar or better accuracy as the existing state-of-the-art but with significantly less time and memory.
Xu Zhang 0022, Felix X. Yu, Sanjiv Kumar, Shengjin Wang, Shih-Fu Chang
ICCV6
2015 Regrasping and unfolding of garments using predictive thin shell modeling
abstract
Deformable objects such as garments are highly unstructured, making them difficult to recognize and manipulate. In this paper, we propose a novel method to teach a two-arm robot to efficiently track the states of a garment from an unknown state to a known state by iterative regrasping. The problem is formulated as a constrained weighted evaluation metric for evaluating the two desired grasping points during regrasping, which can also be used for a convergence criterion The result is then adopted as an estimation to initialize a regrasping, which is then considered as a new state for evaluation. The process stops when the predicted thin shell conclusively agrees with reconstruction. We show experimental results for regrasping a number of different garments including sweater, knitwear, pants, and leggings, etc.
Yinxiao Li, Danfei Xu, Yonghao Yue, Yan Wang 0059, Shih-Fu Chang, Eitan Grinspun, Peter K. Allen
ICRA5
2015 Encoding Concept Prototypes for Video Event Detection and Summarization
abstract
This paper proposes a new semantic video representation for few and zero example event detection and unsupervised video event summarization. Different from existing works, which obtain a semantic representation by training concepts over images or entire video clips, we propose an algorithm that learns a set of relevant frames as the concept prototypes from web video examples, without the need for frame-level annotations, and use them for representing an event video. We formulate the problem of learning the concept prototypes as seeking the frames closest to the densest region in the feature space of video frames from both positive and negative training videos of a target concept. We study the behavior of our video event representation based on concept prototypes by performing three experiments on challenging web videos from the TRECVID 2013 multimedia event detection task and the MED-summaries dataset. Our experiments establish that i) Event detection accuracy increases when mapping each video into concept prototype space. ii) Zero-example event detection increases by analyzing each frame of a video individually in concept prototype space, rather than considering the holistic videos. iii) Unsupervised video event summarization using concept prototypes is more accurate than using video-level concept detectors.
Masoud Mazloom, AmirHossein Habibian, Dong Liu 0001, Cees Snoek, Shih-Fu Chang
ICMR5
2015 2nd Workshop on Computational Models of Social Interactions: Human-Computer-Media Communication (HCMC2015)
abstract
Communicating ideas and information from and to humans is a very important subject. In our daily life, human interact with variety of entities, such as, other humans, machines, media. Constructive interactions are needed for good communication, which would result in successful outcomes, such as answering a query, learning a new skill, getting a service done, and communicating emotions. Each of these entities invokes a set of signals. Current research has focused on analyzing one entity's signals with no respect to the other entities in a unidirectional manner. The computer vision community focused on detection, classification and recognition of humans and their poses and gestures progressing onto actions, activities, and events but it does not go beyond that. The signal processing community focused on emotion recognition from facial expressions or audio or both combined. The HCI community focused on making easier interfaces for machines to ease their usage. The goal of this workshop is to bring multiple disciplines together, to process human directed signals holistically, in a bidirectional manner, rather than isolation. This workshop is positioned to display this rich domain of applications, which will provide the necessary next boost for these technologies. At the same time, it seeks to ground computational models on theory that would help achieve the technology goals. This would allow us to leverage decades of research in different fields and to spur interdisciplinary research thereby opening up new problem domains for the multimedia community.
Mohamed R. Amer, Ajay Divakaran, Shih-Fu Chang, Nicu Sebe
ACM Multimedia3
2015 Opportunities and Challenges of Industry-Academic Collaborations in Multimedia Research
abstract
This ACM MM panel aims to redefine the state of research between Academia and Industry.
Shih-Fu Chang, Matthew Cooper 0002, Denver Dash, Funda Kivran-Swaine, Li-Jia Li 0001, David A. Shamma
ACM Multimedia1
2015 Image Popularity Prediction in Social Media Using Sentiment and Context Features
abstract
Images in social networks share different destinies: some are going to become popular while others are going to be completely unnoticed. In this paper we propose to use visual sentiment features together with three novel context features to predict a concise popularity score of social images. Experiments on large scale datasets show the benefits of proposed features on the performance of image popularity prediction. Exploiting state-of-the-art sentiment features, we report a qualitative analysis of which sentiments seem to be related to good or poor popularity. To the best of our knowledge, this is the first work understanding specific visual sentiments that positively or negatively influence the eventual popularity of images.
Francesco Gelli, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo, Shih-Fu Chang
ACM Multimedia5
2015 Visual Affect Around the World: A Large-scale Multilingual Visual Sentiment Ontology
abstract
Every culture and language is unique. Our work expressly focuses on the uniqueness of culture and language in relation to human affect, specifically sentiment and emotion semantics, and how they manifest in social multimedia. We develop sets of sentiment- and emotion-polarized visual concepts by adapting semantic structures called adjective-noun pairs, originally introduced by Borth et al. (2013), but in a multilingual context. We propose a new language-dependent method for automatic discovery of these adjective-noun constructs. We show how this pipeline can be applied on a social multimedia platform for the creation of a large-scale multilingual visual sentiment concept ontology (MVSO). Unlike the flat structure in Borth et al. (2013), our unified ontology is organized hierarchically by multilingual clusters of visually detectable nouns and subclusters of emotionally biased versions of these nouns. In addition, we present an image-based prediction task to show how generalizable language-specific models are in a multilingual context. A new, publicly available dataset of >15.6K sentiment-biased visual concepts across 12 languages with language-specific detector banks, >7.36M images and their metadata is also released.
Brendan Jou, Tao Chen 0015, Nikolaos Pappas 0002, Miriam Redi, Mercan Topkara, Shih-Fu Chang
ACM Multimedia6
2015 ASM'15: The 1st International Workshop on Affect and Sentiment in Multimedia
abstract
No abstract available.
Mohammad Soleymani 0001, Yi-Hsuan Yang, Yu-Gang Jiang 0001, Shih-Fu Chang
ACM Multimedia4
2015 Large Video Event Ontology Browsing, Search and Tagging (EventNet Demo)
abstract
In this demo we present PITAGORA\footnote{Demo video available at http://bit.ly/1GgtUrN}: a mobile web contextual social network designed for the check-in area of an airport. The app provides recommendation of potential friends, local experts and targeted services. Recommendation is hybrid and combines social media analysis and collaborative filtering techniques. Users' recommendation has been evaluated through a user study with good results.
Hongliang Xu, Guangnan Ye, Dong Liu 0001, Shih-Fu Chang
ACM Multimedia5
2015 EventNet: A Large Scale Structured Concept Library for Complex Event Detection in Video
abstract
Event-specific concepts are the semantic concepts specifically designed for the events of interest, which can be used as a mid-level representation of complex events in videos. Existing methods only focus on defining event-specific concepts for a small number of pre-defined events, but cannot handle novel unseen events. This motivates us to build a large scale event-specific concept library that covers as many real-world events and their concepts as possible. Specifically, we choose WikiHow, an online forum containing a large number of how-to articles on human daily life events. We perform a coarse-to-fine event discovery process and discover 500 events from WikiHow articles. Then we use each event name as query to search YouTube and discover event-specific concepts from the tags of returned videos. After an automatic filter process, we end up with 95,321 videos and 4,490 concepts. We train a Convolutional Neural Network (CNN) model on the 95,321 videos over the 500 events, and use the model to extract deep learning feature from video content. With the learned deep learning feature, we train 4,490 binary SVM classifiers as the event-specific concept library. The concepts and events are further organized in a hierarchical structure defined by WikiHow, and the resultant concept library is called EventNet. Finally, the EventNet concept library is used to generate concept based representation of event videos. To the best of our knowledge, EventNet represents the first video event ontology that organizes events and their concepts into a semantic structure. It offers great potential for event retrieval and browsing. Extensive experiments over the zero-shot event retrieval task when no training samples are available show that the proposed EventNet concept library consistently and significantly outperforms the state-of-the-art (such as the 20K ImageNet concepts trained with CNN) by a large margin up to 207%. We will also show that EventNet structure can help users find relevant concepts for novel event queries that cannot be well handled by conventional text based semantic analysis alone. The unique two-step approach of first applying event detection models followed by detection of event-specific concepts also provides great potential to improve the efficiency and accuracy of Event Recounting since only a very small number of event-specific concept classifiers need to be fired after event detection.
Guangnan Ye, Hongliang Xu, Dong Liu 0001, Shih-Fu Chang
ACM Multimedia5
2015 Building and Using a Knowledge Graph to Combat Human Trafficking
Pedro A. Szekely, Craig A. Knoblock, Jason Slepicka, Andrew Philpot, Chengye Yin, Dipsy Kapoor, Premkumar Natarajan, Daniel Marcu, Kevin Knight, David Stallard, Subessware S. Karunamoorthy, Rajagopal Bojanapalli, Steven Minton, Brian Amanatullah, Todd Hughes, Mike Tamayo, David Flynt, Rachel Artiss, Shih-Fu Chang, Tao Chen 0015, Gerald Hiebel, Lidia Silva Ferreira
ISWC (2)20
2015 Spherical Hashing: Binary Code Embedding with Hyperspheres
abstract
Many binary code embedding schemes have been actively studied recently, since they can provide efficient similarity search, and compact data representations suitable for handling large scale image databases. Existing binary code embedding techniques encode high-dimensional data by using hyperplane-based hashing functions. In this paper we propose a novel hypersphere-based hashing function, spherical hashing, to map more spatially coherent data points into a binary code compared to hyperplane-based hashing functions. We also propose a new binary code distance function, spherical Hamming distance, tailored for our hypersphere-based binary coding scheme, and design an efficient iterative optimization process to achieve both balanced partitioning for each hash function and independence between hashing functions. Furthermore, we generalize spherical hashing to support various similarity measures defined by kernel functions. Our extensive experiments show that our spherical hashing technique significantly outperforms state-of-the-art techniques based on hyperplanes across various benchmarks with sizes ranging from one to 75 million of GIST, BoW and VLAD descriptors. The performance gains are consistent and large, up to 100 percent improvements over the second best method among tested methods. These results confirm the unique merits of using hyperspheres to encode proximity regions in high-dimensional spaces. Finally, our method is intuitive and easy to implement.
Jae-Pil Heo, Youngwoon Lee, Junfeng He, Shih-Fu Chang, Sung-Eui Yoon
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Assistive Image Comment Robot - A Novel Mid-Level Concept-Based Representation
abstract
We present a general framework and working system for predicting likely affective responses of the viewers in the social media environment after an image is posted online. Our approach emphasizes a mid-level concept representation, in which intended affects of the image publisher is characterized by a large pool of visual concepts (termed PACs) detected from image content directly instead of textual metadata, evoked viewer affects are represented by concepts (termed VACs) mined from online comments, and statistical methods are used to model the correlations among these two types of concepts. We demonstrate the utilities of such approaches by developing an end-to-end Assistive Comment Robot application, which further includes components for multi-sentence comment generation, interactive interfaces, and relevance feedback functions. Through user studies, we showed machine suggested comments were accepted by users for online posting in 90 percent of completed user sessions, while very favorable results were also observed in various dimensions (plausibility, preference, and realism) when assessing the quality of the generated image comments.
Yan-Ying Chen, Tao Chen 0015, Taikun Liu, Hong-Yuan Mark Liao, Shih-Fu Chang
IEEE Trans. Affect. Comput.5
2015 Learning Sample Specific Weights for Late Fusion
abstract
Late fusion is one of the most effective approaches to enhance recognition accuracy through combining prediction scores of multiple classifiers, each of which is trained by a specific feature or model. The existing methods generally use a fixed fusion weight for one classifier over all samples, and ignore the fact that each classifier may perform better or worse for different subsets of samples. In order to address this issue, we propose a novel sample specific late fusion (SSLF) method. Specifically, we cast late fusion into an information propagation process that diffuses the fusion weights of labeled samples to the individual unlabeled samples, and enforce positive samples to have higher fusion scores than negative samples. Upon this process, the optimal fusion weight for each sample is identified, while positive samples are pushed toward the top at the fusion score rank list to achieve better accuracy. In this paper, two SSLF methods are presented. The first method is ranking SSLF (R-SSLF), which is based on graph Laplacian with RankSVM style constraints. We formulate and solve the problem with a fast gradient projection algorithm; the second method is infinite push SSLF (I-SSLF), which combines graph Laplacian with infinite push constraints. I-SSLF is a l∞ norm constrained optimization problem and can be solved by an efficient alternating direction method of multipliers method. Extensive experiments on both large-scale image and video data sets demonstrate the effectiveness of our methods. In addition, in order to make our method scalable to support large data sets, the AnchorGraph model is employed to propagate information on a subset of samples (anchor points) and then reconstruct the entire graph to get the weights of all samples. To the best of our knowledge, this is the first method that supports learning of sample specific fusion weights for late fusion.
Kuan-Ting Lai, Dong Liu 0001, Shih-Fu Chang, Ming-Syan Chen
IEEE Trans. Image Process.3
2015 Super Fast Event Recognition in Internet Videos
abstract
Techniques for recognizing high-level events in consumer videos on the Internet have many applications. Systems that produced state-of-the-art recognition performance usually contain modules requiring extensive computation, such as the extraction of the temporal motion trajectories, which cannot be deployed on large-scale datasets. In this paper, we provide a comprehensive study on efficient methods in this area and identify technical options for super fast event recognition in Internet videos. We start from analyzing a multimodal baseline that has produced good performance on popular benchmarks, by systematically evaluating each component in terms of both computational cost and contribution to recognition accuracy. After that, we identify alternative features, classifiers, and fusion strategies that can all be efficiently computed. In addition, we also provide a study on the following interesting question: for event recognition in Internet videos, what is the minimum number of visual and audio frames needed to obtain a comparable accuracy to that of using all the frames? Results on two rigorously designed datasets indicate that similar results can be maintained by using only a small portion of the visual frames. We also find that, different from the visual frames, the soundtracks contain little redundant information and thus sampling is always harmful. Integrating all the findings, our suggested recognition system is 2,350-fold faster than a baseline approach with even higher recognition accuracies. It recognizes 20 classes on a 120-second video sequence in just 1.78 seconds, using a regular desktop computer.
Yu-Gang Jiang 0001, Qi Dai 0001, Tao Mei 0001, Yong Rui, Shih-Fu Chang
IEEE Trans. Multim.5
2015 Uploader Intent for Online Video: Typology, Inference, and Applications
abstract
We investigate automatic inference of uploader intent for online video, i.e., prediction of the reason for which a user has uploaded a particular video to the Internet. Users upload video for specific reasons, but rarely state these reasons explicitly in the video metadata. Information about the reasons motivating uploaders has the potential ultimately to benefit a wide range of application areas, including video production, video-based advertising , and video search. In this paper, we apply a combination of social-Web mining and crowdsourcing to arrive at a typology that characterizes the uploader intent of a broad range of videos. We then use a set of multimodal features, including visual semantic features, found to be indicative of uploader intent in order to classify videos automatically into uploader intent classes. We evaluate our approach on a dataset containing ca. 3K crowdsourcing-annotated videos and demonstrate its usefulness in prediction tasks relevant to common application areas.
Christoph Kofler, Subhabrata Bhattacharya, Martha A. Larson, Tao Chen 0015, Alan Hanjalic, Shih-Fu Chang
IEEE Trans. Multim.6
2014 Locally Linear Hashing for Extracting Non-linear Manifolds
abstract
Previous efforts in hashing intend to preserve data variance or pairwise affinity, but neither is adequate in capturing the manifold structures hidden in most visual data. In this paper, we tackle this problem by reconstructing the locally linear structures of manifolds in the binary Hamming space, which can be learned by locality-sensitive sparse coding. We cast the problem as a joint minimization of reconstruction error and quantization loss, and show that, despite its NP-hardness, a local optimum can be obtained efficiently via alternative optimization. Our method distinguishes itself from existing methods in its remarkable ability to extract the nearest neighbors of the query from the same manifold, instead of from the ambient space. On extensive experiments on various image benchmarks, our results improve previous state-of-the-art by 28-74% typically, and 627% on the Yale face data.
Go Irie, Zhenguo Li, Xiao-Ming Wu 0003, Shih-Fu Chang
CVPR4
2014 Video Event Detection by Inferring Temporal Instance Labels
abstract
Video event detection allows intelligent indexing of video content based on events. Traditional approaches extract features from video frames or shots, then quantize and pool the features to form a single vector representation for the entire video. Though simple and efficient, the final pooling step may lead to loss of temporally local information, which is important in indicating which part in a long video signifies presence of the event. In this work, we propose a novel instance-based video event detection approach. We represent each video as multiple 'instances', defined as video segments of different temporal intervals. The objective is to learn an instance-level event detection model based on only video-level labels. To solve this problem, we propose a large-margin formulation which treats the instance labels as hidden latent variables, and simultaneously infers the instance labels as well as the instance-level classification model. Our framework infers optimal solutions that assume positive videos have a large number of positive instances while negative videos have the fewest ones. Extensive experiments on large-scale video event datasets demonstrate significant performance gains. The proposed method is also useful in explaining the detection results by localizing the temporal segments in a video which is responsible for the positive detection.
Kuan-Ting Lai, Felix X. Yu, Ming-Syan Chen, Shih-Fu Chang
CVPR4
2014 Hash-SVM: Scalable Kernel Machines for Large-Scale Visual Classification
abstract
This paper presents a novel algorithm which uses compact hash bits to greatly improve the efficiency of non-linear kernel SVM in very large scale visual classification problems. Our key idea is to represent each sample with compact hash bits, over which an inner product is defined to serve as the surrogate of the original nonlinear kernels. Then the problem of solving the nonlinear SVM can be transformed into solving a linear SVM over the hash bits. The proposed Hash-SVM enjoys dramatic storage cost reduction owing to the compact binary representation, as well as a (sub-)linear training complexity via linear SVM. As a critical component of Hash-SVM, we propose a novel hashing scheme for arbitrary non-linear kernels via random subspace projection in reproducing kernel Hilbert space. Our comprehensive analysis reveals a well behaved theoretic bound of the deviation between the proposed hashing-based kernel approximation and the original kernel function. We also derive requirements on the hash bits for achieving a satisfactory accuracy level. Several experiments on large-scale visual classification benchmarks are conducted, including one with over 1 million images. The results show that Hash-SVM greatly reduces the computational complexity (more than ten times faster in many cases) while keeping comparable accuracies.
Yadong Mu, Gang Hua 0001, Wei Fan 0001, Shih-Fu Chang
CVPR4
2014 Recognizing Complex Events in Videos by Learning Key Static-Dynamic Evidences
Kuan-Ting Lai, Dong Liu 0001, Ming-Syan Chen, Shih-Fu Chang
ECCV (3)4
2014 Discriminative Indexing for Probabilistic Image Patch Priors
Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang
ECCV (4)4
2014 From Low-Cost Depth Sensors to CAD: Cross-Domain 3D Shape Retrieval via Regression Tree Fields
Yan Wang 0059, Jun Wang 0006, Shih-Fu Chang
ECCV (1)5
2014 Why We Watch the News: A Dataset for Exploring Sentiment in Broadcast Video News
abstract
We present a multimodal sentiment study performed on a novel collection of videos mined from broadcast and cable television news programs. To the best of our knowledge, this is the first dataset released for studying sentiment in the domain of broadcast video news. We describe our algorithm for the processing and creation of person-specific segments from news video, yielding 929 sentence-length videos, and are annotated via Amazon Mechanical Turk. The spoken transcript and the video content itself are each annotated for their expression of positive, negative or neutral sentiment. Based on these gathered user annotations, we demonstrate for news video the importance of taking into account multimodal information for sentiment prediction, and in particular, challenging previous text-based approaches that rely solely on available transcripts. We show that as much as 21.54% of the sentiment annotations for transcripts differ from their respective sentiment annotations when the video clip itself is presented. We present audio and visual classification baselines over a three-way sentiment prediction of positive, negative and neutral, as well as person-dependent versus person-independent classification influence on performance. Finally, we release the News Rover Sentiment dataset to the greater research community.
Joseph G. Ellis, Brendan Jou, Shih-Fu Chang
ICMI3
2014 Circulant Binary Embedding
abstract
Binary embedding of high-dimensional data requires long codes to preserve the discriminative power of the input space. Traditional binary coding methods often suffer from very high computation and storage costs in such a scenario. To address this problem, we propose Circulant Binary Embedding (CBE) which generates binary codes by projecting the data with a circulant matrix. The circulant structure enables the use of Fast Fourier Transformation to speed up the computation. Compared to methods that use unstructured matrices, the proposed method improves the time complexity from \mathcalO(d^2) to \mathcalO(d\logd), and the space complexity from \mathcalO(d^2) to \mathcalO(d) where d is the input dimensionality. We also propose a novel time-frequency alternating optimization to learn data-dependent circulant projections, which alternatively minimizes the objective in original and Fourier domains. We show by extensive experiments that the proposed approach gives much better performance than the state-of-the-art approaches for fixed time, and provides much faster computation with no performance degradation for fixed number of bits.
Felix X. Yu, Sanjiv Kumar, Yunchao Gong, Shih-Fu Chang
ICML4
2014 Real-time pose estimation of deformable objects using a volumetric approach
abstract
Pose estimation of deformable objects is a fundamental and challenging problem in robotics. We present a novel solution to this problem by first reconstructing a 3D model of the object from a low-cost depth sensor such as Kinect, and then searching a database of simulated models in different poses to predict the pose. Given noisy depth images from 360-degree views of the target object acquired from the Kinect sensor, we reconstruct a smooth 3D model of the object using depth image segmentation and volumetric fusion. Then with an efficient feature extraction and matching scheme, we search the database, which contains a large number of deformable objects in different poses, to obtain the most similar model, whose pose is then adopted as the prediction. Extensive experiments demonstrate better accuracy and orders of magnitude speed-up compared to our previous work. An additional benefit of our method is that it produces a high-quality mesh model and camera pose, which is necessary for other tasks such as regrasping and object manipulation.
Yinxiao Li, Yan Wang 0059, Michael Case, Shih-Fu Chang, Peter K. Allen
IROS4
2014 Predicting Evoked Emotions in Video
abstract
Understanding how human emotion is evoked from visual content is a task that we as people do every day, but machines have not yet mastered. In this work we address the problem of predicting the intended evoked emotion at given points within movie trailers. Movie Trailers are carefully curated to elicit distinct and specific emotional responses from viewers, and are therefore well-suited for emotion prediction. However, current emotion recognition systems struggle to bridge the "affective gap", which refers to the difficulty in modeling high-level human emotions with low-level audio and visual features. To address this problem, we propose a mid-level concept feature, which is based on detectable movie shot concepts which we believe to be tied closely to emotions. Examples of these concepts are "Fight", "Rock Music", and "Kiss". We also create 2 datasets, the first with shot-level concept annotations for learning our concept detectors, and a separate, second dataset with emotion annotations taken throughout the trailers using the two dimensional arousal and valence model for emotion annotation. We report the performance of our concept detectors, and show that by using the output of these detectors as a mid-level representation for the movie shots we are able to more accurately predict the evoked emotion throughout a trailer than by using low-level features.
Joseph G. Ellis, Wan-Yi Sabrina Lin, Ching-Yung Lin, Shih-Fu Chang
ISM4
2014 Minimally Needed Evidence for Complex Event Recognition in Unconstrained Videos
abstract
This paper addresses the fundamental question -- How do humans recognize complex events in videos? Normally, humans view videos in a sequential manner. We hypothesize that humans can make high-level inference such as an event is present or not in a video, by looking at a very small number of frames not necessarily in a linear order. We attempt to verify this cognitive capability of humans and to discover the Minimally Needed Evidence (MNE) for each event. To this end, we introduce an online game based event quiz facilitating selection of minimal evidence required by humans to judge the presence or absence of a complex event in an open source video. Each video is divided into a set of temporally coherent microshots (1.5 secs in length) which are revealed only on player request. The player's task is to identify the positive and negative occurrences of the given target event with minimal number of requests to reveal evidence. Incentives are given to players for correct identification with the minimal number of requests.
Subhabrata Bhattacharya, Felix X. Yu, Shih-Fu Chang
ICMR3
2014 Predicting Viewer Affective Comments Based on Image Content in Social Media
abstract
Visual sentiment analysis is getting increasing attention because of the rapidly growing amount of images in online social interactions and several emerging applications such as online propaganda and advertisement. Recent studies have shown promising progress in analyzing visual affect concepts intended by the media content publisher. In contrast, this paper focuses on predicting what viewer affect concepts will be triggered when the image is perceived by the viewers. For example, given an image tagged with "yummy food," the viewers are likely to comment "delicious" and "hungry," which we refer to as viewer affect concepts (VAC) in this paper. To the best of our knowledge, this is the first work explicitly distinguishing intended publisher affect concepts and induced viewer affect concepts associated with social visual content, and aiming at understanding their correlations. We present around 400 VACs automatically mined from million-scale real user comments associated with images in social media. Furthermore, we propose an automatic visual based approach to predict VACs by first detecting publisher affect concepts in image content and then applying statistical correlations between such publisher affect concepts and the VACs. We demonstrate major benefits of the proposed methods in several real-world tasks - recommending images to invoke certain target VACs among viewers, increasing the accuracy of predicting VACs by 20.1% and finally developing a social assistant tool that may suggest plausible, content-specific and desirable comments when users view new images.
Yan-Ying Chen, Tao Chen 0015, Winston H. Hsu, Hong-Yuan Mark Liao, Shih-Fu Chang
ICMR5
2014 Event-Driven Semantic Concept Discovery by Exploiting Weakly Tagged Internet Images
abstract
Analysis and detection of complex events in videos require a semantic representation of the video content. Existing video semantic representation methods typically require users to pre-define an exhaustive concept lexicon and manually annotate the presence of the concepts in each video, which is infeasible for real-world video event detection problems. In this paper, we propose an automatic semantic concept discovery scheme by exploiting Internet images and their associated tags. Given a target event and its textual descriptions, we crawl a collection of images and their associated tags by performing text based image search using the noun and verb pairs extracted from the event textual descriptions. The system first identifies the candidate concepts for an event by measuring whether a tag is a meaningful word and visually detectable. Then a concept visual model is built for each candidate concept using a SVM classifier with probabilistic output. Finally, the concept models are applied to generate concept based video representations. We use the TRECVID Multimedia Event Detection (MED) 2013 as our video test set and crawl 400K Flickr images to automatically discover 2, 000 visual concepts. We show significant performance gains of the proposed concept discovery method over different video event detection tasks including supervised event modeling over concept space and semantic based zero-shot retrieval without training examples. Importantly, we show the proposed method of automatic concept discovery outperforms other well-known concept library construction approaches such as Classemes and ImageNet by a large margin (228%) in zero-shot event retrieval. Finally, subjective evaluation by humans also confirms clear superiority of the proposed method in discovering concepts for event representation.
Yin Cui, Guangnan Ye, Dong Liu 0001, Shih-Fu Chang
ICMR5
2014 Object-Based Visual Sentiment Concept Analysis and Application
abstract
This paper studies the problem of modeling object-based visual concepts such as "crazy car" and "shy dog" with a goal to extract emotion related information from social multimedia content. We focus on detecting such adjective-noun pairs because of their strong co-occurrence relation with image tags about emotions. This problem is very challenging due to the highly subjective nature of the adjectives like "crazy" and "shy" and the ambiguity associated with the annotations. However, associating adjectives with concrete physical nouns makes the combined visual concepts more detectable and tractable. We propose a hierarchical system to handle the concept classification in an object specific manner and decompose the hard problem into object localization and sentiment related concept modeling. In order to resolve the ambiguity of concepts we propose a novel classification approach by modeling the concept similarity, leveraging on online commonsense knowledgebase. The proposed framework also allows us to interpret the classifiers by discovering discriminative features. The comparisons between our method and several baselines show great improvement in classification performance. We further demonstrate the power of the proposed system with a few novel applications such as sentiment-aware music slide shows of personal albums.
Tao Chen 0015, Felix X. Yu, Yin Cui, Yan-Ying Chen, Shih-Fu Chang
ACM Multimedia6
2014 Predicting Viewer Perceived Emotions in Animated GIFs
abstract
Animated GIFs are everywhere on the Web. Our work focuses on the computational prediction of emotions perceived by viewers after they are shown animated GIF images. We evaluate our results on a dataset of over 3,800 animated GIFs gathered from MIT's GIFGIF platform, each with scores for 17 discrete emotions aggregated from over 2.5M user annotations - the first computational evaluation of its kind for content-based prediction on animated GIFs to our knowledge. In addition, we advocate a conceptual paradigm in emotion prediction that shows delineating distinct types of emotion is important and is useful to be concrete about the emotion target. One of our objectives is to systematically compare different types of content features for emotion prediction, including low-level, aesthetics, semantic and face features. We also formulate a multi-task regression problem to evaluate whether viewer perceived emotion prediction can benefit from jointly learning across emotion classes compared to disjoint, independent learning.
Brendan Jou, Subhabrata Bhattacharya, Shih-Fu Chang
ACM Multimedia3
2014 Modeling Attributes from Category-Attribute Proportions
abstract
Attribute-based representation has been widely used in visual recognition and retrieval due to its interpretability and cross-category generalization properties. However, classic attribute learning requires manually labeling attributes on the images, which is very expensive, and not scalable. In this paper, we propose to model attributes from category-attribute proportions. The proposed framework can model attributes without attribute labels on the images. Specifically, given a multi-class image datasets with N categories, we model an attribute, based on an N-dimensional category-attribute proportion vector, where each element of the vector characterizes the proportion of images in the corresponding category having the attribute. The attribute learning can be formulated as a learning from label proportion (LLP) problem. Our method is based on a newly proposed machine learning algorithm called $\propto$SVM. Finding the category-attribute proportions is much easier than manually labeling images, but it is still not a trivial task. We further propose to estimate the proportions from multiple modalities such as human commonsense knowledge, NLP tools, and other domain knowledge. The value of the proposed approach is demonstrated by various applications including modeling animal attributes, visual sentiment attributes, and scene attributes.
Felix X. Yu, Liangliang Cao, Michele Merler, Noel Codella, Tao Chen 0015, John R. Smith, Shih-Fu Chang
ACM Multimedia7
2014 Scalable Visual Instance Mining with Threads of Features
abstract
We address the problem of visual instance mining, which is to extract frequently appearing visual instances automatically from a multimedia collection. We propose a scalable mining method by exploiting Thread of Features (ToF). Specifically, ToF, a compact representation that links consistent features across images, is extracted to reduce noises, discover patterns, and speed up processing. Various instances, especially small ones, can be discovered by exploiting correlated ToFs. Our approach is significantly more effective than other methods in mining small instances. At the same time, it is also more efficient by requiring much fewer hash tables. We compared with several state-of-the-art methods on two fully annotated datasets: MQA and Oxford, showing large performance gain in mining (especially small) visual instances. We also run our method on another Flickr dataset with one million images for scalability test. Two applications, instance search and multimedia summarization, are developed from the novel perspective of instance mining, showing great potential of our method in multimedia analysis.
Wei Zhang 0031, Hongzhi Li 0001, Chong-Wah Ngo, Shih-Fu Chang
ACM Multimedia4
2014 Discrete Graph Hashing
Wei Liu 0005, Cun Mu, Sanjiv Kumar, Shih-Fu Chang
NIPS4
2014 Discovering joint audio-visual codewords for video event detection
I-Hong Jhuo, Guangnan Ye, Shenghua Gao, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang
Mach. Vis. Appl.7
2014 Mixed image-keyword query adaptive hashing over multilabel images
abstract
This article defines a new hashing task motivated by real-world applications in content-based image retrieval, that is, effective data indexing and retrieval given mixed query (query image together with user-provided keywords). Our work is distinguished from state-of-the-art hashing research by two unique features: (1) Unlike conventional image retrieval systems, the input query is a combination of an exemplar image and several descriptive keywords, and (2) the input image data are often associated with multiple labels. It is an assumption that is more consistent with the realistic scenarios. The mixed image-keyword query significantly extends traditional image-based query and better explicates the user intention. Meanwhile it complicates semantics-based indexing on the multilabel data. Though several existing hashing methods can be adapted to solve the indexing task, unfortunately they all prove to suffer from low effectiveness. To enhance the hashing efficiency, we propose a novel scheme “boosted shared hashing”. Unlike prior works that learn the hashing functions on either all image labels or a single label, we observe that the hashing function can be more effective if it is designed to index over an optimal label subset. In other words, the association between labels and hash bits are moderately sparse. The sparsity of the bit-label association indicates greatly reduced computation and storage complexities for indexing a new sample, since only limited number of hashing functions will become active for the specific sample. We develop a Boosting style algorithm for simultaneously optimizing both the optimal label subsets and hashing functions in a unified formulation, and further propose a query-adaptive retrieval mechanism based on hash bit selection for mixed queries, no matter whether or not the query words exist in the training data. Moreover, we show that the proposed method can be easily extended to the case where the data similarity is gauged by nonlinear kernel functions. Extensive experiments are conducted on standard image benchmarks like CIFAR-10, NUS-WIDE and a-TRECVID. The results validate both the sparsity of the bit-label association and the convergence of the proposed algorithm, and demonstrate that the proposed hashing scheme achieves substantially superior performances over state-of-the-art methods under the same hash bit budget.
Xianglong Liu 0001, Yadong Mu, Bo Lang, Shih-Fu Chang
ACM Trans. Multim. Comput. Commun. Appl.4
2013 Robust Object Co-detection
abstract
Object co-detection aims at simultaneous detection of objects of the same category from a pool of related images by exploiting consistent visual patterns present in candidate objects in the images. The related image set may contain a mixture of annotated objects and candidate objects generated by automatic detectors. Co-detection differs from the conventional object detection paradigm in which detection over each test image is determined one-by-one independently without taking advantage of common patterns in the data pool. In this paper, we propose a novel, robust approach to dramatically enhance co-detection by extracting a shared low-rank representation of the object instances in multiple feature spaces. The idea is analogous to that of the well-known Robust PCA~\cite{rpca}, but has not been explored in object co-detection so far. The representation is based on a linear reconstruction over the entire data set and the low-rank approach enables effective removal of noisy and outlier samples. The extracted low-rank representation can be used to detect the target objects by spectral clustering. Extensive experiments over diverse benchmark datasets demonstrate consistent and significant performance gains of the proposed method over the state-of-the-art object co-detection method and the generic object detection methods without co-detection formulations.
Dong Liu 0001, Brendan Jou, Mojun Zhu, Anni Cai, Shih-Fu Chang
CVPR6
2013 A Bayesian Approach to Multimodal Visual Dictionary Learning
abstract
Despite significant progress, most existing visual dictionary learning methods rely on image descriptors alone or together with class labels. However, Web images are often associated with text data which may carry substantial information regarding image semantics, and may be exploited for visual dictionary learning. This paper explores this idea by leveraging relational information between image descriptors and textual words via co-clustering, in addition to information of image descriptors. Existing co-clustering methods are not optimal for this problem because they ignore the structure of image descriptors in the continuous space, which is crucial for capturing visual characteristics of images. We propose a novel Bayesian co-clustering model to jointly estimate the underlying distributions of the continuous image descriptors as well as the relationship between such distributions and the textual words through a unified Bayesian inference. Extensive experiments on image categorization and retrieval have validated the substantial value of the proposed joint modeling in improving visual dictionary learning, where our model shows superior performance over several recent methods.
Go Irie, Dong Liu 0001, Zhenguo Li, Shih-Fu Chang
CVPR4
2013 Hash Bit Selection: A Unified Solution for Selection Problems in Hashing
abstract
Recent years have witnessed the active development of hashing techniques for nearest neighbor search over big datasets. However, to apply hashing techniques successfully, there are several important issues remaining open in selecting features, hashing algorithms, parameter settings, kernels, etc. In this work, we unify all these selection problems into a hash bit selection framework, i.e., selecting the most informative hash bits from a pool of candidate bits generated by different types of hashing methods using different feature spaces and/or parameter settings, etc. We represent the bit pool as a vertex- and edge-weighted graph with the candidate bits as vertices. The vertex weight represents the bit quality in terms of similarity preservation, and the edge weight reflects independence (non-redundancy) between bits. Then we formulate the bit selection problem as quadratic programming on the graph, and solve it efficiently by replicator dynamics. Moreover, a theoretical study is provided to reveal a very interesting insight: the selected bits actually are the normalized dominant set of the candidate bit graph. We conducted extensive large-scale experiments for three important application scenarios of hash techniques, i.e., hashing with multiple features, multiple hashing algorithms, and multiple bit hashing. We demonstrate that our bit selection approach can achieve superior performance over both naive selection methods and state-of-the-art hashing methods under each scenario, with significant accuracy gains ranging from 10% to 50% relatively.
Xianglong Liu 0001, Junfeng He, Bo Lang, Shih-Fu Chang
CVPR4
2013 Sample-Specific Late Fusion for Visual Category Recognition
abstract
Late fusion addresses the problem of combining the prediction scores of multiple classifiers, in which each score is predicted by a classifier trained with a specific feature. However, the existing methods generally use a fixed fusion weight for all the scores of a classifier, and thus fail to optimally determine the fusion weight for the individual samples. In this paper, we propose a sample-specific late fusion method to address this issue. Specifically, we cast the problem into an information propagation process which propagates the fusion weights learned on the labeled samples to individual unlabeled samples, while enforcing that positive samples have higher fusion scores than negative samples. In this process, we identify the optimal fusion weights for each sample and push positive samples to top positions in the fusion score rank list. We formulate our problem as a L∞norm constrained optimization problem and apply the Alternating Direction Method of Multipliers for the optimization. Extensive experiment results on various visual categorization tasks show that the proposed method consistently and significantly beats the state-of-the-art late fusion methods. To the best knowledge, this is the first method supporting sample-specific fusion weight learning.
Dong Liu 0001, Kuan-Ting Lai, Guangnan Ye, Ming-Syan Chen, Shih-Fu Chang
CVPR5
2013 Label Propagation from ImageNet to 3D Point Clouds
abstract
Recent years have witnessed a growing interest in understanding the semantics of point clouds in a wide variety of applications. However, point cloud labeling remains an open problem, due to the difficulty in acquiring sufficient 3D point labels towards training effective classifiers. In this paper, we overcome this challenge by utilizing the existing massive 2D semantic labeled datasets from decade-long community efforts, such as Image Net and Label Me, and a novel ``cross-domain'' label propagation approach. Our proposed method consists of two major novel components, Exemplar SVM based label propagation, which effectively addresses the cross-domain issue, and a graphical model based contextual refinement incorporating 3D constraints. Most importantly, the entire process does not require any training data from the target scenes, also with good scalability towards large scale applications. We evaluate our approach on the well-known Cornell Point Cloud Dataset, achieving much greater efficiency and comparable accuracy even without any 3D training data. Our approach shows further major gains in accuracy when the training data from the target scenes is used, outperforming state-of-the-art approaches with far better efficiency.
Yan Wang 0059, Rongrong Ji, Shih-Fu Chang
CVPR3
2013 Designing Category-Level Attributes for Discriminative Visual Recognition
abstract
Attribute-based representation has shown great promises for visual recognition due to its intuitive interpretation and cross-category generalization property. However, human efforts are usually involved in the attribute designing process, making the representation costly to obtain. In this paper, we propose a novel formulation to automatically design discriminative "category-level attributes", which can be efficiently encoded by a compact category-attribute matrix. The formulation allows us to achieve intuitive and critical design criteria (category-separability, learn ability) in a principled way. The designed attributes can be used for tasks of cross-category knowledge transfer, achieving superior performance over well-known attribute dataset Animals with Attributes (AwA) and a large-scale ILSVRC2010 dataset (1.2M images). This approach also leads to state-of-the-art performance on the zero-shot learning task on AwA.
Felix X. Yu, Liangliang Cao, Rogério Feris, John R. Smith, Shih-Fu Chang
CVPR5
2013 Distributed Low-Rank Subspace Segmentation
abstract
Vision problems ranging from image clustering to motion segmentation to semi-supervised learning can naturally be framed as subspace segmentation problems, in which one aims to recover multiple low-dimensional subspaces from noisy and corrupted input data. Low-Rank Representation (LRR), a convex formulation of the subspace segmentation problem, is provably and empirically accurate on small problems but does not scale to the massive sizes of modern vision datasets. Moreover, past work aimed at scaling up low-rank matrix factorization is not applicable to LRR given its non-decomposable constraints. In this work, we propose a novel divide-and-conquer algorithm for large-scale subspace segmentation that can cope with LRR's non-decomposable constraints and maintains LRR's strong recovery guarantees. This has immediate implications for the scalability of subspace segmentation, which we demonstrate on a benchmark face recognition dataset and in simulations. We then introduce novel applications of LRR-based subspace segmentation to large-scale semi-supervised learning for multimedia event detection, concept detection, and image tagging. In each case, we obtain state-of-the-art results and order-of-magnitude speed ups.
Ameet Talwalkar, Lester Mackey, Yadong Mu, Shih-Fu Chang, Michael I. Jordan
ICCV4
2013 Large-Scale Video Hashing via Structure Learning
abstract
Recently, learning based hashing methods have become popular for indexing large-scale media data. Hashing methods map high-dimensional features to compact binary codes that are efficient to match and robust in preserving original similarity. However, most of the existing hashing methods treat videos as a simple aggregation of independent frames and index each video through combining the indexes of frames. The structure information of videos, e.g., discriminative local visual commonality and temporal consistency, is often neglected in the design of hash functions. In this paper, we propose a supervised method that explores the structure learning techniques to design efficient hash functions. The proposed video hashing method formulates a minimization problem over a structure-regularized empirical loss. In particular, the structure regularization exploits the common local visual patterns occurring in video frames that are associated with the same semantic class, and simultaneously preserves the temporal consistency over successive frames from the same video. We show that the minimization objective can be efficiently solved by an Accelerated Proximal Gradient (APG) method. Extensive experiments on two large video benchmark datasets (up to around 150K video clips with over 12 million frames) show that the proposed method significantly outperforms the state-of-the-art hashing methods.
Guangnan Ye, Dong Liu 0001, Jun Wang 0006, Shih-Fu Chang
ICCV4
2013 \(\propto\)SVM for Learning with Label Proportions
abstract
We study the problem of learning with label proportions in which the training data is provided in groups and only the proportion of each class in each group is known. We propose a new method called proportion-SVM, or \proptoSVM, which explicitly models the latent unknown instance labels together with the known group label proportions in a large-margin framework. Unlike the existing works, our approach avoids making restrictive assumptions about the data. The \proptoSVM model leads to a non-convex integer programming problem. In order to solve it efficiently, we propose two algorithms: one based on simple alternating optimization and the other based on a convex relaxation. Extensive experiments on standard datasets show that \proptoSVM outperforms the state-of-the-art, especially for larger group sizes.
Felix X. Yu, Dong Liu 0001, Sanjiv Kumar, Tony Jebara, Shih-Fu Chang
ICML (3)5
2013 Towards a comprehensive computational model foraesthetic assessment of videos
abstract
In this paper we propose a novel aesthetic model emphasizing psycho-visual statistics extracted from multiple levels in contrast to earlier approaches that rely only on descriptors suited for image recognition or based on photographic principles. At the lowest level, we determine dark-channel, sharpness and eye-sensitivity statistics over rectangular cells within a frame. At the next level, we extract Sentibank features (1,200 pre-trained visual classifiers) on a given frame, that invoke specific sentiments such as "colorful clouds", "smiling face" etc. and collect the classifier responses as frame-level statistics. At the topmost level, we extract trajectories from video shots. Using viewer's fixation priors, the trajectories are labeled as foreground, and background/camera on which statistics are computed. Additionally, spatio-temporal local binary patterns are computed that capture texture variations in a given shot. Classifiers are trained on individual feature representations independently. On thorough evaluation of 9 different types of features, we select the best features from each level -- dark channel, affect and camera motion statistics. Next, corresponding classifier scores are integrated in a sophisticated low-rank fusion framework to improve the final prediction scores. Our approach demonstrates strong correlation with human prediction on 1,000 broadcast quality videos released by NHK as an aesthetic evaluation dataset.
Subhabrata Bhattacharya, Behnaz Nojavanasghari, Tao Chen 0015, Dong Liu 0001, Shih-Fu Chang, Mubarak Shah
ACM Multimedia5
2013 SentiBank: large-scale ontology and classifiers for detecting sentiment and emotions in visual content
abstract
A picture is worth one thousand words, but what words should be used to describe the sentiment and emotions conveyed in the increasingly popular social multimedia? We demonstrate a novel system which combines sound structures from psychology and the folksonomy extracted from social multimedia to develop a large visual sentiment ontology consisting of 1,200 concepts and associated classifiers called SentiBank. Each concept, defined as an Adjective Noun Pair (ANP), is made of an adjective strongly indicating emotions and a noun corresponding to objects or scenes that have a reasonable prospect of automatic detection. We believe such large-scale visual classifiers offer a powerful mid-level semantic representation enabling high-level sentiment analysis of social multimedia. We demonstrate novel applications made possible by SentiBank including live sentiment prediction of social media and visualization of visual content in a rich intuitive semantic space.
Damian Borth, Tao Chen 0015, Rongrong Ji, Shih-Fu Chang
ACM Multimedia4
2013 Large-scale visual sentiment ontology and detectors using adjective noun pairs
abstract
We address the challenge of sentiment analysis from visual content. In contrast to existing methods which infer sentiment or emotion directly from visual low-level features, we propose a novel approach based on understanding of the visual concepts that are strongly related to sentiments. Our key contribution is two-fold: first, we present a method built upon psychological theories and web mining to automatically construct a large-scale Visual Sentiment Ontology (VSO) consisting of more than 3,000 Adjective Noun Pairs (ANP). Second, we propose SentiBank, a novel visual concept detector library that can be used to detect the presence of 1,200 ANPs in an image. The VSO and SentiBank are distinct from existing work and will open a gate towards various applications enabled by automatic sentiment analysis. Experiments on detecting sentiment of image tweets demonstrate significant improvement in detection accuracy when comparing the proposed SentiBank based predictors with the text-based approaches. The effort also leads to a large publicly available resource consisting of a visual sentiment ontology, a large detector library, and the training/testing benchmark for visual sentiment analysis.
Damian Borth, Rongrong Ji, Tao Chen 0015, Thomas M. Breuel, Shih-Fu Chang
ACM Multimedia5
2013 Structured exploration of who, what, when, and where in heterogeneous multimedia news sources
abstract
We present a fully automatic system from raw data gathering to navigation over heterogeneous news sources, including over 18k hours of broadcast video news, 3.58M online articles, and 430M public Twitter messages. Our system addresses the challenge of extracting "who," "what," "when," and "where" from a truly multimodal perspective, leveraging audiovisual information in broadcast news and those embedded in articles, as well as textual cues in both closed captions and raw document content in articles and social media. Performed over time, we are able to extract and study the trend of topics in the news and detect interesting peaks in news coverage over the life of the topic. We visualize these peaks in trending news topics using automatically extracted keywords and iconic images, and introduce a novel multimodal algorithm for naming speakers in the news. We also present several intuitive navigation interfaces for interacting with these complex topic structures over different news sources.
Brendan Jou, Hongzhi Li 0001, Joseph G. Ellis, Daniel Morozoff-Abegauz, Shih-Fu Chang
ACM Multimedia5
2013 News rover: exploring topical structures and serendipity in heterogeneous multimedia news
abstract
News stories are rarely understood in isolation. Every story is driven by key entities that give the story its context. Persons, places, times, and several surrounding topics can often succinctly represent a news event, but are only useful if they can be both identified and linked together. We introduce a novel architecture called News Rover for re-bundling broadcast video news, online articles, and Twitter content. The system utilizes these many multimodal sources to link and organize content by topics, events, persons and time. We present two intuitive interfaces for navigating content by topics and their related news events as well as serendipitously learning about a news topic. These two interfaces trade-off between user-controlled and serendipitous exploration of news while retaining the story context. The novelty of our work includes the linking of multi-source, multimodal news content to extracted entities and topical structures for contextual understanding, and visualized in intuitive active and passive interfaces.
Hongzhi Li 0001, Brendan Jou, Joseph G. Ellis, Daniel Morozoff-Abegauz, Shih-Fu Chang
ACM Multimedia5
2013 Analyzing the Harmonic Structure in Graph-Based Learning
abstract
We show that either explicitly or implicitly, various well-known graph-based models exhibit a common significant \emph{harmonic} structure in its target function -- the value of a vertex is approximately the weighted average of the values of its adjacent neighbors. Understanding of such structure and analysis of the loss defined over such structure help reveal important properties of the target function over a graph. In this paper, we show that the variation of the target function across a cut can be upper and lower bounded by the ratio of its harmonic loss and the cut cost. We use this to develop an analytical tool and analyze 5 popular models in graph-based learning: absorbing random walks, partially absorbing random walks, hitting times, pseudo-inverse of graph Laplacian, and eigenvectors of the Laplacian matrices. Our analysis well explains several open questions of these models reported in the literature. Furthermore, it provides theoretical justifications and guidelines for their practical use. Simulations on synthetic and real datasets support our analysis.
Xiao-Ming Wu 0003, Zhenguo Li, Shih-Fu Chang
NIPS3
2013 Semi-supervised learning using greedy max-cut
Jun Wang 0006, Tony Jebara, Shih-Fu Chang
J. Mach. Learn. Res.3
2013 Query-Adaptive Image Search With Hash Codes
abstract
Scalable image search based on visual similarity has been an active topic of research in recent years. State-of-the-art solutions often use hashing methods to embed high-dimensional image features into Hamming space, where search can be performed in real-time based on Hamming distance of compact hash codes. Unlike traditional metrics (e.g., Euclidean) that offer continuous distances, the Hamming distances are discrete integer values. As a consequence, there are often a large number of images sharing equal Hamming distances to a query, which largely hurts search results where fine-grained ranking is very important. This paper introduces an approach that enables query-adaptive ranking of the returned images with equal Hamming distances to the queries. This is achieved by firstly offline learning bitwise weights of the hash codes for a diverse set of predefined semantic concept classes. We formulate the weight learning process as a quadratic programming problem that minimizes intra-class distance while preserving inter-class relationship captured by original raw image features. Query-adaptive weights are then computed online by evaluating the proximity between a query and the semantic concept classes. With the query-adaptive bitwise weights, returned images can be easily ordered by weighted Hamming distance at a finer-grained hash code level rather than the original Hamming distance level. Experiments on a Flickr image dataset show clear improvements from our proposed approach.
Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Multim.4
2013 How far we've come: Impact of 20 years of multimedia information retrieval
abstract
This article reviews the major research trends that emerged in the last two decades within the broad area of multimedia information retrieval, with a focus on the ACM Multimedia community. Trends are defined (nonscientifically) to be topics that appeared in ACM multimedia publications and have had a significant number of citations. The article also assesses the impacts of these trends on real-world applications. The views expressed are subjective and likely biased but hopefully useful for understanding the heritage of the community and stimulating new research direction.
Shih-Fu Chang
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Exploiting web images for event recognition in consumer videos: A multiple source domain adaptation approach
abstract
Recent work has demonstrated the effectiveness of domain adaptation methods for computer vision applications. In this work, we propose a new multiple source domain adaptation method called Domain Selection Machine (DSM) for event recognition in consumer videos by leveraging a large number of loosely labeled web images from different sources (e.g., Flickr.com and Photosig.com), in which there are no labeled consumer videos. Specifically, we first train a set of SVM classifiers (referred to as source classifiers) by using the SIFT features of web images from different source domains. We propose a new parametric target decision function to effectively integrate the static SIFT features from web images/video keyframes and the spacetime (ST) features from consumer videos. In order to select the most relevant source domains, we further introduce a new data-dependent regularizer into the objective of Support Vector Regression (SVR) using the ∊-insensitive loss, which enforces the target classifier shares similar decision values on the unlabeled consumer videos with the selected source classifiers. Moreover, we develop an alternating optimization algorithm to iteratively solve the target decision function and a domain selection vector which indicates the most relevant source domains. Extensive experiments on three real-world datasets demonstrate the effectiveness of our proposed method DSM over the state-of-the-art by a performance gain up to 46.41%.
Lixin Duan, Dong Xu 0001, Shih-Fu Chang
CVPR3
2012 Mobile product search with Bag of Hash Bits and boundary reranking
abstract
Rapidly growing applications on smartphones have provided an excellent platform for mobile visual search. Most of previous visual search systems adopt the framework of ”Bag of Words”, in which words indicate quantized codes of visual features. In this work, we propose a novel visual search system based on ”Bag of Hash Bits” (BoHB), in which each local feature is encoded to a very small number of hash bits, instead of quantized to visual words, and the whole image is represented as bag of hash bits. The proposed BoHB method offers unique benefits in solving the challenges associated with mobile visual search, e.g., low transmission cost, cheap memory and computation on the mobile side, etc. Moreover, our BoHB method leverages the distinct properties of hashing bits such as multi-table indexing, multiple bucket probing, bit reuse, and hamming distance based ranking to achieve efficient search over gigantic visual databases. The proposed method significantly outperforms state-of-the-art mobile visual search methods like CHoG, and other (conventional desktop) visual search approaches like bag of words via vocabulary tree, or product quantization. The proposed BoHB approach is easy to implement on mobile devices, and general in the sense that it can be applied to different types of local features, hashing algorithms and image databases. We also incorporate a boundary feature in the reranking step to describe the object shapes, complementing the local features that are usually used to characterize the local details. The boundary feature can further filter out noisy results and improve the search performance, especially at the coarse category level. Extensive experiments over large-scale data sets up to 400k product images demonstrate the effectiveness of our approach.
Junfeng He, Jinyuan Feng, Xianglong Liu 0001, Tai-Hsu Lin, Hyunjin Chung, Shih-Fu Chang
CVPR7
2012 Spherical hashing
abstract
Many binary code encoding schemes based on hashing have been actively studied recently, since they can provide efficient similarity search, especially nearest neighbor search, and compact data representations suitable for handling large scale image databases in many computer vision problems. Existing hashing techniques encode high-dimensional data points by using hyperplane-based hashing functions. In this paper we propose a novel hypersphere-based hashing function, spherical hashing, to map more spatially coherent data points into a binary code compared to hyperplane-based hashing functions. Furthermore, we propose a new binary code distance function, spherical Hamming distance, that is tailored to our hypersphere-based binary coding scheme, and design an efficient iterative optimization process to achieve balanced partitioning of data points for each hash function and independence between hashing functions. Our extensive experiments show that our spherical hashing technique significantly outperforms six state-of-the-art hashing techniques based on hyperplanes across various image benchmarks of sizes ranging from one to 75 million of GIST descriptors. The performance gains are consistent and large, up to 100% improvements. The excellent results confirm the unique merits of the proposed idea in using hyperspheres to encode proximity regions in high-dimensional spaces. Finally, our method is intuitive and easy to implement.
Jae-Pil Heo, Youngwoon Lee, Junfeng He, Shih-Fu Chang, Sung-Eui Yoon
CVPR4
2012 Robust visual domain adaptation with low-rank reconstruction
abstract
Visual domain adaptation addresses the problem of adapting the sample distribution of the source domain to the target domain, where the recognition task is intended but the data distributions are different. In this paper, we present a low-rank reconstruction method to reduce the domain distribution disparity. Specifically, we transform the visual samples in the source domain into an intermediate representation such that each transformed source sample can be linearly reconstructed by the samples of the target domain. Unlike the existing work, our method captures the intrinsic relatedness of the source samples during the adaptation process while uncovering the noises and outliers in the source domain that cannot be adapted, making it more robust than previous methods. We formulate our problem as a constrained nuclear norm and ℓ2, 1norm minimization objective and then adopt the Augmented Lagrange Multiplier (ALM) method for the optimization. Extensive experiments on various visual adaptation tasks show that the proposed method consistently and significantly beats the state-of-the-art domain adaptation methods.
I-Hong Jhuo, Dong Liu 0001, D. T. Lee, Shih-Fu Chang
CVPR4
2012 Segmentation using superpixels: A bipartite graph partitioning approach
abstract
Grouping cues can affect the performance of segmentation greatly. In this paper, we show that superpixels (image segments) can provide powerful grouping cues to guide segmentation, where superpixels can be collected easily by (over)-segmenting the image using any reasonable existing segmentation algorithms. Generated by different algorithms with varying parameters, superpixels can capture diverse and multi-scale visual patterns of a natural image. Successful integration of the cues from a large multitude of superpixels presents a promising yet not fully explored direction. In this paper, we propose a novel segmentation framework based on bipartite graph partitioning, which is able to aggregate multi-layer superpixels in a principled and very effective manner. Computationally, it is tailored to unbalanced bipartite graph structure and leads to a highly efficient, linear-time spectral algorithm. Our method achieves significantly better performance on the Berkeley Segmentation Database compared to state-of-the-art techniques.
Zhenguo Li, Xiao-Ming Wu 0003, Shih-Fu Chang
CVPR3
2012 Supervised hashing with kernels
abstract
Recent years have witnessed the growing popularity of hashing in large-scale vision problems. It has been shown that the hashing quality could be boosted by leveraging supervised information into hash function learning. However, the existing supervised methods either lack adequate performance or often incur cumbersome model training. In this paper, we propose a novel kernel-based supervised hashing model which requires a limited amount of supervised information, i.e., similar and dissimilar data pairs, and a feasible training cost in achieving high quality hashing. The idea is to map the data to compact binary codes whose Hamming distances are minimized on similar pairs and simultaneously maximized on dissimilar pairs. Our approach is distinct from prior works by utilizing the equivalence between optimizing the code inner products and the Hamming distances. This enables us to sequentially and efficiently train the hash functions one bit at a time, yielding very short yet discriminative codes. We carry out extensive experiments on two image benchmarks with up to one million samples, demonstrating that our approach significantly outperforms the state-of-the-arts in searching both metric distance neighbors and semantically similar neighbors, with accuracy gains ranging from 13% to 46%.
Wei Liu 0005, Jun Wang 0006, Rongrong Ji, Yu-Gang Jiang 0001, Shih-Fu Chang
CVPR5
2012 Robust late fusion with rank minimization
abstract
In this paper, we propose a rank minimization method to fuse the predicted confidence scores of multiple models, each of which is obtained based on a certain kind of feature. Specifically, we convert each confidence score vector obtained from one model into a pairwise relationship matrix, in which each entry characterizes the comparative relationship of scores of two test samples. Our hypothesis is that the relative score relations are consistent among component models up to certain sparse deviations, despite the large variations that may exist in the absolute values of the raw scores. Then we formulate the score fusion problem as seeking a shared rank-2 pairwise relationship matrix based on which each original score matrix from individual model can be decomposed into the common rank-2 matrix and sparse deviation errors. A robust score vector is then extracted to fit the recovered low rank score relation matrix. We formulate the problem as a nuclear norm and ℓ1norm optimization objective function and employ the Augmented Lagrange Multiplier (ALM) method for the optimization. Our method is isotonic (i.e., scale invariant) to the numeric scales of the scores originated from different models. We experimentally show that the proposed method achieves significant performance gains on various tasks including object categorization and video event detection.
Guangnan Ye, Dong Liu 0001, I-Hong Jhuo, Shih-Fu Chang
CVPR4
2012 Weak attributes for large-scale image retrieval
abstract
Attribute-based query offers an intuitive way of image retrieval, in which users can describe the intended search targets with understandable attributes. In this paper, we develop a general and powerful framework to solve this problem by leveraging a large pool of weak attributes comprised of automatic classifier scores or other mid-level representations that can be easily acquired with little or no human labor. We extend the existing retrieval model of modeling dependency within query attributes to modeling dependency of query attributes on a large pool of weak attributes, which is more expressive and scalable. To efficiently learn such a large dependency model without overfitting, we further propose a semi-supervised graphical model to map each multiattribute query to a subset of weak attributes. Through extensive experiments over several attribute benchmarks, we demonstrate consistent and significant performance improvements over the state-of-the-art techniques. In addition, we compile the largest multi-attribute image retrieval dateset to date, including 126 fully labeled query attributes and 6,000 weak attributes of 0.26 million images.
Felix X. Yu, Rongrong Ji, Ming-Hen Tsai, Guangnan Ye, Shih-Fu Chang
CVPR5
2012 Scene Aligned Pooling for Complex Video Recognition
Liangliang Cao, Yadong Mu, Apostol Natsev, Shih-Fu Chang, Gang Hua 0001, John R. Smith
ECCV (2)4
2012 Accelerated Large Scale Optimization by Concomitant Hashing
Yadong Mu, John Wright 0001, Shih-Fu Chang
ECCV (1)3
2012 On the Difficulty of Nearest Neighbor Search
Junfeng He, Sanjiv Kumar, Shih-Fu Chang
ICML3
2012 Compact Hyperplane Hashing with Bilinear Functions
Wei Liu 0005, Jun Wang 0006, Yadong Mu, Sanjiv Kumar, Shih-Fu Chang
ICML5
2012 Compact hashing for mixed image-keyword query over multi-label images
abstract
Recently locality-sensitive hashing (LSH) algorithms have attracted much attention owing to its empirical success and theoretic guarantee in large-scale visual search. In this paper we address the new topic of hashing with multi-label data, in which images in the database are assumed to be associated with missing or noisy multiple labels and each query consists of a query image and several textual search terms, similar to the new "Search with Image" function introduced by the Google Image Search. The returned images are judged based on the combination of visual similarity and semantic information conveyed by search terms. In most of the state-of-the-art approaches, the learned hashing functions are universal for all labels. To further enhance the hashing efficiency for such multi-label data, we propose a novel scheme "boosted shared hashing". Our basic observation is that image labels typically form cliques in the feature space. Hashing efficacy can be greatly improved by making each hashing function more targeted at and only shared across such cliques instead of all labels in conventional hashing methods. In other words, each hashing function is deliberately designed such that it is especially effective for a subset of labels. The targeted, but sparse association between labels and hash bits reduces the computation and storage when indexing a new datum, since only a small number of relevant hashing functions become active given the labels. We develop a Boosting-style algorithm for simultaneously optimizing the label subset and hashing function in a unified framework. Experimental results on standard image benchmarks like CIFAR-10 and NUS-WIDE show that the proposed hashing scheme achieves substantially superior performances over conventional methods in terms of accuracy under the same hash bit budget.
Xianglong Liu 0001, Yadong Mu, Bo Lang, Shih-Fu Chang
ICMR4
2012 Joint audio-visual bi-modal codewords for video event detection
abstract
Joint audio-visual patterns often exist in videos and provide strong multi-modal cues for detecting multimedia events. However, conventional methods generally fuse the visual and audio information only at a superficial level, without adequately exploring deep intrinsic joint patterns. In this paper, we propose a joint audio-visual bi-modal representation, called bi-modal words. We first build a bipartite graph to model relation across the quantized words extracted from the visual and audio modalities. Partitioning over the bipartite graph is then applied to construct the bi-modal words that reveal the joint patterns across modalities. Finally, different pooling strategies are employed to re-quantize the visual and audio words into the bi-modal words and form bi-modal Bag-of-Words representations that are fed to subsequent multimedia event classifiers. We experimentally show that the proposed multi-modal feature achieves statistically significant performance gains over methods using individual visual and audio features alone and alternative multi-modal fusion methods. Moreover, we found that average pooling is the most suitable strategy for bi-modal feature generation.
Guangnan Ye, I-Hong Jhuo, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang
ICMR6
2012 Submodular video hashing: a unified framework towards video pooling and indexing
abstract
This paper develops a novel framework for efficient large-scale video retrieval. We aim to find video according to higher level similarities, which is beyond the scope of traditional near duplicate search. Following the popular hashing technique we employ compact binary codes to facilitate nearest neighbor search. Unlike the previous methods which capitalize on only one type of hash code for retrieval, this paper combines heterogeneous hash codes to effectively describe the diverse and multi-scale visual contents in videos. Our method integrates feature pooling and hashing in a single framework. In the pooling stage, we cast video frames into a set of pre-specified components, which capture a variety of semantics of video contents. In the hashing stage, we represent each video component as a compact hash code, and combine multiple hash codes into hash tables for effective search. To speed up the retrieval while retaining most informative codes, we propose a graph-based influence maximization method to bridge the pooling and hashing stages. We show that the influence maximization problem is submodular, which allows a greedy optimization method to achieve a nearly optimal solution. Our method works very efficiently, retrieving thousands of video clips from TRECVID dataset in about 0.001 second. For a larger scale synthetic dataset with 1M samples, it uses less than 1 second in response to 100 queries. Our method is extensively evaluated in both unsupervised and supervised scenarios, and the results on TRECVID Multimedia Event Detection and Columbia Consumer Video datasets demonstrate the success of our proposed technique.
Liangliang Cao, Zhenguo Li, Yadong Mu, Shih-Fu Chang
ACM Multimedia4
2012 Hybrid social media network
abstract
Analysis and recommendation of multimedia information can be greatly improved if we know the interactions between the content, user, and concept, which can be easily observed from the social media networks. However, there are many heterogeneous entities and relations in such networks, making it difficult to fully represent and exploit the diverse array of information. In this paper, we develop a hybrid social media network, through which the heterogeneous entities and relations are seamlessly integrated and a joint inference procedure across the heterogeneous entities and relations can be developed. The network can be used to generate personalized information recommendation in response to specific targets of interests, e.g., personalized multimedia albums, target advertisement and friend/topic recommendation. In the proposed network, each node denotes an entity and the multiple edges between nodes characterize the diverse relations between the entities (e.g., friends, similar contents, related concepts, favorites, tags, etc). Given a query from a user indicating his/her information needs, a propagation over the hybrid social media network is employed to infer the utility scores of all the entities in the network while learning the edge selection function to activate only a sparse subset of relevant edges, such that the query information can be best propagated along the activated paths. Driven by the intuition that much redundancy exists among the diverse relations, we have developed a robust optimization framework based on several sparsity principles. We show significant performance gains of the proposed method over the state of the art in multimedia retrieval and recommendation using data crawled from social media sites. To the best of our knowledge, this is the first model supporting not only aggregation but also judicious selection of heterogeneous relations in the social media networks.
Dong Liu 0001, Guangnan Ye, Ching-Ting Chen, Shuicheng Yan, Shih-Fu Chang
ACM Multimedia5
2012 Learning with Partially Absorbing Random Walks
abstract
We propose a novel stochastic process that is with probability $\alpha_i$ being absorbed at current state $i$, and with probability $1-\alpha_i$ follows a random edge out of it. We analyze its properties and show its potential for exploring graph structures. We prove that under proper absorption rates, a random walk starting from a set $\mathcal{S}$ of low conductance will be mostly absorbed in $\mathcal{S}$. Moreover, the absorption probabilities vary slowly inside $\mathcal{S}$, while dropping sharply outside $\mathcal{S}$, thus implementing the desirable cluster assumption for graph-based learning. Remarkably, the partially absorbing process unifies many popular models arising in a variety of contexts, provides new insights into them, and makes it possible for transferring findings from one paradigm to another. Simulation results demonstrate its promising applications in graph-based learning.
Xiao-Ming Wu 0003, Zhenguo Li, Anthony Man-Cho So, John Wright 0001, Shih-Fu Chang
NIPS5
2012 Semi-Supervised Hashing for Large-Scale Search
abstract
Hashing-based approximate nearest neighbor (ANN) search in huge databases has become popular due to its computational and memory efficiency. The popular hashing methods, e.g., Locality Sensitive Hashing and Spectral Hashing, construct hash functions based on random or principal projections. The resulting hashes are either not very accurate or are inefficient. Moreover, these methods are designed for a given metric similarity. On the contrary, semantic similarity is usually given in terms of pairwise labels of samples. There exist supervised hashing methods that can handle such semantic similarity, but they are prone to overfitting when labeled data are small or noisy. In this work, we propose a semi-supervised hashing (SSH) framework that minimizes empirical error over the labeled set and an information theoretic regularizer over both labeled and unlabeled sets. Based on this framework, we present three different semi-supervised hashing methods, including orthogonal hashing, nonorthogonal hashing, and sequential hashing. Particularly, the sequential hashing method generates robust codes in which each hash function is designed to correct the errors made by the previous ones. We further show that the sequential learning paradigm can be extended to unsupervised domains where no labeled pairs are available. Extensive experiments on four large datasets (up to 80 million samples) demonstrate the superior performance of the proposed SSH methods over state-of-the-art supervised and unsupervised hashing techniques.
Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Web-Scale Multimedia Processing and Applications [Scanning the Issue]
abstract
The articles in this special issue focus on web-scale multimedia processing as well as applications for its use.
Edward Y. Chang, Shih-Fu Chang, Alex Hauptmann 0001, Thomas S. Huang, Malcolm Slaney
Proc. IEEE2
2012 Robust and Scalable Graph-Based Semisupervised Learning
abstract
Graph-based semisupervised learning (GSSL) provides a promising paradigm for modeling the manifold structures that may exist in massive data sources in high-dimensional spaces. It has been shown effective in propagating a limited amount of initial labels to a large amount of unlabeled data, matching the needs of many emerging applications such as image annotation and information retrieval. In this paper, we provide reviews of several classical GSSL methods and a few promising methods in handling challenging issues often encountered in web-scale applications. First, to successfully incorporate the contaminated noisy labels associated with web data, label diagnosis and tuning techniques applied to GSSL are surveyed. Second, to support scalability to the gigantic scale (millions or billions of samples), recent solutions based on anchor graphs are reviewed. To help researchers pursue new ideas in this area, we also summarize a few popular data sets and software tools publicly available. Important open issues are discussed at the end to stimulate future research.
Wei Liu 0005, Jun Wang 0006, Shih-Fu Chang
Proc. IEEE3
2012 Fast Semantic Diffusion for Large-Scale Context-Based Image and Video Annotation
abstract
Exploring context information for visual recognition has recently received significant research attention. This paper proposes a novel and highly efficient approach, which is named semantic diffusion, to utilize semantic context for large-scale image and video annotation. Starting from the initial annotation of a large number of semantic concepts (categories), obtained by either machine learning or manual tagging, the proposed approach refines the results using a graph diffusion technique, which recovers the consistency and smoothness of the annotations over a semantic graph. Different from the existing graph-based learning methods that model relations among data samples, the semantic graph captures context by treating the concepts as nodes and the concept affinities as the weights of edges. In particular, our approach is capable of simultaneously improving annotation accuracy and adapting the concept affinities to new test data. The adaptation provides a means to handle domain change between training and test data, which often occurs in practice. Extensive experiments are conducted to improve concept annotation results using Flickr images and TV program videos. Results show consistent and significant performance gain (10 +% on both image and video data sets). Source codes of the proposed algorithms are available online.
Yu-Gang Jiang 0001, Qi Dai 0001, Jun Wang 0006, Chong-Wah Ngo, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Image Process.6
2012 Active query sensing: Suggesting the best query view for mobile visual search
abstract
While much exciting progress is being made in mobile visual search, one important question has been left unexplored in all current systems. When searching objects or scenes in the 3D world, which viewing angle is more likely to be successful? More particularly, if the first query fails to find the right target, how should the user control the mobile camera to form the second query? In this article, we propose a novel Active Query Sensing system for mobile location search, which actively suggests the best subsequent query view to recognize the physical location in the mobile environment. The proposed system includes two unique components: (1) an offline process for analyzing the saliencies of different views associated with each geographical location, which predicts the location search precisions of individual views by modeling their self-retrieval score distributions. (2) an online process for estimating the view of an unseen query, and suggesting the best subsequent view change. Specifically, the optimal viewing angle change for the next query can be formulated as an online information theoretic approach. Using a scalable visual search system implemented over a NYC street view dataset (0.3 million images), we show a performance gain by reducing the failure rate of mobile location search to only 12% after the second query. We have also implemented an end-to-end functional system, including user interfaces on iPhones, client-server communication, and a remote search server. This work may open up an exciting new direction for developing interactive mobile media applications through the innovative exploitation of active sensing and query formulation.
Rongrong Ji, Felix X. Yu, Tongtao Zhang, Shih-Fu Chang
ACM Trans. Multim. Comput. Commun. Appl.4
2011 Compact hashing with joint optimization of search accuracy and time
abstract
Similarity search, namely, finding approximate nearest neighborhoods, is the core of many large scale machine learning or vision applications. Recently, many research results demonstrate that hashing with compact codes can achieve promising performance for large scale similarity search. However, most of the previous hashing methods with compact codes only model and optimize the search accuracy. Search time, which is an important factor for hashing in practice, is usually not addressed explicitly. In this paper, we develop a new scalable hashing algorithm with joint optimization of search accuracy and search time simultaneously. Our method generates compact hash codes for data of general formats with any similarity function. We evaluate our method using diverse data sets up to 1 million samples (e.g., web images). Our comprehensive results show the proposed method significantly outperforms several state-of-the-art hashing approaches.
Junfeng He, Shih-Fu Chang, Regunathan Radhakrishnan, Claus Bauer
CVPR2
2011 Noise resistant graph ranking for improved web image search
abstract
In this paper, we exploit a novel ranking mechanism that processes query samples with noisy labels, motivated by the practical application of web image search re-ranking where the originally highest ranked images are usually posed as pseudo queries for subsequent re-ranking. Availing ourselves of the low-frequency spectrum of a neighborhood graph built on the samples, we propose a graph-theoretical framework amenable to noise resistant ranking. The proposed framework consists of two components: spectral filtering and graph-based ranking. The former leverages sparse bases, progressively selected from a pool of smooth eigenvectors of the graph Laplacian, to reconstruct the noisy label vector associated with the query sample set and accordingly filter out the query samples with less authentic positive labels. The latter applies a canonical graph ranking algorithm with respect to the filtered query sample set. Quantitative image re-ranking experiments carried out on two public web image databases bear out that our re-ranking approach compares favorably with the state-of-the-arts and improves web image search engines by a large margin though we harvest the noisy queries from the top-ranked images returned by these search engines.
Wei Liu 0005, Yu-Gang Jiang 0001, Jiebo Luo 0001, Shih-Fu Chang
CVPR4
2011 Learning component-level sparse representation using histogram information for image classification
abstract
A novel component-level dictionary learning framework which exploits image group characteristics within sparse coding is introduced in this work. Unlike previous methods, which select the dictionaries that best reconstruct the data, we present an energy minimization formulation that jointly optimizes the learning of both sparse dictionary and component level importance within one unified framework to give a discriminative representation for image groups. The importance measures how well each feature component represents the image group property with the dictionary by using histogram information. Then, dictionaries are updated iteratively to reduce the influence of unimportant components, thus refining the sparse representation for each image group. In the end, by keeping the top K important components, a compact representation is derived for the sparse coding dictionary. Experimental results on several public datasets are shown to demonstrate the superior performance of the proposed algorithm compared to the-state-of-the-art methods.
Chen-Kuo Chiang, Chih-Hsueh Duan, Shang-Hong Lai, Shih-Fu Chang
ICCV4
2011 Towards Optimal Discriminating Order for Multiclass Classification
abstract
In this paper, we investigate how to design an optimized discriminating order for boosting multiclass classification. The main idea is to optimize a binary tree architecture, referred to as Sequential Discriminating Tree (SDT), that performs the multiclass classification through a hierarchical sequence of coarse-to-fine binary classifiers. To infer such a tree architecture, we employ the constrained large margin clustering procedure which enforces samples belonging to the same class to locate at the same side of the hyper plane while maximizing the margin between these two partitioned class subsets. The proposed SDT algorithm has a theoretic error bound which is shown experimentally to effectively guarantee the generalization performance. Experiment results indicate that SDT clearly beats the state-of-the-art multiclass classification algorithms.
Dong Liu 0001, Shuicheng Yan, Yadong Mu, Xian-Sheng Hua 0001, Shih-Fu Chang, HongJiang Zhang
ICDM5
2011 Hashing with Graphs
Wei Liu 0005, Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
ICML4
2011 Lost in binarization: query-adaptive ranking for similar image search with compact codes
abstract
With the proliferation of images on the Web, fast search of visually similar images has attracted significant attention. State-of-the-art techniques often embed high-dimensional visual features into low-dimensional Hamming space, where search can be performed in real-time based on Hamming distance of compact binary codes. Unlike traditional metrics (e.g., Euclidean) of raw image features that produce continuous distance, the Hamming distances are discrete integer values. In practice, there are often a large number of images sharing equal Hamming distances to a query, resulting in a critical issue for image search where ranking is very important. In this paper, we propose a novel approach that facilitates query-adaptive ranking for the images with equal Hamming distance. We achieve this goal by firstly offline learning bit weights of the binary codes for a diverse set of predefined semantic concept classes. The weight learning process is formulated as a quadratic programming problem that minimizes intra-class distance while preserving interclass relationship in the original raw image feature space. Query-adaptive weights are then rapidly computed by evaluating the proximity between a query and the concept categories. With the adaptive bit weights, the returned images can be ordered by weighted Hamming distance at a finer-grained binary code level rather than at the original integer Hamming distance level. Experimental results on a Flickr image dataset show clear improvements from our query-adaptive ranking approach.
Yu-Gang Jiang 0001, Jun Wang 0006, Shih-Fu Chang
ICMR3
2011 Consumer video understanding: a benchmark database and an evaluation of human and machine performance
abstract
Recognizing visual content in unconstrained videos has become a very important problem for many applications. Existing corpora for video analysis lack scale and/or content diversity, and thus limited the needed progress in this critical area. In this paper, we describe and release a new database called CCV, containing 9,317 web videos over 20 semantic categories, including events like "baseball" and "parade", scenes like "beach", and objects like "cat". The database was collected with extra care to ensure relevance to consumer interest and originality of video content without post-editing. Such videos typically have very little textual annotation and thus can benefit from the development of automatic content analysis techniques.
Yu-Gang Jiang 0001, Guangnan Ye, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ICMR3
2011 Content based multimedia retrieval: lessons learned from two decades of research
abstract
In the past two decades, we have witnessed bourgeoning research on content based multimedia information retrieval, covering a wide range of topics such as feature extraction, content matching, structure parsing, semantic annotation, multimodal analysis, 3D content retrieval, and user-in-the-loop interaction. More than ten years have also passed since the publication of the influential survey paper by Smeulders et al on content based image retrieval. Recently, exciting solutions are emerging in several practical contexts such as mobile media search, augmented reality, and Web-scale copy detection. However, many fundamental problems remain open, including but not limited to large-scale semantic annotation, multimedia ontological organization, and human-machine interaction for searching complex events. In this talk, I will discuss lessons learned from our past research, drawing from successes and failures in developing and deploying a few image/video search systems in different domains, and then share views about promising future directions.
Shih-Fu Chang
ACM Multimedia1
2011 Mobile product search with bag of hash bits
abstract
The advent of smart phones has provided an excellent plat- form for mobile visual search. Most of previous mobile visual search systems adopt the framework of "bag of words",in which words indicate quantized codes of visual features. In this work, we propose a novel mobile visual search system based on "bag of hash bits". Using new ideas for hash bit selection, multi-hash table generation, and hamming-distance soft scoring, we overcome the problem of bit inefficiency affecting the traditional hashing approaches, and achieve promising accuracy outperforming state of the art. The framework is also general in that general feature type can be used for generating the hash bits. Demos and experiments over a large scale product image set demonstrate the effectiveness of our approach.
Junfeng He, Tai-Hsu Lin, Jinyuan Feng, Shih-Fu Chang
ACM Multimedia4
2011 Towards low bit rate mobile visual search with multiple-channel coding
abstract
In this paper, we propose a multiple-channel coding scheme to extract compact visual descriptors for low bit rate mobile visual search. Different from previous visual search scenarios that send the query image, we make use of the ever growing mobile computational capability to directly extract compact visual descriptors at the mobile end. Meanwhile, stepping forward from the state-of-the-art compact descriptor extractions, we exploit the rich contextual cues at the mobile end (such as GPS tags for mobile visual search and 2D barcodes or RFID tags for mobile product search), together with the visual statistics at the reference database, to learn multiple coding channels. Therefore, we describe the query with one of many forms of high-dimensional visual signature, which is subsequently mapped to one or more channels and compressed. The compression function within each channel is learnt based on a novel robust PCA scheme, with specific consideration to preserve the retrieval ranking capability of the original signature. We have deployed our scheme on both iPhone4 and HTC DESIRE 7 to search ten million landmark images in a low bit rate setting. Quantitative comparisons to the state-of-the-arts demonstrate our significant advantages in descriptor compactness (with orders of magnitudes improvement) and retrieval mAP in mobile landmark, product, and CD/book cover search.
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Yong Rui, Shih-Fu Chang, Wen Gao 0001
ACM Multimedia6
2011 Active query sensing for mobile location search
abstract
While much exciting progress is being made in mobile visual search, one important question has been left unexplored in all current systems. When the first query fails to find the right target (up to 50% likelihood), how should the user form his/her search strategy in the subsequent interaction? In this paper, we propose a novel Active Query Sensing system to suggest the best way for sensing the surrounding scenes while forming the second query for location search. We accomplish the goal by developing several unique components -- an offline process for analyzing the saliency of the views associated with each geographical location based on score distribution modeling, predicting the visual search precision of individual views and locations, estimating the view of an unseen query, and suggesting the best subsequent view change. Using a scalable visual search system implemented over a NYC street view data set (0.3 million images), we show a performance gain as high as two folds, reducing the failure rate of mobile location search to only 12% after the second query. This work may open up an exciting new direction for developing interactive mobile media applications through innovative exploitation of active sensing and query formulation.
Felix X. Yu, Rongrong Ji, Shih-Fu Chang
ACM Multimedia3
2011 A mobile location search system with active query sensing
abstract
How should the second query be taken once the first query fails in mobile location search based on visual recognition? In this demo, we describe a mobile search system with a unique Active Query Sensing (AQS) function to intelligently guide the mobile user to take a successful second query. This suggestion is built upon a scalable visual matching system covering over 0.3 million street view reference images in New York City, where each location is associated with multiple surrounding views and panorama. In online search, once the initial search result fails, the system will perform online analysis and suggest the mobile user to turn to the most discriminative viewing angle to take the second visual query, from which the search performance is expected to greatly improve. The AQS suggestion is based on both offline salient view discovery and online viewing angle prediction and intelligent turning decision. Our experiments show our AVS can improve the mobile location search with a performance gain as high as 100%, reducing the failure rate to only 12% after taking the second visual query.
Felix X. Yu, Rongrong Ji, Tongtao Zhang, Shih-Fu Chang
ACM Multimedia4
2011 Modeling Scene and Object Contexts for Human Action Retrieval With Few Examples
abstract
The use of context knowledge is critical for understanding human actions, which typically occur under particular scene settings with certain object interactions. For instance, driving car usually happens outdoors, and kissing involves two people moving toward each other. In this paper, we investigate the problem of context modeling for human action retrieval. We first identify ten simple object-level action atoms relevant to many human actions, e.g., people getting closer. With the action atoms and several background scene classes, we show that action retrieval can be improved through modeling action-scene-object dependency. An algorithm inspired by the popular semi-supervised learning paradigm is introduced for this purpose. One important contribution of this paper is to show that modeling the dependencies among actions, objects, and scenes can be efficiently achieved with very few examples. Such a solution has tremendous potential in practice as it is often expensive to acquire large sets of training data. Experiments were performed on the challenging Hollywood2 dataset containing 89 movies. The results validate the effectiveness of our approach, achieving a mean average precision of 26% with just ten examples per action.
Yu-Gang Jiang 0001, Zhenguo Li, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.3
2010 Semi-supervised hashing for scalable image retrieval
abstract
Large scale image search has recently attracted considerable attention due to easy availability of huge amounts of data. Several hashing methods have been proposed to allow approximate but highly efficient search. Unsupervised hashing methods show good performance with metric distances but, in image search, semantic similarity is usually given in terms of labeled pairs of images. There exist supervised hashing methods that can handle such semantic similarity but they are prone to overfitting when labeled data is small or noisy. Moreover, these methods are usually very slow to train. In this work, we propose a semi-supervised hashing method that is formulated as minimizing empirical error on the labeled data while maximizing variance and independence of hash bits over the labeled and unlabeled data. The proposed method can handle both metric as well as semantic similarity. The experimental results on two large datasets (up to one million samples) demonstrate its superior performance over state-of-the-art supervised and unsupervised methods.
Jun Wang 0006, Ondrej Kumar, Shih-Fu Chang
CVPR3
2010 Single-view recaptured image detection based on physics-based features
abstract
In daily life, we can see images of real-life objects on posters, television, or virtually any type of smooth physical surfaces. We seldom confuse these images with the objects per se mainly with the help of the contextual information from the surrounding environment and nearby objects. Without this contextual information, distinguishing an object from an image of the object becomes subtle; it is precisely an effect that a large immersive display aims at achieving. In this work, we study and address a problem that mirrors the above-mentioned recognition problem, i.e., distinguishing images of true natural scenes and those from recapturing. Being able to detect recaptured images, robot vision can be more intelligent and a single-image-based counter-measure for re-broadcast attack on a face authentication system becomes feasible. This work is timely as the face authentication system is getting common on consumer mobile devices such as smart phones and laptop computers. In this work, we present a physical model for image recapturing and the features derived from the model are used in a recaptured image detector. Our physics-based method out-performs a statistics-based method by a significant margin on images of VGA (640×480) and QVGA (320×240) resolutions which are common for mobile devices. In our study, we find that apart from the contextual information, the unique properties for the recaptured image rendering process are crucial for the recognition problem.
Xinting Gao, Tian-Tsong Ng, Shih-Fu Chang
ICME4
2010 Large Graph Construction for Scalable Semi-Supervised Learning
Wei Liu 0005, Junfeng He, Shih-Fu Chang
ICML3
2010 Sequential Projection Learning for Hashing with Compact Codes
Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
ICML3
2010 Scalable similarity search with optimized kernel hashing
abstract
Scalable similarity search is the core of many large scale learning or data mining applications. Recently, many research results demonstrate that one promising approach is creating compact and efficient hash codes that preserve data similarity. By efficient, we refer to the low correlation (and thus low redundancy) among generated codes. However, most existing hash methods are designed only for vector data. In this paper, we develop a new hashing algorithm to create efficient codes for large scale data of general formats with any kernel function, including kernels on vectors, graphs, sequences, sets and so on. Starting with the idea analogous to spectral hashing, novel formulations and solutions are proposed such that a kernel based hash function can be explicitly represented and optimized, and directly applied to compute compact hash codes for new samples of general formats. Moreover, we incorporate efficient techniques, such as Nystrom approximation, to further reduce time and space complexity for indexing and search, making our algorithm scalable to huge data sets. Another important advantage of our method is the ability to handle diverse types of similarities according to actual task requirements, including both feature similarities and semantic similarities like label consistency. We evaluate our method using both vector and non-vector data sets at a large scale up to 1 million samples. Our comprehensive results show the proposed method outperforms several state-of-the-art approaches for all the tasks, with a significant gain for most tasks.
Junfeng He, Wei Liu 0005, Shih-Fu Chang
KDD3
2010 In a Blink of an Eye and a Switch of a Transistor: Cortically Coupled Computer Vision
abstract
Our society's information technology advancements have resulted in the increasingly problematic issue of information overload—i.e., we have more access to information than we can possibly process. This is nowhere more apparent than in the volume of imagery and video that we can access on a daily basis—for the general public, availability of YouTube video and Google Images, or for the image analysis professional tasked with searching security video or satellite reconnaissance. Which images to look at and how to ensure we see the images that are of most interest to us, begs the question of whether there are smart ways to triage this volume of imagery. Over the past decade, computer vision research has focused on the issue of ranking and indexing imagery. However, computer vision is limited in its ability to identify interesting imagery, particularly as “interesting” might be defined by an individual. In this paper we describe our efforts in developing brain–computer interfaces (BCIs) which synergistically integrate computer vision and human vision so as to construct a system for image triage. Our approach exploits machine learning for real-time decoding of brain signals which are recorded noninvasively via electroencephalography (EEG). The signals we decode are specific for events related to imagery attracting a user's attention. We describe two architectures we have developed for this type of cortically coupled computer vision and discuss potential applications and challenges for the future.
Paul Sajda, Eric Pohlmeyer, Jun Wang 0006, Lucas C. Parra, Christoforos Christoforou, Jacek Dmochowski, Barbara Hanna, Claus Bahlmann, Maneesh Kumar Singh 0001, Shih-Fu Chang
Proc. IEEE10
2010 Near Duplicate Identification With Spatially Aligned Pyramid Matching
abstract
A new framework, termed spatially aligned pyramid matching, is proposed for near duplicate image identification. The proposed method robustly handles spatial shifts as well as scale changes, and is extensible for video data. Images are divided into both overlapped and non-overlapped blocks over multiple levels. In the first matching stage, pairwise distances between blocks from the examined image pair are computed using earth mover's distance (EMD) or the visual word with$\chi^{2}$distance based method with scale-invariant feature transform (SIFT) features. In the second stage, multiple alignment hypotheses that consider piecewise spatial shifts and scale variation are postulated and resolved using integer-flow EMD. Moreover, to compute the distances between two videos, we conduct the third step matching (i.e., temporal matching) after spatial matching. Two application scenarios are addressed—near duplicate retrieval (NDR) and near duplicate detection (NDD). For retrieval ranking, a pyramid-based scheme is constructed to fuse matching results from different partition levels. For NDD, we also propose a dual-sample approach by using the multilevel distances as features and support vector machine for binary classification. The proposed methods are shown to clearly outperform existing methods through extensive testing on the Columbia Near Duplicate Image Database and two new datasets. In addition, we also discuss in depth our framework in terms of the extension for video NDR and NDD, the sensitivity to parameters, the utilization of multiscale dense SIFT descriptors, and the test of scalability in image NDD.
Dong Xu 0001, Tat-Jen Cham, Shuicheng Yan, Lixin Duan, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.5
2010 Camera Response Functions for Image Forensics: An Automatic Algorithm for Splicing Detection
abstract
We present a fully automatic method to detect doctored digital images. Our method is based on a rigorous consistency checking principle of physical characteristics among different arbitrarily shaped image regions. In this paper, we specifically study the camera response function (CRF), a fundamental property in cameras mapping input irradiance to output image intensity. A test image is first automatically segmented into distinct arbitrarily shaped regions. One CRF is estimated from each region using geometric invariants from locally planar irradiance points (LPIPs). To classify a boundary segment between two regions as authentic or spliced, CRF-based cross fitting and local image features are computed and fed to statistical classifiers. Such segment level scores are further fused to infer the image level authenticity. Tests on two data sets reach performance levels of 70% precision and 70% recall, showing promising potential for real-world applications. Moreover, we examine individual features and discover the key factor in splicing detection. Our experiments show that the anomaly introduced around splicing boundaries plays the major role in detecting splicing. Such finding is important for designing effective and efficient solutions to image splicing detection.
Yu-Feng Hsu, Shih-Fu Chang
IEEE Trans. Inf. Forensics Secur.2
2010 Semi-supervised distance metric learning for collaborative image retrieval and clustering
abstract
Learning a good distance metric plays a vital role in many multimedia retrieval and data mining tasks. For example, a typical content-based image retrieval (CBIR) system often relies on an effective distance metric to measure similarity between any two images. Conventional CBIR systems simply adopting Euclidean distance metric often fail to return satisfactory results mainly due to the well-known semantic gap challenge. In this article, we present a novel framework of Semi-Supervised Distance Metric Learning for learning effective distance metrics by exploring the historical relevance feedback log data of a CBIR system and utilizing unlabeled data when log data are limited and noisy. We formally formulate the learning problem into a convex optimization task and then present a new technique, named as “Laplacian Regularized Metric Learning” (LRML). Two efficient algorithms are then proposed to solve the LRML task. Further, we apply the proposed technique to two applications. One direct application is for Collaborative Image Retrieval (CIR), which aims to explore the CBIR log data for improving the retrieval performance of CBIR systems. The other application is for Collaborative Image Clustering (CIC), which aims to explore the CBIR log data for enhancing the clustering performance of image pattern clustering tasks. We conduct extensive evaluation to compare the proposed LRML method with a number of competing methods, including 2 standard metrics, 3 unsupervised metrics, and 4 supervised metrics with side information. Encouraging results validate the effectiveness of the proposed technique.
Steven C. H. Hoi, Wei Liu 0005, Shih-Fu Chang
ACM Trans. Multim. Comput. Commun. Appl.3
2010 Audio-visual atoms for generic video concept classification
abstract
We investigate the challenging issue of joint audio-visual analysis of generic videos targeting at concept detection. We extract a novel local representation, Audio-Visual Atom (AVA), which is defined as a region track associated with regional visual features and audio onset features. We develop a hierarchical algorithm to extract visual atoms from generic videos, and locate energy onsets from the corresponding soundtrack by time-frequency analysis. Audio atoms are extracted around energy onsets. Visual and audio atoms form AVAs, based on which discriminative audio-visual codebooks are constructed for concept detection. Experiments over Kodak's consumer benchmark videos confirm the effectiveness of our approach.
Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ACM Trans. Multim. Comput. Commun. Appl.3
2009 Robust multi-class transductive learning with graphs
abstract
Graph-based methods form a main category of semi-supervised learning, offering flexibility and easy implementation in many applications. However, the performance of these methods is often sensitive to the construction of a neighborhood graph, which is non-trivial for many real-world problems. In this paper, we propose a novel framework that builds on learning the graph given labeled and unlabeled data. The paper has two major contributions. Firstly, we use a nonparametric algorithm to learn the entire adjacency matrix of a symmetry-favored k-NN graph, assuming that the matrix is doubly stochastic. The nonparametric algorithm makes the constructed graph highly robust to noisy samples and capable of approximating underlying submanifolds or clusters. Secondly, to address multi-class semi-supervised classification, we formulate a constrained label propagation problem on the learned graph by incorporating class priors, leading to a simple closed-form solution. Experimental results on both synthetic and real-world datasets show that our approach is significantly better than the state-of-the-art graph-based semi-supervised learning algorithms in terms of accuracy and robustness.
Wei Liu 0005, Shih-Fu Chang
CVPR2
2009 Label diagnosis through self tuning forweb image search
abstract
Semi-supervised learning (SSL) relies on partial supervision information for prediction, where only a small set of samples are associated with labels. Performance of SSL is significantly degraded if the given labels are not reliable. Such problems arise in realistic applications such as web image search using noisy textual tags. This paper proposes a novel and efficient graph based SSL method with the unique capacity of pruning contradictory labels and inferring new labels through a bidirectional and alternating optimization process. The objective is to automatically identify the most suitable samples for manipulation, labeling or unlabeling, and meanwhile estimate a smooth classification function over a weighted graph. Different from other graph based SSL approaches, the proposed method employs a bivariate objective function and iteratively modifies label variables on both labeled and unlabeled samples. Starting from such a SSL setting, we present a relearning framework to improve the performance of base learner, particularly for the application of web image search. Besides the toy demonstration on artificial data, we evaluated the proposed method on Flickr image search with unreliable textual labels. Experimental results confirm the significant improvements of the method over the baseline text based search engine and the state-of-the-art SSL methods.
Jun Wang 0006, Yu-Gang Jiang 0001, Shih-Fu Chang
CVPR3
2009 Visual saliency with side information
abstract
We propose novel algorithms for organizing large image and video datasets using both the visual content and the associated side-information, such as time, location, authorship, and so on. Earlier research have used side-information as pre-filter before visual analysis is performed, and we design a machine learning algorithm to model the join statistics of the content and the side information. Our algorithm, diverse-density contextual clustering (D2C2), starts by finding unique patterns for each sub-collection sharing the same side-info, e.g., scenes from winter. It then finds the common patterns that are shared among all subsets, e.g., persistent scenes across all seasons. These unique and common prototypes are found with multiple instance learning and subsequent clustering steps. We evaluate D2C2on two Web photo collections from Flickr and one news video collection from TRECVID. Results show that not only the visual patterns found by D2C2are intuitively salient across different seasons, locations and events, classifiers constructed from the unique and common patterns also outperform state-of-the-art bag-of-features classifiers.
Wei Jiang 0001, Lexing Xie, Shih-Fu Chang
ICASSP3
2009 Domain adaptive semantic diffusion for large scale context-based video annotation
abstract
Learning to cope with domain change has been known as a challenging problem in many real-world applications. This paper proposes a novel and efficient approach, named domain adaptive semantic diffusion (DASD), to exploit semantic context while considering the domain-shift-of-context for large scale video concept annotation. Starting with a large set of concept detectors, the proposed DASD refines the initial annotation results using graph diffusion technique, which preserves the consistency and smoothness of the annotation over a semantic graph. Different from the existing graph learning methods which capture relations among data samples, the semantic graph treats concepts as nodes and the concept affinities as the weights of edges. Particularly, the DASD approach is capable of simultaneously improving the annotation results and adapting the concept affinities to new test data. The adaptation provides a means to handle domain change between training and test data, which occurs very often in video annotation task. We conduct extensive experiments to improve annotation results of 374 concepts over 340 hours of videos from TRECVID 2005-2007 data sets. Results show consistent and significant performance gain over various baselines. In addition, the proposed approach is very efficient, completing DASD over 374 concepts within just 2 milliseconds for each video shot on a regular PC.
Yu-Gang Jiang 0001, Jun Wang 0006, Shih-Fu Chang, Chong-Wah Ngo
ICCV3
2009 Muti-scale temporal segmentation and outlier detection in sensor networks
abstract
Monitoring multimodal data generated by sensor networks for extracting information is a challenging task for the human observer. To manage the barrage of data, one needs to create mechanisms for identifying only those time intervals which are informative and worthy of further highlevel analysis either by machine or the human observer. We regard a time interval to be informative and contain an event if it is uncommon or distinct from routine background. Different events in general may unfold at different temporal scales. Here, we present a non-parametric distribution based approach for event detection in sensor network data. In this approach we employ multiple sliding windows at different scales to obtain the distribution of the data. We segment the temporal data stream and identify the potential event bearing candidates by comparing the present and past statistical behavior of the data. In the experiments we demonstrate the effect of optimum bandwidth selection on accuracy and the range of allowable window sizes and therefore time scales. We analyze the computational speed as well as the supporting empirical results on the bin width.
Mandis Beigi, Shih-Fu Chang, Shahram Ebadollahi, Dinesh C. Verma
ICME2
2009 Graph construction and b-matching for semi-supervised learning
abstract
Graph based semi-supervised learning (SSL) methods play an increasingly important role in practical machine learning systems. A crucial step in graph based SSL methods is the conversion of data into a weighted graph. However, most of the SSL literature focuses on developing label inference algorithms without extensively studying the graph building method and its effect on performance. This article provides an empirical study of leading semi-supervised methods under a wide range of graph construction algorithms. These SSL inference algorithms include the Local and Global Consistency (LGC) method, the Gaussian Random Field (GRF) method, the Graph Transduction via Alternating Minimization (GTAM) method as well as other techniques. Several approaches for graph construction, sparsification and weighting are explored including the popular k-nearest neighbors method (kNN) and the b-matching method. As opposed to the greedily constructed kNN graph, the b-matched graph ensures each node in the graph has the same number of edges and produces a balanced or regular graph. Experimental results on both artificial data and real benchmark datasets indicate that b-matching produces more robust graphs and therefore provides significantly better prediction accuracy without any significant change in computation time.
Tony Jebara, Jun Wang 0006, Shih-Fu Chang
ICML3
2009 Mobile media search: has media search finally found its perfect platform? part II
abstract
Recently, many exciting media search applications have been introduced to take advantage of smart phones' audiovisual capture capabilities and their being always on and connected. These applications address a real pain point for most mobile users and allow them to search with minimal text entry, if any. Is the mobile platform an ideal fit for media search? Are audio and visual signal processing technologies sufficiently accurate to support most mobile search applications? What are the killer applications of mobile media search? Earlier in 2009 at ICASSP, a panel on this topic stirred up great interest and enthusiasm while leaving many questions untouched due to the limited time.
Berna Erol, Jiebo Luo 0001, Shih-Fu Chang, Minoru Etoh, Hsiao-Wuen Hon, Qian Lin 0001, Vidya Setlur
ACM Multimedia3
2009 Short-term audio-visual atoms for generic video concept classification
abstract
We investigate the challenging issue of joint audio-visual analysis of generic videos targeting at semantic concept detection. We propose to extract a novel representation, the Short-term Audio-Visual Atom (S-AVA), for improved concept detection. An S-AVA is defined as a short-term region track associated with regional visual features and background audio features. An effective algorithm, named Short-Term Region tracking with joint Point Tracking and Region Segmentation (STR-PTRS), is developed to extract S-AVAs from generic videos under challenging conditions such as uneven lighting, clutter, occlusions, and complicated motions of both objects and camera. Discriminative audio-visual codebooks are constructed on top of S-AVAs using Multiple Instance Learning. Codebook-based features are generated for semantic concept detection. We extensively evaluate our algorithm over Kodak's consumer benchmark video set from real users. Experimental results confirm significant performance improvements - over 120% MAP gain compared to alternative approaches using static region segmentation without temporal tracking. The joint audio-visual features also outperform visual features alone by an average of 8.5% (in terms of AP) over 21 concepts, with many concepts achieving more than 20%.
Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ACM Multimedia3
2009 Semantic context transfer across heterogeneous sources for domain adaptive video search
abstract
Automatic video search based on semantic concept detectors has recently received significant attention. Since the number of available detectors is much smaller than the size of human vocabulary, one major challenge is to select appropriate detectors to response user queries. In this paper, we propose a novel approach that leverages heterogeneous knowledge sources for domain adaptive video search. First, instead of utilizing WordNet as most existing works, we exploit the context information associated with Flickr images to estimate query-detector similarity. The resulting measurement, named Flickr context similarity (FCS), reflects the co-occurrence statistics of words in image context rather than textual corpus. Starting from an initial detector set determined by FCS, our approach novelly transfers semantic context learned from test data domain to adaptively refine the query-detector similarity. The semantic context transfer process provides an effective means to cope with the domain shift between external knowledge source (e.g., Flickr context) and test data, which is a critical issue in video search. To the best of our knowledge, this work represents the first research aiming to tackle the challenging issue of domain change in video search. Extensive experiments on 120 textual queries over TRECVID 2005-2008 data sets demonstrate the effectiveness of semantic context transfer for domain adaptive video search. Results also show that the FCS is suitable for measuring query-detector similarity, producing better performance to various other popular measures.
Yu-Gang Jiang 0001, Chong-Wah Ngo, Shih-Fu Chang
ACM Multimedia3
2009 Brain state decoding for rapid image retrieval
abstract
Human visual perception is able to recognize a wide range of targets under challenging conditions, but has limited throughput. Machine vision and automatic content analytics can process images at a high speed, but suffers from inadequate recognition accuracy for general target classes. In this paper, we propose a new paradigm to explore and combine the strengths of both systems. A single trial EEG-based brain machine interface (BCI) subsystem is used to detect objects of interest of arbitrary classes from an initial subset of images. The EEG detection outcomes are used as input to a graph-based pattern mining subsystem to identify, refine, and propagate the labels to retrieve relevant images from a much larger pool. The combined strategy is unique in its generality, robustness, and high throughput. It has great potential for advancing the state of the art in media retrieval applications. We have evaluated and demonstrated significant performance gains of the proposed system with multiple and diverse image classes over several data sets, including those from Internet (Caltech 101) and remote sensing images. In this paper, we will also present insights learned from the experiments and discuss future research directions.
Jun Wang 0006, Eric Pohlmeyer, Barbara Hanna, Yu-Gang Jiang 0001, Paul Sajda, Shih-Fu Chang
ACM Multimedia6
2009 An image score inference system for RNAi genome-wide screening based on fuzzy mixture regression modeling
Jun Wang 0006, Xiaobo Zhou 0001, Fuhai Li 0001, Pamela Bradley, Shih-Fu Chang, Norbert Perrimon, Stephen T. C. Wong
J. Biomed. Informatics5
2009 Enhancing Bilinear Subspace Learning by Element Rearrangement
abstract
The success of bilinear subspace learning heavily depends on reducing correlations among features along rows and columns of the data matrices. In this work, we study the problem of rearranging elements within a matrix in order to maximize these correlations so that information redundancy in matrix data can be more extensively removed by existing bilinear subspace learning algorithms. An efficient iterative algorithm is proposed to tackle this essentially integer programming problem. In each step, the matrix structure is refined with a constrained Earth Mover's Distance procedure that incrementally rearranges matrices to become more similar to their low-rank approximations, which have high correlation among features along rows and columns. In addition, we present two extensions of the algorithm for conducting supervised bilinear subspace learning. Experiments in both unsupervised and supervised bilinear subspace learning demonstrate the effectiveness of our proposed algorithms in improving data compression performance and classification accuracy.
Dong Xu 0001, Shuicheng Yan, Stephen Lin 0001, Thomas S. Huang, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.5
2008 Fast kernel learning for spatial pyramid matching
abstract
Spatial pyramid matching (SPM) is a simple yet effective approach to compute similarity between images. Similarity kernels at different regions and scales are usually fused by some heuristic weights. In this paper, we develop a novel and fast approach to improve SPM by finding the optimal kernel fusing weights from multiple scales, locations, as well as codebooks. One unique contribution of our approach is the novel formulation of kernel matrix learning problem leading to an efficient quadratic programming solution, with much lower complexity than those associated with existing solutions (e.g., semidefinite programming). We demonstrate performance gains of the proposed methods by evaluations over well-known public data sets such as natural scenes and TRECVID 2007.
Junfeng He, Shih-Fu Chang, Lexing Xie
CVPR2
2008 Semi-supervised distance metric learning for Collaborative Image Retrieval
abstract
Typical content-based image retrieval (CBIR) solutions with regular Euclidean metric usually cannot achieve satisfactory performance due to the semantic gap challenge. Hence, relevance feedback has been adopted as a promising approach to improve the search performance. In this paper, we propose a novel idea of learning with historical relevance feedback log data, and adopt a new paradigm called ldquoCollaborative Image Retrievalrdquo (CIR). To effectively explore the log data, we propose a novel semi-supervised distance metric learning technique, called ldquoLaplacian Regularized Metric Learningrdquo (LRML), for learning robust distance metrics for CIR. Different from previous methods, the proposed LRML method integrates both log data and unlabeled data information through an effective graph regularization framework. We show that reliable metrics can be learned from real log data even they may be noisy and limited at the beginning stage of a CIR system. We conducted extensive evaluation to compare the proposed method with a large number of competing methods, including 2 standard metrics, 3 unsupervised metrics, and 4 supervised metrics with side information.
Steven C. H. Hoi, Wei Liu 0005, Shih-Fu Chang
CVPR3
2008 Active microscopic cellular image annotation by superposable graph transduction with imbalanced labels
abstract
Systematic content screening of cell phenotypes in microscopic images has been shown promising in gene function understanding and drug design. However, manual annotation of cells and images in genome-wide studies is cost prohibitive. In this paper, we propose a highly efficient active annotation framework, in which a small amount of expert input is leveraged to rapidly and effectively infer the labels over the remaining unlabeled data. We formulate this as a graph based transductive learning problem and develop a novel method for label propagation. Specifically, a label regularizer method is proposed to handle the important label imbalance issue, typically seen in the cellular image screening applications. We also design a new scheme which breaks the graph into linear superposition of contributions from individual labeled samples. We take advantage of such a superposable representation to achieve fast annotation in an interactive setting. Extensive evaluations over toy data and realistic cellular images confirm the superiority of the proposed method over existing alternatives.
Jun Wang 0006, Shih-Fu Chang, Xiaobo Zhou 0001, Stephen T. C. Wong
CVPR2
2008 Near duplicate image identification with patially Aligned Pyramid Matching
abstract
A new framework, termed spatially aligned pyramid matching, is proposed for near duplicate image identification. The proposed method robustly handles spatial shifts as well as scale changes. Images are divided into both overlapped and non-overlapped blocks over multiple levels. In the first matching stage, pairwise distances between blocks from the examined image pair are computed using SIFT features and Earth Moverpsilas distance (EMD). In the second stage, multiple alignment hypotheses that consider piecewise spatial shifts and scale variation are postulated and resolved using integer-flow EMD. Two application scenarios are addressed - retrieval ranking and binary classification. For retrieval ranking, a pyramid-based scheme is constructed to fuse matching results from different partition levels. For binary classification, a novel generalized neighborhood component analysis method is formulated that can be effectively used in tandem with SVMs to select the most critical matching components. The proposed methods are shown to clearly outperform existing methods through extensive testing on the Columbia near duplicate image database and another new dataset.
Dong Xu 0001, Tat-Jen Cham, Shuicheng Yan, Shih-Fu Chang
CVPR4
2008 Semantic Concept Classification by Joint Semi-supervised Learning of Feature Subspaces and Support Vector Machines
Wei Jiang 0001, Shih-Fu Chang, Tony Jebara, Alexander C. Loui
ECCV (4)2
2008 Cross-domain learning methods for high-level visual concept classification
abstract
Exploding amounts of multimedia data increasingly require automatic indexing and classification, e.g. training classifiers to produce high-level features, or semantic concepts, chosen to represent image content, like car, person, etc. When changing the applied domain (i.e. from news domain to consumer home videos), the classifiers trained in one domain often perform poorly in the other domain due to changes in feature distributions. Additionally, classifiers trained on the new domain alone may suffer from too few positive training samples. Appropriately adapting data/models from an old domain to help classify data in a new domain is an important issue. In this work, we develop a new cross-domain SVM (CDSVM) algorithm for adapting previously learned support vectors from one domain to help classification in another domain. Better precision is obtained with almost no additional computational cost. Also, we give a comprehensive summary and comparative study of the state-of-the-art SVM-based cross-domain learning methods. Evaluation over the latest large-scale TRECVID benchmark data set shows that our CDSVM method can improve mean average precision over 36 concepts by 7.5%. For further performance gain, we also propose an intuitive selection criterion to determine which cross-domain learning method to use for each concept.
Wei Jiang 0001, Eric Zavesky, Shih-Fu Chang, Alexander C. Loui
ICIP3
2008 Graph transduction via alternating minimization
abstract
Graph transduction methods label input data by learning a classification function that is regularized to exhibit smoothness along a graph over labeled and unlabeled samples. In practice, these algorithms are sensitive to the initial set of labels provided by the user. For instance, classification accuracy drops if the training set contains weak labels, if imbalances exist across label classes or if the labeled portion of the data is not chosen at random. This paper introduces a propagation algorithm that more reliably minimizes a cost function over both a function on the graph and a binary label matrix. The cost function generalizes prior work in graph transduction and also introduces node normalization terms for resilience to label imbalances. We demonstrate that global minimization of the function is intractable but instead provide an alternating minimization scheme that incrementally adjusts the function and the labels towards a reliable local minimum. Unlike prior methods, the resulting propagation of labels does not prematurely commit to an erroneous labeling and obtains more consistent labels. Experiments are shown for synthetic and real classification tasks including digit and text recognition. A substantial improvement in accuracy compared to state of the art semi-supervised methods is achieved. The advantage are even more dramatic when labeled instances are limited.
Jun Wang 0006, Tony Jebara, Shih-Fu Chang
ICML3
2008 Internet image archaeology: automatically tracing the manipulation history of photographs on the web
abstract
We propose a system for automatically detecting the ways in which images have been copied and edited or manipulated. We draw upon these manipulation cues to construct probable parent-child relationships between pairs of images, where the child image was derived through a series of visual manipulations on the parent image. Through the detection of these relationships across a plurality of images, we can construct a history of the image, called the visual migration map (VMM), which traces the manipulations applied to the image through past generations. We propose to apply VMMs as part of a larger internet image archaeology system (IIAS), which can process a given set of related images and surface many interesting instances of images from within the set. In particular, the image closest to the "original" photograph might be among the images with the most descendants in the VMM. Or, the images that are most deeply descended from the original may exhibit unique differences and changes in the perspective being conveyed by the author. We evaluate the system across a set of photographs crawled from the web and find that many types of image manipulations can be automatically detected and used to construct plausible VMMs. These maps can then be successfully mined to find interesting instances of images and to suppress uninteresting or redundant ones, leading to a better understanding of how images are used over different times, sources, and contexts.
Lyndon S. Kennedy, Shih-Fu Chang
ACM Multimedia2
2008 SIFT-Bag kernel for video event analysis
abstract
In this work, we present a SIFT-Bag based generative-todiscriminative framework for addressing the problem of video event recognition in unconstrained news videos. In the generative stage, each video clip is encoded as a bag of SIFT feature vectors, the distribution of which is described by a Gaussian Mixture Models (GMM). In the discriminative stage, the SIFT-Bag Kernel is designed for characterizing the property of Kullback-Leibler divergence between the specialized GMMs of any two video clips, and then this kernel is utilized for supervised learning in two ways. On one hand, this kernel is further refined in discriminating power for centroid-based video event classification by using the Within-Class Covariance Normalization approach, which depresses the kernel components with high-variability for video clips of the same event. On the other hand, the SIFT-Bag Kernel is used in a Support Vector Machine for margin-based video event classification. Finally, the outputs from these two classifiers are fused together for final decision. The experiments on the TRECVID 2005 corpus demonstrate that the mean average precision is boosted from the best reported 38.2 % in [36] to 60.4 % based on our new framework.
Xiaodan Zhuang, Shuicheng Yan, Shih-Fu Chang, Mark Hasegawa-Johnson, Thomas S. Huang
ACM Multimedia4
2008 Video Event Recognition Using Kernel Methods with Multilevel Temporal Alignment
abstract
In this work, we systematically study the problem of event recognition in unconstrained news video sequences. We adopt the discriminative kernel-based method for which video clip similarity plays an important role. First, we represent a video clip as a bag of orderless descriptors extracted from all of the constituent frames and apply the earth mover's distance (EMD) to integrate similarities among frames from two clips. Observing that a video clip is usually comprised of multiple subclips corresponding to event evolution over time, we further build a multilevel temporal pyramid. At each pyramid level, we integrate the information from different subclips with Integer-value-constrained EMD to explicitly align the subclips. By fusing the information from the different pyramid levels, we develop Temporally Aligned Pyramid Matching (TAPM) for measuring video similarity. We conduct comprehensive experiments on the TRECVID 2005 corpus, which contains more than 6,800 clips. Our experiments demonstrate that 1) the TAPM multilevel method clearly outperforms single-level EMD (SLEMD) and 2) SLEMD outperforms keyframe and multiframe-based detection methods by a large margin. In addition, we conduct in-depth investigation of various aspects of the proposed techniques such as weight selection in SLEMD, sensitivity to temporal clustering, the effect of temporal alignment, and possible approaches for speed up. Extensive analysis of the results also reveals intuitive interpretation of video event recognition through video subclip alignment at different levels.
Dong Xu 0001, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Query-Adaptive Fusion for Multimodal Search
abstract
We conduct a broad survey of query-adaptive search strategies in a variety of application domains, where the internal retrieval mechanisms used for search are adapted in response to the anticipated needs for each individual query experienced by the system. While these query-adaptive approaches can range from meta-search over text collections to multimodal search over video databases, we propose that all such systems can be framed and discussed in the context of a single, unified framework. In our paper, we keep an eye towards the domain of video search, where search cues are available from a rich set of modalities, including textual speech transcripts, low-level visual features, and high-level semantic concept detectors. The relative efficacy of each of the modalities is highly variant between many types of queries. We observe that the state of the art in query-adaptive retrieval frameworks for video collections is highly dependent upon the definition of classes of queries, which are groups of queries that share similar optimal search strategies, while many applications in text and Web retrieval have included many advanced strategies, such as direct prediction of search method performance and inclusion of contextual cues from the searcher. We conclude that such advanced strategies previously developed for text retrieval have a broad range of possible applications in future research in multimodal video search.
Lyndon S. Kennedy, Shih-Fu Chang, Apostol Natsev
Proc. IEEE2
2008 Quality-Optimized and Secure End-to-End Authentication for Media Delivery
abstract
The need for security services, such as confidentiality and authentication, has become one of the major concerns in multimedia communication applications, such as video on demand and peer-to-peer content delivery. Conventional data authentication cannot be directly applied for streaming media when an unreliable channel is used and packet loss may occur. This paper begins by reviewing existing end-to-end media authentication schemes, which can be classified into stream-based and content-based techniques. We then motivate and describe how to design authentication schemes for multimedia delivery that exploit the unequal importance of different packets. By applying conventional cryptographic hashes and digital signatures to the media packets, the system security is similar to that achievable in conventional data security. However, instead of optimizing packet verification probability, we optimize the quality of the authenticated media, which is determined by the packets that are received and able to be decoded and authenticated. The quality of the authenticated media is optimized by allocating the authentication resources unequally across streamed packets based on their relative importance, thereby providing unequal authenticity protection. The effectiveness of this approach is demonstrated through experimental results on different media types (image and video), different compression standards (JPEG, JPEG2000, and H.264), and different channels (wired with packet erasures and wireless with bit errors).
Qibin Sun, John G. Apostolopoulos, Chang Wen Chen, Shih-Fu Chang
Proc. IEEE4
2008 An Introduction to the Special Issue on Event Analysis in Videos
abstract
I NTEREST from industry and academia has increased dramatically over recent years in the challenging area of event analysis and recognition from various video sources including sports, surveillance, user-generated video, etc. Video event analysis and recognition is a critical task in many applications such as detection of sporting highlights, incident detection in surveillance video, indexing, retrieval and summarization of video databases, and human-computer interaction. This special issue aims to capture the latest advances by the research community working in the area of video event analysis. The call for papers was enthusiastically greeted by the research community and we received over seventy submissions. The special issue presents 16 articles which provide fundamental contributions in a wide range of topics in video event analysis: 1) human action and activity recognition; 2) motion trajectory analysis; 3) video content analysis and pattern mining; 4) audio-visual multi-modal analysis; and 5) video analysis applications. An overview of the organization and a brief summary of the articles selected for publication in the special issue are provided below.
Shih-Fu Chang, Jiebo Luo 0001, Stephen J. Maybank, Dan Schonfeld, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2007 Kernel Sharing With Joint Boosting For Multi-Class Concept Detection
abstract
Object/scene detection by discriminative kernel-based classification has gained great interest due to its promising performance and flexibility. In this paper, unlike traditional approaches that independently build binary classifiers to detect individual concepts, we proposed a new framework for multi-class concept detection based on kernel sharing and joint learning. By sharing "good" kernels among concepts, accuracy of individual weak detectors can be greatly improved; by joint learning of common detectors among classes, the required kernels and the computational complexity for detecting each individual concept can be reduced. We demonstrated our approach by developing an extended JointBoost framework, which was used to choose the optimal kernel and subset of sharing classes in an iterative boosting process. In addition, we constructed multi-resolution visual vocabularies by hierarchical clustering and computed kernels based on spatial matching. We tested our method in detecting 12 concepts (objects, scenes, etc) over 80+ hours of broadcast news videos from the challenging TRECVID 2005 corpus. Significant performance gains were achieved -10% in mean average precision (MAP) and up to 34% average precision (AP) for some concepts like maps, building, and boat-ship. Extensive analysis of the results also revealed interesting and important underlying relations among concepts.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
CVPR2
2007 Using Geometry Invariants for Camera Response Function Estimation
abstract
In this paper, we present a new single-image camera response function (CRF) estimation method using geometry invariants (GI). We derive mathematical properties and geometric interpretation for GI, which lend insight to addressing various algorithm implementation issues in a principled way. In contrast to the previous single-image CRF estimation methods, our method provides a constraint equation for selecting the potential target data points. Comparing to the prior work, our experiment is conducted over more extensive data and our method is flexible in that its estimation accuracy and stability can be improved whenever more than one image is available. The geometry invariance theory is novel and may be of wide interest.
Tian-Tsong Ng, Shih-Fu Chang, Mao-Pei Tsui
CVPR2
2007 Visual Event Recognition in News Video using Kernel Methods with Multi-Level Temporal Alignment
abstract
In this work, we systematically study the problem of visual event recognition in unconstrained news video sequences. We adopt the discriminative kernel-based method for which video clip similarity plays an important role. First, we represent a video clip as a bag of orderless descriptors extracted from all of the constituent frames and apply Earth mover's distance (EMD) to integrate similarities among frames from two clips. Observing that a video clip is usually comprised of multiple sub-clips corresponding to event evolution over time, we further build a multilevel temporal pyramid. At each pyramid level, we integrate the information from different sub-clips with Integer-value-constrained EMD to explicitly align the sub-clips. By fusing the information from the different pyramid levels, we develop temporally aligned pyramid matching (TAPM) for measuring video similarity. We conduct comprehensive experiments on the Trecvid 2005 corpus, which contains more than 6,800 clips. Our experiments demonstrate that 1) the TAPM multi-level method clearly outperforms single-level EMD, and 2) single-level EMD outperforms by a large margin (43.0% in Mean Average Precision) basic detection methods that use only a single key-frame. Extensive analysis of the results also reveals an intuitive interpretation of subclip alignment at different levels.
Dong Xu 0001, Shih-Fu Chang
CVPR2
2007 Element Rearrangement for Tensor-Based Subspace Learning
abstract
The success of tensor-based subspace learning depends heavily on reducing correlations along the column vectors of the mode-k flattened matrix. In this work, we study the problem of rearranging elements within a tensor in order to maximize these correlations, so that information redundancy in tensor data can be more extensively removed by existing tensor-based dimensionality reduction algorithms. An efficient iterative algorithm is proposed to tackle this essentially integer optimization problem. In each step, the tensor structure is refined with a spatially-constrained Earth Mover's Distance procedure that incrementally rearranges tensors to become more similar to their low rank approximations, which have high correlation among features along certain tensor dimensions. Monotonic convergence of the algorithm is proven using an auxiliary function analogous to that used for proving convergence of the Expectation-Maximization algorithm. In addition, we present an extension of the algorithm for conducting supervised subspace learning with tensor data. Experiments in both unsupervised and supervised subspace learning demonstrate the effectiveness of our proposed algorithms in improving data compression performance and classification accuracy.
Shuicheng Yan, Dong Xu 0001, Stephen Lin 0001, Thomas S. Huang, Shih-Fu Chang
CVPR5
2007 Recent Advances and Challenges of Semantic Image/Video Search
abstract
We present an overview of recent advances and major challenges in image and video search, with a specific focus on large-scale semantic concept detection and indexing. Such semantic indexing paradigm has been driven by the increasing availability of the large resources of corpora, novel labeling approaches, innovative image features, and machine learning techniques for visual content recognition. We will discus key approaches, recent results, and novel applications in text-to-concept semantic search and multi-modal retrieval models. Open issues and major opportunities are also presented.
Shih-Fu Chang, Wei-Ying Ma, Arnold W. M. Smeulders
ICASSP (4)1
2007 Context-Based Concept Fusion with Boosted Conditional Random Fields
abstract
The contextual relationships among different semantic concepts provide important information for automatic concept detection in images/videos. We propose a new context-based concept fusion (CBCF) method for semantic concept detection. Our work includes two folds. (1) We model the inter-conceptual relationships by a conditional random field (CRF) that improves detection results from independent detectors by taking into account the inter-correlation among concepts. CRF directly models the posterior probability of concept labels and is more accurate for the discriminative concept detection than previous statistical inferencing techniques. The boosted CRF framework is incorporated to further enhance performance by combining the power of boosting with CRF. (2) We develop an effective criterion to predict which concepts may benefit from CBCF. As reported in previous works, CBCF has inconsistent performance gain on different concepts. With accurate prediction, computational and data resources can be allocated to enhance concepts that are promising to gain performance. Evaluation on TRECVID2005 development set demonstrates the effectiveness of our algorithm.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
ICASSP (1)2
2007 Image Splicing Detection using Camera Response Function Consistency and Automatic Segmentation
abstract
We propose a fully automatic spliced image detection method based on consistency checking of camera characteristics among different areas in an image. A test image is first segmented into distinct areas. One camera response function (CRF) is estimated from each area using geometric invariants from locally planar irradiance points (LPIPs). To classify a boundary segment between two areas as authentic or spliced, CRF cross fitting scores and area intensity features are computed and fed to SVM-based classifiers. Such segment-level scores are further fused to form the image-level decision. Tests on both the benchmark data set and an unseen high-quality spliced data set reach promising performance levels with 70 % precision and 70 % recall. 1.
Yu-Feng Hsu, Shih-Fu Chang
ICME2
2007 Video search reranking through random walk over document-level context graph
abstract
Multimedia search over distributed sources often result in recurrent images or videos which are manifested beyond the textual modality. To exploit such contextual patterns and keep the simplicity of the keyword-based search, we propose novel reranking methods to leverage the recurrent patterns to improve the initial text search results. The approach, context reranking, is formulated as a random walk problem along the context graph, where video stories are nodes and the edges between them are weighted by multimodal contextual similarities. The random walk is biased with the preference towards stories with higher initial text search scores - a principled way to consider both initial text search results and their implicit contextual relationships. When evaluated on TRECVID 2005 video benchmark, the proposed approach can improve retrieval on the average up to 32% relative to the baseline text search method in terms of story-level Mean Average Precision. In the people-related queries, which usually have recurrent coverage across news sources, we can have up to 40% relative improvement. Most of all, the proposed method does not require any additional input from users (e.g., example images), or complex search models for special queries (e.g., named person search).
Winston H. Hsu, Lyndon S. Kennedy, Shih-Fu Chang
ACM Multimedia3
2007 Enabling MPEG-7 structural and semantic descriptions in retrieval applications
abstract
Abstract The MPEG‐7 standard supports the description of both the structure and the semantics of multimedia; however, the generation and consumption of MPEG‐7 structural and semantic descriptions are outside the scope of the standard. This article presents two research prototype systems that demonstrate the generation and consumption of MPEG‐7 structural and semantic descriptions in retrieval applications. The active system for MPEG‐4 video object simulation (AMOS) is a video object segmentation and retrieval system that segments, tracks, and models objects in videos (e.g., person, car) as a set of regions with corresponding visual features and spatiotemporal relations. The region‐based model provides an effective base for similarity retrieval of video objects. The second system, the Intelligent Multimedia Knowledge Application (IMKA), uses the novel MediaNet framework for representing semantic and perceptual information about the world using multimedia. MediaNet knowledge bases can be constructed automatically from annotated collections of multimedia data and used to enhance the retrieval of multimedia.
Ana B. Benitez, Di Zhong, Shih-Fu Chang
J. Assoc. Inf. Sci. Technol.3
2007 Utility-Based Video Adaptation for Universal Multimedia Access (UMA) and Content-Based Utility Function Prediction for Real-Time Video Transcoding
abstract
Many techniques exist for adapting videos to satisfy heterogeneous resource conditions or user preferences, whereas selection of the best adaptation operation among various choices usually is either ad hoc or inefficient. To provide a systematic solution, we present a conceptual framework based on utility function (UF), which models video entity, adaptation, resource, utility, and the relations among them. In order to support real-time video adaptation, we present a content-based statistical paradigm to facilitate the prediction of UF for real-time transcoding of live videos. Instead of modelling the UF through analytical models, as in the conventional rate-distortion framework, the proposed approach formulates the prediction as a classification and regression problem. Each video clip is classified into one of distinctive categories and then local regression is used to accurately predict the utility value. Our extensive experiment results based on MPEG-4 transcoding demonstrate that the proposed method achieves very promising performance - up to 89% accuracy in choosing the optimal transcoding operation (in both spatial and temporal dimensions) with the highest quality over a diverse range of target bit rates
Jae-Gon Kim, Shih-Fu Chang, Hyung-Myung Kim
IEEE Trans. Multim.3
2006 A Generative-Discriminative Hybrid Method for Multi-View Object Detection
abstract
We present a novel discriminative-generative hybrid approach in this paper, with emphasis on application in multiview object detection. Our method includes a novel generative model called Random Attributed Relational Graph (RARG) which is able to capture the structural and appearance characteristics of parts extracted from objects. We develop new variational learning methods to compute the approximation of the detection likelihood ratio function. The variaitonal likelihood ratio function can be shown to be a linear combination of the individual generative classifiers defined at nodes and edges of the RARG. Such insight inspires us to replace the generative classifiers at nodes and edges with discriminative classifiers, such as support vector machines, to further improve the detection performance. Our experiments have shown the robustness of the hybrid approach - the combined detection method incorporating the SVM-based discriminative classifiers yields superior detection performances compared to prior works in multiview object detection.
DongQing Zhang, Shih-Fu Chang
CVPR (2)2
2006 Complexity Adaptive H.264 Encoding for Light Weight Streams
abstract
Emerging video coding standard H.264 achieves significant efficiency improvements, at the expense of greatly increased computational complexity at both the encoder and the decoder. Most prior works focus on reducing the encoder complexity only. In this paper, we develop a novel approach to reduce the decoder complexity without changing any implementation of standard-compliant decoders. Our solution, called complexity adaptive motion estimation and mode decision (CAMED), involves several core components, including a rigorous rate-distortion-complexity optimization framework, complexity cost modeling, and a complexity control algorithm. Such components are incorporated in the encoder side. Experiments over diverse videos and bit rates have shown that our method can achieve significant decoding complexity reduction (up to 60% of interpolation computation) with little degradation in video quality (les 0.2 dB) and very minor impact on the encoder complexity. Results also show that our method can be readily combined with other methods to reduce the complexity at both encoder and decoder
Shih-Fu Chang
ICASSP (2)2
2006 Topic Tracking Across Broadcast News Videos with Visual Duplicates and Semantic Concepts
abstract
Videos from distributed sources (e.g., broadcasts, podcasts, blogs, etc.) have grown exponentially. Topic threading is very useful for organizing such large-volume information sources. Current solutions primarily rely on text features only but encounter difficulty when text is noisy or unavailable. In this paper, we propose new representations and similarity measures for news videos based on low-level features, visual near-duplicates, and high-level semantic concepts automatically detected from videos. We develop a multi-modal fusion framework for estimating relevance of a new story to a known topic. Our extensive experiments using TRECVID 2005 data set (171 hours, 6 channels, 3 languages) confirm that near-duplicates consistently and significantly boost the tracking performance by up to 25%. In addition, we present information-theoretic analysis to assess the complexity of each semantic topic and determine the best subset of concepts for tracking each topic.
Winston H. Hsu, Shih-Fu Chang
ICIP2
2006 Active Context-Based Concept Fusionwith Partial User Labels
abstract
In this paper we propose a new framework, called active context-based concept fusion, for effectively improving the accuracy of semantic concept detection in images and videos. Our approach solicits user annotations for a small number of concepts, which are used to refine the detection of the rest of concepts. In contrast with conventional methods, our approach is active, by using information theoretic criteria to automatically determine the optimal concepts for user annotation. Our experiments over TRECVID 2005 development set (about 80 hours) show significant performance gains. In addition, we have developed an effective method to predict concepts that may benefit from context-based fusion.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
ICIP2
2006 Visual Event Detection using Multi-Dimensional Concept Dynamics
abstract
A novel framework is introduced for visual event detection. Visual events are viewed as stochastic temporal processes in the semantic concept space. In this concept-centered approach to visual event modeling, the dynamic pattern of an event is modeled through the collective evolution patterns of the individual semantic concepts in the course of the visual event. Video clips containing different events are classified by employing information about how well their dynamics in the direction of each semantic concept matches those of a given event. Results indicate that such a data-driven statistical approach is in fact effective in detecting different visual events such as exiting car, riot, and airplane flying
Shahram Ebadollahi, Lexing Xie, Shih-Fu Chang, John R. Smith
ICME3
2006 Detecting Image Splicing using Geometry Invariants and Camera Characteristics Consistency
abstract
Recent advances in computer technology have made digital image tampering more and more common. In this paper, we propose an authentic vs. spliced image classification method making use of geometry invariants in a semi-automatic manner. For a given image, we identify suspicious splicing areas, compute the geometry invariants from the pixels within each region, and then estimate the camera response function (CRF) from these geometry invariants. The cross-fitting errors are fed into a statistical classifier. Experiments show a very promising accuracy, 87%, over a large data set of 363 natural and spliced images. To the best of our knowledge, this is the first work detecting image splicing by verifying camera characteristic consistency from a single-channel image
Yu-Feng Hsu, Shih-Fu Chang
ICME2
2006 Pattern Mining in Visual Concept Streams
abstract
Pattern mining algorithms are often much easier applied than quantitatively assessed. In this paper we address the pattern evaluation problem by looking at both the capability of models and the difficulty of target concepts. We use four different data mining models: frequent itemset mining, k-means clustering, hidden Markov model, and hierarchical hidden Markov model to mine 39 concept streams from the a 137-video broadcast news collection from TRECVID-2005. We hypothesize that the discovered patterns can reveal semantics beyond the input space, and thus evaluate the patterns against a much larger concept space containing 192 concepts defined by LSCOM. Results show that HHMM has the best average prediction among all models, however different models seem to excel in different concepts depending on the concept prior and the ontological relationship. Results also show that the majority of the target concepts are better predicted with temporal or combination hypotheses, and there are novel concepts found that are not part of the original lexicon. This paper presents the first effort on temporal pattern mining in the large concept space. There are many promising directions to use concept mining to help construct better concept detectors or to guide the design of multimedia ontology
Lexing Xie, Shih-Fu Chang
ICME2
2006 Concept-based electronic health records: opportunities and challenges
abstract
Healthcare is a data-rich but information-poor domain. Terabytes of multimedia medical data are being generated on a monthly basis in a typical healthcare organization in order to document patients' health status and care process. Government and health-related organizations are pushing for fully electronic, cross-institution, integrated Electronic Health Records to provide a better, cost effective and more complete access to this data. However, provision of efficient access to the content of such records for timely and decision-enabling information extraction will not be available. Such a capability is essential for providing efficient decision support and objective evidence to clinicians. In addition researchers, medical students, patients, and payers could also benefit from it. We present the idea of concept-based multimedia health records, which aims at organizing the health records at the information level. We will explore the opportunities and possibilities that such an organization will provide, what role the field of multimedia content management could play to materialize this type of health record organization, and what the challenges will be in the quest for realizing the idea.We believe that the field of multimedia can play a very active role in taking healthcare information systems to the next level by facilitating the access to decision-enabling information for different types of users in healthcare. Our goal is to share with the community our thoughts on where the field of multimedia content management research should be focusing its attention to have a fundamental impact on the practice of medicine.
Shahram Ebadollahi, Anni Coden, Michael A. Tanenblatt, Shih-Fu Chang, Tanveer F. Syeda-Mahmood, Arnon Amir
ACM Multimedia4
2006 Video search reranking via information bottleneck principle
abstract
We propose a novel and generic video/image reranking algorithm, IB reranking, which reorders results from text-only searches by discovering the salient visual patterns of relevant and irrelevant shots from the approximate relevance provided by text results. The IB reranking method, based on a rigorous Information Bottleneck (IB) principle, finds the optimal clustering of images that preserves the maximal mutual information between the search relevance and the high-dimensional low-level visual features of the images in the text search results. Evaluating the approach on the TRECVID 2003-2005 data sets shows significant improvement upon the text search baseline, with relative increases in average performance of up to 23%. The method requires no image search examples from the user, but is competitive with other state-of-the-art example-based approaches. The method is also highly generic and performs comparably with sophisticated models which are highly tuned for specific classes of queries, such as named-persons. Our experimental analysis has also confirmed the proposed reranking method works well when there exist sufficient recurrent visual patterns in the search results, as often the case in multi-source news videos.
Winston H. Hsu, Lyndon S. Kennedy, Shih-Fu Chang
ACM Multimedia3
2006 New semi-fragile image authentication watermarking techniques using random bias and nonuniform quantization
abstract
Semi-fragile watermarking techniques aim at detecting malicious manipulations on an image, while allowing acceptable manipulations such as lossy compression. Although both of these manipulations are considered to be pixel value changes, semi-fragile watermarks should be sensitive to malicious manipulations but robust to the degradation introduced by lossy compression and other defined acceptable manipulations. In this paper, after studying the characteristics of both natural images and malicious manipulations, we propose two new semi-fragile authentication techniques robust against lossy compression, using random bias and nonuniform quantization, to improve the performance of the methods proposed by Lin and Chang.
Kurato Maeno, Qibin Sun, Shih-Fu Chang, Masayuki Suto
IEEE Trans. Multim.3
2005 Combining text and audio-visual features in video indexing
abstract
We discuss the opportunities, state of the art, and open research issues in using multi-modal features in video indexing. Specifically, we focus on how imperfect text data obtained by automatic speech recognition (ASR) may be used to help solve challenging problems, such as story segmentation, concept detection, retrieval, and topic clustering. We review the frameworks and machine learning techniques that are used to fuse the text features with audio-visual features. Case studies showing promising performance are described, primarily in the broadcast news video domain.
Shih-Fu Chang, R. Manmatha, Tat-Seng Chua
ICASSP (5)1
2005 Commercial Detection in Heterogeneous Video Streams Using Fused Multi-Modal and Temporal Features
abstract
We provide an integrated approach for detecting commercial segments in video streams. This approach systematically fuses the "local" multi-modal characteristics of commercials in the context of their "global" temporal behavior throughout the video stream. Discriminative classifiers are employed to distinguish between commercial and program segments based on their local multi-modal features. The decisions made by different discriminators are fused using a support vector machine. The fusion results are then used as the probabilistic outcomes of a generative model describing the transitions between the commercial and program segments, with explicit models for the inter-arrival times of the commercial segments throughout the video. This approach aims to enhance upon the simple, yet effective, blank frames, which usually indicate the start of commercials. It also provides acceptable performance when such indicators do not exist in the program stream. The results of comprehensive experiments on a heterogeneous data set of 36 hours of video taken from 6 different sources are reported. Our method provides almost 92% correct detection of the commercial segments and 8% enhancement over just using the blank indicators. For the case when blank indicators do not exist, our approach results in almost 85% correct detection.
Masami Mizutani, Shahram Ebadollahi, Shih-Fu Chang
ICASSP (2)3
2005 Layered dynamic mixture model for pattern discovery in asynchronous multi-modal streams [video applications]
abstract
We propose a layered dynamic mixture model for asynchronous multi-modal fusion for unsupervised pattern discovery in video. The lower layer of the model uses generative temporal structures such as a hierarchical hidden Markov model to convert the audiovisual streams into mid-level labels, it also models the correlations in text with probabilistic latent semantic analysis. The upper layer fuses the statistical evidence across diverse modalities with a flexible meta-mixture model that assumes loose temporal correspondence. Evaluation on a large news database shows that multi-modal clusters have better correspondence to news topics than audio-visual clusters alone; novel analysis techniques suggest that meaningful clusters occur when the prediction of salient features by the model concurs with those shown in the story clusters.
Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin
ICASSP (2)3
2005 Automatic discovery of query-class-dependent models for multimodal search
abstract
We develop a framework for the automatic discovery of query classes for query-class-dependent search models in multimodal retrieval. The framework automatically discovers useful query classes by clustering queries in a training set according to the performance of various unimodal search methods, yielding classes of queries which have similar fusion strategies for the combination of unimodal components for multimodal search. We further combine these performance features with the semantic features of the queries during clustering in order to make discovered classes meaningful. The inclusion of the semantic space also makes it possible to choose the correct class for new, unseen queries, which have unknown performance space features. We evaluate the system against the TRECVID 2004 automatic video search task and find that the automatically discovered query classes give an improvement of 18% in MAP over hand-defined query classes used in previous works. We also find that some hand-defined query classes, such as "Named Person" and "Sports" do, indeed, have similarities in search method performance and are useful for query-class-dependent multimodal search, while other hand-defined classes, such as "Named Object" and "General Object" do not have consistent search method performance and should be split apart or replaced with other classes. The proposed framework is general and can be applied to any new domain without expert domain knowledge.
Lyndon S. Kennedy, Apostol Natsev, Shih-Fu Chang
ACM Multimedia3
2005 Physics-motivated features for distinguishing photographic images and computer graphics
abstract
The increasing photorealism for computer graphics has made computer graphics a convincing form of image forgery. Therefore, classifying photographic images and photorealistic computer graphics has become an important problem for image forgery detection. In this paper, we propose a new geometry-based image model, motivated by the physical image generation process, to tackle the above-mentioned problem. The proposed model reveals certain physical differences between the two image categories, such as the gamma correction in photographic images and the sharp structures in computer graphics. For the problem of image forgery detection, we propose two levels of image authenticity definition, i.e., imaging-process authenticity and scene authenticity, and analyze our technique against these definitions. Such definition is important for making the concept of image authenticity computable. Apart from offering physical insights, our technique with a classification accuracy of 83.5% outperforms those in the prior work, i.e., wavelet features at 80.3% and cartoon features at 71.0%. We also consider a recapturing attack scenario and propose a counter-attack measure. In addition, we constructed a publicly available benchmark dataset with images of diverse content and computer graphics of high photorealism.
Tian-Tsong Ng, Shih-Fu Chang, Jessie Hsu, Lexing Xie, Mao-Pei Tsui
ACM Multimedia2
2005 A Framework for Sub-Window Shot Detection
abstract
Browsing a digital video library can be very tedious especially with an ever expanding collection of multimedia material. We present a novel framework for extracting sub-window shots from MPEG encoded news video with the expectation that this will be another tool that can be used by retrieval systems. Sub-windows shots are also useful for tying in relevant material from multiple video sources. The system makes use of Macroblock parameters to extract visual features, which are then combined to identify possible sub-windows in individual frames. The identified sub-widows are then filtered by a non-linear Spatial-Temporal filter to produce sub-window shots. By working only on compressed domain information, this system avoids full frame decoding of MPEG sequences and hence achieves high speeds of up to 11 times real time.
Chuohao Yeo, Yongwei Zhu, Qibin Sun, Shih-Fu Chang
MMM4
2005 Video Adaptation: Concepts, Technologies, and Open Issues
abstract
Video adaptation is an emerging field that offers a rich body of techniques for answering challenging questions in pervasive media applications. It transforms the input video(s) to an output in video or augmented multimedia form by utilizing manipulations at multiple levels (signal, structural, or semantic) in order to meet diverse resource constraints and user preferences while optimizing the overall utility of the video. There has been a vast amount of activity in research and standard development in this area. This paper first presents a general framework that defines the fundamental entities, important concepts (i.e., adaptation, resource, and utility), and formulation of video adaptation as constrained optimization problems. A taxonomy is used to classify different types of adaptation techniques. The state of the art in several active research areas is reviewed with open challenging issues identified. Finally, support of video adaptation from related international standards is discussed.
Shih-Fu Chang, Anthony Vetro
Proc. IEEE1
2005 Classification-based multidimensional adaptation prediction for scalable video coding using subjective quality evaluation
abstract
Scalable video coding offers a flexible representation for video adaptation in multiple dimensions comprising spatial detail and temporal resolution, thus providing great benefits for universal media access (UMA) applications. However, currently most of the approaches address the multidimensional adaptation (MDA) problem in an ad hoc manner. One challenging issue affecting the systematic MDA solution is the difficulty in constructing analytical models in theoretical optimization that capture the relations between video utility and MDA operations. In this paper, we propose a general classification-based prediction framework for selecting the preferred MDA operations based on subjective quality evaluation. For this purpose, we first apply domain-specific knowledge or general unsupervised clustering to construct distinct categories within which the videos share similar preferred MDA operations. Thereafter, a machine learning based method is applied where the low level content features extracted from the compressed video streams are employed to train a framework for the problem of joint signal-to-noise ratio (SNR)-temporal adaptation selection based on the motion compensated three-dimensional subband coding (MC-3DSBC) system. We conduct extensive subjective tests involving 31 subjects, 128 video clips, and formal subjective quality metrics. Statistical analysis of the experimental results confirms the excellent accuracy in using domain knowledge and content features to predict the MDA operation.
Mihaela van der Schaar, Shih-Fu Chang, Alexander C. Loui
IEEE Trans. Circuits Syst. Video Technol.3
2005 A secure and robust digital signature scheme for JPEG2000 image authentication
abstract
In this paper, we present a secure and robust content-based digital signature scheme for verifying the authenticity of JPEG2000 images quantitatively, in terms of a unique concept named lowest authenticable bit rates (LABR). Given a LABR, the authenticity of the watermarked JPEG2000 image will be protected as long as its final transcoded bit rate is not less than the LABR. The whole scheme, which is extended from the crypto data-based digital signature scheme, mainly comprises signature generation/verification, error correction coding (ECC) and watermark embedding/extracting. The invariant features, which are generated from fractionalized bit planes during the procedure of embedded block coding with optimized truncation in JPEG2000, are coded and signed by the sender's private key to generate one crypto signature (hundreds of bits only) per image, regardless of the image size. ECC is employed to tame the perturbations of extracted features caused by processes such as transcoding. Watermarking only serves to store the check information of ECC. The proposed solution can be efficiently incorporated into the JPEG2000 codec (Part 1) and is also compatible with Public Key Infrastructure. After detailing the proposed solution, system performance on security as well as robustness will be evaluated.
Qibin Sun, Shih-Fu Chang
IEEE Trans. Multim.2
2004 Automatic View Recognition in Echocardiogram Videos Using Parts-Based Representation
Shahram Ebadollahi, Shih-Fu Chang, Henry D. Wu
CVPR (2)2
2004 News video story segmentation using fusion of multi-level multi-modal features in TRECVID 2003
abstract
We present our new results in news video story segmentation and classification in the context of the TRECVID video retrieval benchmarking event 2003. We applied and extended the maximum entropy statistical model to fuse diverse features effectively from multiple levels and modalities, including visual, audio, and text. We have included various features such as motion, face, music/speech types, prosody, and high-level text segmentation information. The statistical fusion model is used to discover automatically relevant features contributing to the detection of story boundaries. One novel aspect of our method is the use of a feature wrapper to address different types of features - asynchronous, discrete, continuous and delta ones. We also developed several novel features related to prosody. Using the large news video set from the TRECVID 2003 benchmark, we demonstrate satisfactory performance (F1 measure up to 0.76) and, more importantly, observe an interesting opportunity for further improvement.
Winston H. Hsu, Lyndon S. Kennedy, Chih-Wei Huang, Shih-Fu Chang, Ching-Yung Lin, Giridharan Iyengar
ICASSP (3)4
2004 Video mining: pattern discovery versus pattern recognition
abstract
We examine the significance of video mining as pattern discovery in multimedia content. We examine the underlying issue of pattern discovery versus pattern recognition, since most past work has not drawn such a sharp distinction. We argue that while the term "pattern discovery" implies a purely unsupervised approach, in practice a mixture of unsupervised and supervised techniques will have to be used. We compare conventional data mining with video mining and observe that a key difference is in the multilayered semantics of multimedia content. We then identify significant challenges posed by video mining.
Ajay Divakaran, Kadir A. Peker, Shih-Fu Chang, Regunathan Radhakrishnan, Lexing Xie
ICIP3
2004 A model for image splicing
abstract
The ease of creating image forgery using image-splicing techniques will soon make our naive trust on image authenticity a tiling of the past. In prior work, we observed the capability of the bicoherence magnitude and phase features for image splicing detection. To bridge the gap between empirical observations and theoretical justifications, in this paper, an image-splicing model based on the idea of bipolar signal perturbation is proposed and studied. A theoretical analysis of the model leads to propositions and predictions consistent with the empirical observations.
Tian-Tsong Ng, Shih-Fu Chang
ICIP2
2004 Discovering meaningful multimedia patterns with audio-visual concepts and associated text
abstract
The work presents the first effort to automatically annotate the semantic meanings of temporal video patterns obtained through unsupervised discovery processes. This problem is interesting in domains where neither perceptual patterns nor semantic concepts have simple structures. The patterns in video are modeled with hierarchical hidden Markov models (HHMM), with efficient algorithms to learn the parameters, the model complexity and the relevant features; the meanings are contained in words of the speech transcript of the video. The pattern-word association is obtained via cooccurrence analysis and statistical machine translation models. Promising results are obtained through extensive experiments on 20+ hours of TRECVID news videos: video patterns that associate with distinct topics such as el-nino and politics are identified; the HHMM temporal structure model compares favorably to a nontemporal clustering algorithm.
Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin
ICIP3
2004 Generative, discriminative, and ensemble learning on multi-modal perceptual fusion toward news video story segmentation
abstract
News video story segmentation is a critical task for automatic video indexing and summarization. Our prior work has demonstrated promising performance by using a generative model, called maximum entropy (ME), which models the posterior probability given the multi-modal perceptual features near the candidate points. We investigate alternative statistical approaches based on discriminative models, i.e. support vector machine (SVM), and ensemble learning, i.e. boosting. In addition, we develop a novel approach, called BoostME, which uses the ME classifiers and the associated confidence scores in each boosting iteration. We evaluated these different methods using the TRECVID 2003 broadcast news data set. We found that SVM-based and ME-based techniques both outperformed the pure boosting techniques, with the SVM-based solutions achieving even slightly higher accuracy. Moreover, we summarize extensive analysis results of error sources over distinctive news story types to identify future research opportunities.
Winston H. Hsu, Shih-Fu Chang
ICME2
2004 Understanding and modeling user interests in consumer videos
abstract
The paper analyzes the interests of users in viewing and organizing consumer videos. It proposes a taxonomy of relevant concepts with three basic dimensions of interests (DOIs) and effective models to predict the user interests in each dimension. The three DOIs correspond to the objects, the scenes and the events. Our conclusions are backed with an extensive study, in which users were asked to annotate and score the importance of each DOI in short clips of diverse and real consumer videos. Analysis of the user study data reveals high consistency (70%) of the scores across different users and independence between objects and events. In addition, we show how heuristic rules and neural networks can accurately predict these scores using camera motion, foreground object and audio information. The automatic and effective prediction of user interests has the potential for improving applications for annotating and summarizing consumer videos.
Ryoma Oami, Ana B. Benitez, Shih-Fu Chang, Nevenka Dimitrova
ICME3
2004 A crypto signature scheme for image authentication over wireless channel
abstract
With the ambient use of digital images and the increasing concern on their integrity and originality, consumers are facing an emergent need of authenticating degraded images despite lossy compression and packet loss. In this paper, we propose a scheme to meet this need by incorporating a watermarking solution into a traditional crypto signature scheme to make the digital signatures robust to image degradations. The proposed approach is compatible with traditional crypto signature schemes except that the original image needs to be watermarked in order to guarantee the robustness of its derived digital signature. We demonstrate the effectiveness of this proposed scheme through practical experimental results as well as illustrative analysis.
Qibin Sun, Shuiming Ye, Ching-Yung Lin, Shih-Fu Chang
ICME4
2004 Subjective preference of spatio-temporal rate in video adaptation using multi-dimensional scalable coding
abstract
Video adaptation allows for direct manipulation of existing encoded video streams to meet new resource constraints without having to encode the video from scratch. Multi-dimensional scalable coding, such as motion-compensated subband coding (MCSBC), offers an effective and flexible representation for video adaptation. In order to develop robust criteria for selecting optimal spatio-temporal rates used in adaptation, knowledge about subjective preference of spatio-temporal rates is needed. We study the optimal temporal frame rate over a wide range of bandwidth (50 kbps to 1 Mbps) using subjective quality evaluation with 128 clips and 31 subjects. We analyze the results using statistical testing methods and investigate the dependence of optimal frame rate on user, bandwidth, and video content characteristics. Our findings indicate agreement among most users and the existence of switching bandwidths at which preferred frame rates change. Dependence of the preference on video content types is also revealed.
Shih-Fu Chang, Alexander C. Loui
ICME2
2004 Color-mood analysis of films based on syntactic and psychological models
abstract
The emergence of peer-to-peer networking and the increase of home PC storage capacity are necessitating efficient scaleable methods for video clustering, recommending and browsing. Based on film theories and psychological models, color-mood is an important factor affecting user emotional preferences. We propose a compact set of features for color-mood analysis and subgenre discrimination. We introduce two color representations for scenes and full films in order to extract the essential moods from the films: a global measure for the color palette and a discriminative measure for the transitions of the moods in the movie. We captured the dominant color ratio and the pace of the movie. Despite the simplicity and efficiency of the features, the classification accuracy was surprisingly good, about 80%, possibly thanks to the prevalence of the color-mood association in feature films.
Cheng-Yu Wei, Nevenka Dimitrova, Shih-Fu Chang
ICME3
2004 Semantic video clustering across sources using bipartite spectral clustering
abstract
Data clustering is an important technique for visual data management. Most previous work focuses on clustering video data within single sources. We address the problem of clustering across sources, and propose novel spectral clustering algorithms for multisource clustering problems. Spectral clustering is a new discriminative method realizing clustering by partitioning data graphs. We represent multi-source data as bipartite or K-partite graphs, and investigate the spectral clustering algorithm under these representations. The algorithms are evaluated using the TRECVID-2003 corpus with semantic features extracted from speech transcripts and visual concept recognition results from videos. The experiments show that the proposed bipartite clustering algorithm significantly outperforms the regular spectral clustering algorithm in capturing cross-source associations.
DongQing Zhang, Ching-Yung Lin, Shih-Fu Chang, John R. Smith
ICME3
2004 Story boundary detection in large broadcast news video archives: techniques, experience and trends
abstract
The segmentation of news video into story units is an important step towards effective processing and management of large news video archives. In the story segmentation task in TRECVID 2003, a wide variety of techniques were employed by many research groups to segment over 120-hour of news video. The techniques employed range from simple anchor person detector to soisticated machine learning models based on HMM and Maximum Entropy (ME) approaches. The general results indicate that the judicious use of multi-modality features coupled with rigorous machine learning models could produce effective solutions. This paper presents the algorithms and experience learned in TRECVID evaluations. It also points the way towards the development of scalable technology to process large news video corpuses.
Tat-Seng Chua, Shih-Fu Chang, Lekha Chaisorn, Winston H. Hsu
ACM Multimedia2
2004 Detecting image near-duplicate by stochastic attributed relational graph matching with learning
abstract
Detecting Image Near-Duplicate (IND) is an important problem in a variety of applications, such as copyright infringement detection and multimedia linking. Traditional image similarity models are often difficult to identify IND due to their inability to capture scene composition and semantics. We present a part-based image similarity measure derived from stochastic matching of Attributed Relational Graphs that represent the compositional parts and part relations of image scenes. Such a similarity model is fundamentally different from traditional approaches using low-level features or image alignment. The advantage of this model is its ability to accommodate spatial attributed relations and support supervised and unsupervised learning from training data. The experiments compare the presented model with several prior similarity models, such as color histogram, local edge descriptor, etc. The presented model outperforms the prior approaches with large margin.
DongQing Zhang, Shih-Fu Chang
ACM Multimedia2
2004 Predicting optimal operation of MC-3DSBC multidimensional scalable video coding using subjective quality measurement
abstract
Recently we have witnessed a growing interest in the development of the subband/wavelet coding (SBC) technology, partly due to the superior scalability of SBC. Scalable coding provides great synergy with the universal media access applications, where media content is delivered to client devices of diverse types through heterogeneous channels. In this respect, SBC system provides flexibility in realizing different ways of media scaling, including scaling dimensions of SNR, spatial, and temporal. However, the selection of specific scalability operations given the bit rate constraint has always been ad hoc - a systematic methodology is missing. In this paper, we address this open issue by applying our content-based optimal scalability selection framework and adopting subjective quality evaluation. For this purpose we firstly explore the behavior of SNR-Spatial-Temporal scalability using Motion Compensated (MC) SBC systems. Based on the system behavior, we propose an efficient method for the optimal selection of scalability operator through content-based prediction. Our experiment results demonstrate that the proposed method can efficiently predict the optimal scalability operation with an excellent accuracy.
Tian-Tsong Ng, Mihaela van der Schaar, Shih-Fu Chang
VCIP4
2004 Multimedia database management systems
John R. Smith, Tong Zhang 0007, Shih-Fu Chang
J. Vis. Commun. Image Represent.4
2004 Real-time view recognition and event detection for sports video
Di Zhong, Shih-Fu Chang
J. Vis. Commun. Image Represent.2
2004 Structure analysis of soccer video with domain knowledge and hidden Markov models
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
Pattern Recognit. Lett.3
2003 A Bayesian Framework for Fusing Multiple Word Knowledge Models in Videotext Recognition
abstract
Videotext recognition is challenging due to low resolution, diverse fonts/styles, and cluttered background. Past methods enhanced recognition by using multiple frame averaging, image interpolation and lexicon correction, but recognition using multi-modality language models has not been explored. In this paper, we present a formal Bayesian framework for videotext recognition by combining multiple knowledge using mixture models, and describe a learning approach based on Expectation-Maximization (EM). In order to handle unseen words, a back-off smoothing approach derived from the Bayesian model is also presented. We exploited a prototype that fuses the model from closed caption and that from the British National Corpus. The model from closed caption is based on a unique time distance distribution model of videotext words and closed caption words. Our method achieves a significant performance gain, with word recognition rate of 76.8% and character recognition rate of 86.7%. The proposed methods also reduce false videotext detection significantly, with a false alarm rate of 8.2% without substantial loss of recall.
DongQing Zhang, Shih-Fu Chang
CVPR (2)2
2003 Image classification using multimedia knowledge networks
abstract
This paper presents novel methods for classifying images based on knowledge discovered from annotated images using WordNet. The novelty of this work is the automatic class discovery and the classifier combination using the extracted knowledge. The extracted knowledge is a network of concepts (e.g., image clusters and word-senses) with associated image and text examples. Concepts that are similar statistically are merged to reduce the size of the concept network. Our knowledge classifier is constructed by training a meta-classifier to predict the presence of each concept in images. A Bayesian network is then learned using the meta-classifiers and the concept network. For a new image, the presence of concepts is first detected using the meta-classifiers and refined using Bayesian inference. Experiments have shown that combining classifiers using knowledge-based Bayesian networks results in superior (up to 15%) or comparable accuracy to individual classifiers and purely statistically learned classifier structures. Another contribution of this work is the analysis of the role of visual and text features in image classification. As text or joint text + visual features perform better in classifying images than visual features, we tried to predict text features for images without annotations; however, the accuracy of visual + predicted text features did not consistently improve over visual features.
Ana B. Benitez, Shih-Fu Chang
ICIP (3)2
2003 Content-based utility function prediction for real-time MPEG-4 video transcoding
abstract
Utility function based transcoding is an efficient systematic solution for choosing optimal media transcoding operation to meet dynamic resource constraints (such as bandwidth). However, to date the real-time generation of utility function is not feasible due to computational complexity. In this paper we present a content-based utility function prediction framework for real-time MPEG-4 video transcoding. We develop a statistical approach combining real-time compressed-domain feature extraction, content-based pattern classification and regression. Our extensive experiment results demonstrate that the proposed method achieves very promising prediction accuracy - up to 89% in choosing the optimal transcoding operation with the highest quality from multiple alternatives meeting the same target bitrate.
Jae-Gon Kim, Shih-Fu Chang
ICIP (1)3
2003 Feature selection for unsupervised discovery of statistical temporal structures in video
abstract
In this paper, we present algorithms for automatic feature selection for of structure discovery from video sequences. Feature selection in this scenario is hard because of the absence of class labels to evaluate against, and the temporal correlation among samples that prevents the direct estimation of posterior probabilities of the cluster given the sequence. The overall problem of structure discovery is formulated as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. Under Markov assumptions among events, structures in the video are modelled with hierarchical hidden Markov models, with efficient algorithms to jointly learn the model parameters and the optimal model complexity. Feature selection iterates between a wrapper step that partitions the large feature pool into consistent subsets, and a filter step that eliminate redundancy within these subsets, respectively. The feature subsets are then ranked according to the normalized Bayesian Information criteria, and the learning results from these ranked subsets can be evaluated and interpreted by a human observer. Results on soccer and baseball videos show that the automatically selected feature set coincides with those selected with domain knowledge and intuition, while achieving a correspondence comparable to that of supervised learning against manually labelled ground truth.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICIP (1)2
2003 A statistical framework for fusing mid-level perceptual features in news story segmentation
abstract
News story segmentation is essential for video indexing, summarization and intelligence exploitation. In this paper, we present a general statistical framework, called exponential model or maximum entropy model that can systematically select the most significant mid-level features of various types (visual, audio, and semantic) and learn the optimal ways in fusing their combinations in story segmentation. The model utilizes a family of weighted, exponential functions to account for the contributions from different features. The Kullbak-Leibler divergence measure is used in an optimization procedure to iteratively estimate the model parameters, and automatically select the optimal features. The framework is scalable in incorporating new features and adapting to new domains and also discovers how these feature sets contribute to the segmentation work. When tested on foreign news programs, the proposed techniques achieve significant performance improvement over prior work using ad hoc algorithms and slightly better gain over the state of the art using HMM-based models.
Winston H. Hsu, Shih-Fu Chang
ICME2
2003 Content-adaptive utility-based video adaptation
abstract
Many techniques exist for adapting videos to satisfy heterogeneous resource conditions or user preferences. However, selections of appropriate adaptations among various choices are often ad hoc. To provide a systematic solution, we present a general conceptual framework to model video entity, adaptation, resource, utility, and relations among them. The framework extends the conventional rate-distortion model in terms of flexibility and generality. It allows for formulations of various adaptation problems as resource-constrained utility maximization. It also facilitates new approaches to predicting critical information about resource-utility relations. We apply the framework to a practical case in dynamic bit rate adaptation. Furthermore, we present a description tool, which has been accepted as a part of the MPEG-21 digital item adaptation (DIA), to support utility-based adaptation.
Jae-Gon Kim, Shih-Fu Chang
ICME3
2003 Unsupervised discovery of multilevel statistical video structures using hierarchical hidden Markov models
abstract
Structure elements in a time sequence (e.g. video) are repetitive segments with consistent deterministic or stochastic characteristics. While most existing work in detecting structures follows a supervised paradigm, we propose a fully unsupervised statistical solution in this paper. We present a unified approach to structure discovery from long video sequences as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. We model the multilevel statistical structure as hierarchical hidden Markov models, and present efficient algorithms for learning both the parameters and the model structure. When tested on a specific domain, soccer video, the unsupervised learning scheme achieves very promising results: it automatically discovers the statistical descriptions of high-level structures, and at the same time achieves even slightly better accuracy in detecting discovered structures in unlabelled videos than a supervised approach designed with domain knowledge and trained with comparable hidden Markov models.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICME2
2003 Survey of compressed-domain features used in audio-visual indexing and analysis
Hualu Wang, Ajay Divakaran, Anthony Vetro, Shih-Fu Chang, Huifang Sun
J. Vis. Commun. Image Represent.4
2003 Special issue on multimedia adaptation
Fernando Pereira 0001, Ian S. Burnett, Shih-Fu Chang
Signal Process. Image Commun.3
2002 Structure analysis of soccer video with hidden Markov models
abstract
In this paper, we present algorithms for parsing the structure of produced soccer programs. The problem is important in the context of a personalized video streaming and browsing system. While prior work focuses on the detection of special events such as goals or corner kicks, this paper is concerned with generic structural elements of the game. We begin by defining two mutually exclusive states of the game, play and break based on the rules of soccer. We select a domain-tuned feature set, dominant color ratio and motion intensity, based on the special syntax and content characteristics of soccer videos. Each state of the game has a stochastic structure that is modeled with a set of hidden Markov models. Finally, standard dynamic programming techniques are used to obtain the maximum likelihood segmentation of the game into the two states. The system works well, with 83.5% classification accuracy and good boundary timing from extensive tests over diverse data sets.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICASSP2
2002 Echocardiogram videos: summarization, temporal segmentation and browsing
abstract
We present a system for the temporal segmentation, summarization, and browsing of the echocardiogram videos. Echocardiogram videos are video sequences produced by the ultrasound scanning of the heart, and are one of the main modalities of imaging the heart structure. Our approach combines the domain-specific knowledge and the automatic analysis of the spatio-temporal structure of the echocardiogram videos. The videos are temporally sampled using the embedded electrocardiogram graph. The consecutive sampled frames are compared based on the shape of the region of interest and the presence/absence of color to detect the boundaries between the different segments. The content of each segment of the video is summarized into two forms: the static and the dynamic summaries. Finally the summary is displayed in the user interface in an intuitive form for the purpose of browsing. Applications include digital medical image libraries, medical image management, and telemedicine.
Shahram Ebadollahi, Shih-Fu Chang, Henry D. Wu
ICIP (1)2
2002 Semi-fragile image authentication using generic wavelet domain features and ECC
abstract
We present a generic content-based solution targeting at authenticating image in a semi-fragile way, which integrates watermarking-based approach with signature-based approach. Robust signatures are cryptographically generated based on invariant features called significance-linked connected component (SLCC) extracted from image content and are then signed and embedded back into the image again as watermarks, all in the wavelet domain. De-noising and morphological filtering are applied as pre-processing to tame some small perturbations on extracted features caused by various incidental distortions introduced in acceptable manipulations such as lossy compression, common image processing (bluffing, sharpening, etc.) as well as watermarking. Error correcting coding is employed to further bridge between generated signatures and watermarks in a novel way: message bits are formed based on SLCC features, and parity check bits are taken as the seeds of watermarks. The generated signature is hashable and can be incorporated into a PKI framework.
Qibin Sun, Shih-Fu Chang
ICIP (2)2
2002 A quantitative semi-fragile JPEG2000 image authentication system
abstract
We propose a novel integrated approach to quantitative semi-fragile authentication of JPEG2000 images under a generic framework which combines ECC and PKI infrastructures. Firstly acceptable manipulations (e.g., re-encoding) which should pass authentication are defined based on considerations of some target applications. We propose a unique concept of lowest authenticable bit rate - images undergoing repetitive re-encoding are guaranteed to pass authentication provided the re-encoding rates are above the lowest allowable bit rate. Our solutions include computation of content-based features derived from the EBCOT encoding procedure of JPEG2000, error correction coding of the derived features, PKI cryptographic signing, and finally robust embedding of the feature codes into image as watermarks.
Qibin Sun, Shih-Fu Chang, Kurato Maeno, Masayuki Suto
ICIP (2)2
2002 Video skims: taxonomies and an optimal generation framework
abstract
This paper presents a new conceptual framework for summarization that considers the relationship between entities, device properties and user information needs. We summarize using a skim - an audio-visual clip that is a drastically condensed version of the original video. An entity is defined to be a sequence of elements that are related to each other by a certain property. In this paper we discuss the causes, the different entity types and also present a skim taxonomy. Each entity is associated with a utility. The skim is generated by a constrained utility maximization over those entity-utilities that satisfy the user information needs as well as the device rendering capabilities. We construct an optimal skim within this framework that retains a particular subset of entities. These entities have been chosen since they can be automatically computed in a robust manner. The user studies show that the optimal skims perform well in a statistically significant sense, at compression rates as high as 90%.
Hari Sundaram, Shih-Fu Chang
ICIP (2)2
2002 General and domain-specific techniques for detecting and recognizing superimposed text in video
abstract
We have developed generic and domain-specific video algorithms for caption text extraction and recognition in digital video. Our system includes several unique features: for caption box location, we combine the compressed-domain features derived from DCT coefficients and motion vectors. Long-term temporal consistency is employed to enhance localization performance. For character segmentation, we use a single-pass threshold free approach combining classification and projection to address noisy segmentation, text intensity variation, and algorithm complexity. In recognition, we use Zernike moments to achieve more accurate recognition performance. Finally, domain knowledge is explored and a statistical transition graph model is used to enhance recognition of domain-specific characters, such as ball counts and game score of baseball videos. The algorithms achieved real-time speed and significantly improved recognition accuracy. Furthermore, although the experiments were conducted in baseball videos only, these algorithms (except the transition model) are general and can be used in other applications, such as news and films.
DongQing Zhang, Raj Kumar Rajendran, Shih-Fu Chang
ICIP (1)3
2002 Perceptual knowledge construction from annotated image collections
abstract
This paper presents and evaluates new methods for extracting perceptual knowledge from collections of annotated images. The proposed methods include automatic techniques for constructing perceptual concepts by clustering the images based on visual and text feature descriptors, and for discovering perceptual relationships among the concepts based on descriptor similarity and statistics between the clusters. There are two main contributions of this work. The first lies on the support and the evaluation of several techniques for visual and text feature descriptor extraction, for visual and text feature descriptor integration, and for data clustering in the extraction of perceptual concepts. The second contribution is in proposing novel ways for discovering perceptual relationships among concepts. Experiments show extraction of useful knowledge from visual and text feature descriptors, high independence between visual and text feature descriptors, and potential performance improvement by integrating both kinds of descriptors compared to using either kind of descriptor alone.
Ana B. Benitez, Shih-Fu Chang
ICME (1)2
2002 Semantic knowledge construction from annotated image collections
abstract
This paper presents new methods for extracting semantic knowledge from collections of annotated images. The proposed methods include novel automatic techniques for extracting semantic concepts by disambiguating the senses of words in annotations using the lexical database WordNet, using both the images and their annotations, and for discovering semantic relations among the detected concepts based on WordNet. Another contribution of this paper is the evaluation of several techniques for visual feature descriptor extraction and data clustering in the extraction of semantic concepts. Experiments show the potential of integrating the analysis of both images and annotations for improving the performance of the word-sense disambiguation process. In particular, the accuracy improves 4-15% with respect to the baselines systems for nature images.
Ana B. Benitez, Shih-Fu Chang
ICME (2)2
2002 Duplicate detection in consumer photography and news video
abstract
Consumers often make more than one photograph of the same scene, creating non-identical duplicates and near duplicates. In Kodak's consumer photography database, on average, 19% of the images, per roll, fall into this category. Automatic detection of duplicates, therefore, is extremely useful in applications that help users organize their image collections. We introduce the challenging problem of non-identical duplicate image detection in consumer photography, describe STELLA (a novel interactive personal image collection organization system), and give an overview of our novel framework for detecting duplicate and near duplicate consumer photographs and news videos.
Alejandro Jaimes, Shih-Fu Chang, Alexander C. Loui
ACM Multimedia2
2002 A utility framework for the automatic generation of audio-visual skims
abstract
In this paper, we present a novel algorithm for generating audio-visual skims from computable scenes. Skims are useful for browsing digital libraries, and for on-demand summaries in set-top boxes. A computable scene is a chunk of data that exhibits consistencies with respect to chromaticity, lighting and sound. There are three key aspects to our approach: (a) visual complexity and grammar, (b) robust audio segmentation and (c) an utility model for skim generation. We define a measure of visual complexity of a shot, and map complexity to the minimum time for comprehending the shot. Then, we analyze the underlying visual grammar, since it makes the shot sequence meaningful. We segment the audio data into four classes, and then detect significant phrases in the speech segments. The utility functions are defined in terms of complexity and duration of the segment. The target skim is created using a general constrained utility maximization procedure that maximizes the information content and the coherence of the resulting skim. The objective function is constrained due to multimedia synchronization constraints, visual syntax and by penalty functions on audio and video segments. The user study results indicate that the optimal skims show statistically significant differences with other skims with compression rates up to 90%.
Hari Sundaram, Lexing Xie, Shih-Fu Chang
ACM Multimedia3
2002 Event detection in baseball video using superimposed caption recognition
abstract
We have developed a novel system for baseball video event detection and summarization using superimposed caption text detection and recognition. The system detects different types of semantic level events in baseball video including scoring and last pitch of each batter. The system has two components: event detection and event boundary detection. Event detection is realized by change detection and recognition of game stat texts (such as text information showing in score box). Event boundary detection is achieved using our previously developed algorithm, which detects the pitch view as the event beginning and nonactive view as potential endings of the event. One unique contribution of the system is its capability to accurately detect the semantic level events by combining video text recognition with camera view recognition. Another unique feature is the real-time processing speed by taking advantage of compressed-domain approaches in part of the algorithms such as caption detection. To the best of our knowledge, this is the first system achieving accurate detection of multiple types of high-level semantic events in baseball videos.
DongQing Zhang, Shih-Fu Chang
ACM Multimedia2
2002 Special issue on multimedia adaptation
Fernando Pereira 0001, Ian S. Burnett, Shih-Fu Chang
Signal Process. Image Commun.3
2002 Computable scenes and structures in films
abstract
We present a computational scene model and also derive novel algorithms for computing audio and visual scenes and within-scene structures in films. We use constraints derived from film-making rules and from experimental results in the psychology of audition, in our computational scene model. Central to the computational model is the notion of a causal, finite-memory viewer model. We segment the audio and video data separately. In each case, we determine the degree of correlation of the most recent data in the memory with the past. The audio and video scene boundaries are determined using local maxima and minima, respectively. We derive four types of computable scenes that arise due to different kinds of audio and video scene boundary synchronizations. We show how to exploit the local topology of an image sequence in conjunction with statistical tests, to determine dialogs. We also derive a simple algorithm to detect silences in audio. An important feature of our work is to introduce semantic constraints based on structure and silence in our computational model. This results in computable scenes that are more consistent with human observations. The algorithms were tested on a difficult data set: three commercial films. We take the first hour of data from each of the three films. The best results: computational scene detection: 94%; dialogue detection: 91%; and recall 100% precision.
Hari Sundaram, Shih-Fu Chang
IEEE Trans. Multim.2
2001 MPEG-7 MDS Content Description Tools and Applications
Ana B. Benitez, Di Zhong, Shih-Fu Chang, John R. Smith
CAIP3
2001 VISMap: an interactive image/video retrieval system using visualization and concept maps
abstract
Images and videos can be indexed by multiple features at different levels, such as color, texture, motion and text annotation. Organizing this information into a system so that users can query effectively is a challenging and important problem. We present VISMap, a visual information seeking system that extends the traditional query paradigms of query-by-example and query-by-sketch and replaces the models of relevance feedback with principles from information visualization and concept representation. Users no longer perform lengthy "one-shot" queries or rely on hidden relevance feedback mechanisms. Instead, we provide a rich set of tools that allow users to construct personal views of the video database and directly visualize and manipulate various views and comprehend effects of individual query criteria on the final search results. The set of tools include: (1) a feature space browser for feature-based exploration and navigation; (2) a distance map for metric comparison and setting and (3) a novel concept map for query representation and creation.
William Chen 0001, Shih-Fu Chang
ICIP (3)2
2001 Zero-error information hiding capacity of digital images
abstract
We derive a theoretical capacity for digital image watermarking with zero transmission errors. We present three different discrete memoryless channel model to represent the watermarking process. Given the magnitude bound of noise set by applications and the acceptable watermark magnitude determined by the just-noticeable distortion, we estimate the zero-error capacity by applying Shannon's (1948) adjacency-reducing mapping technique. The capacity we estimate here corresponds to a deterministic guarantee of zero error, different from the traditional theorem approaching zero error asymptotically.
Ching-Yung Lin, Shih-Fu Chang
ICIP (3)2
2001 Long-term moving object segmentation and tracking using spatio-temporal consistency
abstract
The success of object-based media representation and description (e.g., MPEG-4 and -7) depends largely on effective object segmentation tools. We expand our previous work on automatic video region tracking and develop a robust-moving objects detection system. In our system, we first utilize innovative methods of combining color and edge information in improving the object motion estimation results. Then we use the long-term spatio-temporal constraints to achieve reliable object tracking over long sequences. Our extensive experiments demonstrate excellent results in handling challenging cases in general domains (e.g., stock footage) including depth-varying multi-layer background and fast camera motion.
Di Zhong, Shih-Fu Chang
ICIP (2)2
2001 Condensing Computable Scenes Using Visual Complexity And Film Syntax Analysis
abstract
In this paper, we present a novel algorithm to condense computable scenes. A computable scene is a chunk of data that exhibits consistencies with respect to chromaticity, lighting and sound. We attempt to condense such scenes in two ways. First, we define visual complexity of a shot to be its Kolmogorov complexity. Then, we conduct experiments that help us map the complexity of a shot into the minimum time required for its comprehension. Second, we analyze the grammar of the film language, since it makes the shot sequence meaningful. These grammatical rules are used to condense scenes, in parallel to the shot level condensation. We've implemented a system that generates a skim given a time budget. Our user studies show good results on skims with compression rates between 60-80%.
Hari Sundaram, Shih-Fu Chang
ICME2
2001 Algorithms And System For Segmentation And Structure Analysis In Soccer Video
abstract
In this paper, we present a novel system and effective algorithms for soccer video segmentation. The output, about whether the ball is in play, reveals high-level structure of the content. The first step is to classify each sample frame into 3 kinds of view using a unique domain-specific feature, grass-area-ratio. Here the grass value and classification rules are learned and automatically adjusted to each new clip. Then heuristic rules are used in processing the view label sequence, and obtain play/break status of the game. The results provide good basis for detailed content analysis in next step. We also show that lowlevel features and mid-level view classes can be combined to extract more information about the game, via the example of detecting grass orientation in the field. The results are evaluated under different metrics intended for different applications; the best result in segmentation is 86.5%. 1.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Anthony Vetro, Huifang Sun
ICME3
2001 Structure Analysis of Sports Video Using Domain Models
abstract
In this paper, we present an effective framework for scene detection and structure analysis for sports videos, using tennis and baseball as examples. Sports video can be characterized by its predictable temporal syntax, recurrent events with consistent features, and a fixed number of views. Our approach combines domain-specific knowledge, supervised machine learning techniques, and automatic feature analysis at multiple levels. Real time processing performance is achieved by utilizing compressed-domain processing techniques. High accuracy in view recognition is achieved by using compressed-domain global features as prefilters and object-level refined analysis in the latter verification stage. Applications include high-level structure browsing/navigation, highlight generation, and mobile
Di Zhong, Shih-Fu Chang
ICME2
2001 IMKA: a multimedia organization system combining perceptual and semantic knowledge
abstract
In the demo, we present the IMKA system, which implements the innovative approach of integrating perceptual information such as low-level features and images, and symbolic information such as words to represent the knowlege associated with a large multimedia collection for multimedia organization and retrieval. The IMKA system utilizes the unique MediaNet framework, which greatly extends existing knowlege representation tools in the text domain (e.g., semantic networks and WordNet) and the multimedia domain (e.g. Multimedia Thesaurus) by combining perceptual and semantic concepts in the same network and by supporting perceptual and semantic relationships among concepts exemplified by different media. It also brings the level of multimedia retrieval closer to users' needs by translating low-level feature queries to high-level semantic queries and vice versa. We will demonstrate the process of constructing the MediaNet knowledge base and new ways of searching multimedia in the IMKA system by presenting the current implementation of the IMKA system that uses image collections from online sources.
Ana B. Benitez, Shih-Fu Chang, John R. Smith
ACM Multimedia2
2001 SARI: self-authentication-and-recovery image watermarking system
abstract
In this project, we designed a novel image authentication system based on a our semi-fragile watermarking technique. The system, called SARI, can accept quantization-based lossy compression to a determined degree without any false alarm and can sensitively detect and locate malicious manipulations. It's the first system that has such capability in distinguishing malicious attacks from acceptable operations. Furthermore, the corrupted area can be approximately recovered by the information hidden in the other part of the contentimage. The amount of information embedded in our SARI system has nearly reached the theoretical maximum zero-error information hiding capacity of digital images. The software prototype includes two parts - the watermark embedder that's freely distributed and the authenticator that can be deployed online as a third-party service or used in the recipient side.
Ching-Yung Lin, Shih-Fu Chang
ACM Multimedia2
2001 Real-time personalized sports video filtering and summarization
abstract
We demonstrate a real-time fully automated software system for filtering important events in sports video. Events represent occurrences of actions or state changes in video content. In the current prototype, we demonstrate detection of pitching in baseball and serving in tennis. For wireless video applications, we propose and apply a unique notion of content-based adaptive streaming, in which video encoding rate and media modality is dynamically varied according to the event filtering results. Our system includes an event detection module, an adaptive encoding module, and a buffer management module for adaptive streaming. We achieve the real-time performance by exploring compresseddomain techniques and multi-stage multi-resolution contentanalysis processes.
Di Zhong, Shih-Fu Chang
ACM Multimedia3
2001 A conceptual framework and empirical research for classifying visual descriptors
abstract
Abstract This article presents exploratory research evaluating a conceptual structure for the description of visual content of images. The structure, which was developed from empirical research in several fields (e.g., Computer Science, Psychology, Information Studies, etc.), classifies visual attributes into a “Pyramid” containing four syntactic levels (type/technique, global distribution, local structure, composition), and six semantic levels (generic, specific, and abstract levels of both object and scene, respectively). Various experiments are presented, which address the Pyramid's ability to achieve several tasks: (1) classification of terms describing image attributes generated in a formal and an informal description task, (2) classification of terms that result from a structured approach to indexing, and (3) guidance in the indexing process. Several descriptions, generated by naive users and indexers, are used in experiments that include two image collections: a random Web sample, and a set of news images. To test descriptions generated in a structured setting, an Image Indexing Template (developed independently over several years of this project by one of the authors) was also used. The experiments performed suggest that the Pyramid is conceptually robust (i.e., can accommodate a full range of attributes), and that it can be used to organize visual content for retrieval, to guide the indexing process, and to classify descriptions obtained manually and automatically.
Corinne Jörgensen, Alejandro Jaimes, Ana B. Benitez, Shih-Fu Chang
J. Assoc. Inf. Sci. Technol.4
2001 Introduction to the special issue on MPEG-7
abstract
MPEG-7 has ignited interest in industry with broad interests in media, content management, content provision, broadcasting, computers, telecommunications, as well as researchers and academia. Overall, the MPEG-7 standard is currently in a fairly mature stage and is on its way to final standardization later this year. The pace of progress in development of MPEG-7 has been quite frantic. This special issue provides snapshot of the status of MPEG-7 technology as of around April 2001. This special issue contains only invited papers from experts who are involved in MPEG-7 standardization. It is designed to be useful to beginners as well as experts alike as it covers various topics at different depths.
Shih-Fu Chang, Atul Puri, Thomas Sikora, HongJiang Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2001 Overview of the MPEG-7 standard
abstract
MPEG-7, formally known as the Multimedia Content Description Interface, includes standardized tools (descriptors, description schemes, and language) enabling structural, detailed descriptions of audio-visual information at different granularity levels (region, image, video segment, collection) and in different areas (content description, management, organization, navigation, and user interaction). It aims to support and facilitate a wide range of applications, such as media portals, content broadcasting, and ubiquitous multimedia. We present a high-level overview of the MPEG-7 standard. We first discuss the scope, basic terminology, and potential applications. Next, we discuss the constituent components. Then, we compare the relationship with other standards to highlight its capabilities.
Shih-Fu Chang, Thomas Sikora, Atul Puri
IEEE Trans. Circuits Syst. Video Technol.1
2001 A robust image authentication method distinguishing JPEG compression from malicious manipulation
abstract
Image authentication verifies the originality of an image by detecting malicious manipulations. Its goal is different from that of image watermarking, which embeds into the image a signature surviving most manipulations. Most existing methods for image authentication treat all types of manipulation equally (i.e., as unacceptable). However, some practical applications demand techniques that can distinguish acceptable manipulations (e.g., compression) from malicious ones. In this paper, we present an effective technique for image authentication which can prevent malicious manipulations but allow JPEG lossy compression. The authentication signature is based on the invariance of the relationships between discrete cosine transform (DCT) coefficients at the same position in separate blocks of an image. These relationships are preserved when DCT coefficients are quantized in JPEG compression. Our proposed method can distinguish malicious manipulations from JPEG lossy compression regardless of the compression ratio or the number of compression iterations. We describe adaptive methods with probabilistic guarantee to handle distortions introduced by various acceptable manipulations such as integer rounding, image filtering, image enhancement, or scaling-recaling. We also present theoretical and experimental results to demonstrate the effectiveness of the technique.
Ching-Yung Lin, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.2
2000 Audio scene segmentation using multiple features, models and time scales
abstract
We present an algorithm for audio scene segmentation. An audio scene is a semantically consistent sound segment that is characterized by a few dominant sources of sound. A scene change occurs when a majority of the sources present in the data change. Our segmentation framework has three parts: a definition of an audio scene; multiple feature models that characterize the dominant sources; and a simple, causal listener model, which mimics human audition using multiple time-scales. We define a correlation function that determines correlation with past data to determine segmentation boundaries. The algorithm was tested on a difficult data set, a 1 hour audio segment of a film, with impressive results. It achieves an audio scene change detection accuracy of 97%.
Hari Sundaram, Shih-Fu Chang
ICASSP2
2000 Discovering Recurrent Visual Semantics in Consumer Photographs
abstract
We present techniques to semi-automatically discover recurrent visual semantics (RVS)-the repetitive appearance of visually similar elements such as objects and scenes-in consumer photographs. First, we introduce the detection of "bracketing" (very similar photographs) using an edge-correlation metric, which outperforms the color histogram. Then, we use color and novel composition features (based on automatic region segmentation) to perform scene-level clustering of images. We use a novel sequence-weighted technique, which uses the structure of standard film (only image sequence information), to perform hierarchical clustering. We show performance results of bracketing, explore clustering evaluation, and discuss STELLA, an interactive albuming and story telling application that uses these techniques to assist users in building digital albums. The STELLA system uses a new approach to album creation: instead of automatically creating albums, it provides an interactive environment that assists users in digital album creation.
Alejandro Jaimes, Ana B. Benitez, Shih-Fu Chang, Alexander C. Loui
ICIP3
2000 Experiments in Constructing Belief Networks for Image Classification Systems
abstract
We present procedures and experimental results in constructing belief networks for image classification systems based on probabilistic reasoning. In particular, we compare the performance of systems based on manually constructed and automatically constructed belief networks. The systems exploit existing image descriptions and also exploit interactions between multiple classifiers to improve classification performance. Performance evaluation results for the consumer photography domain are presented.
Seungyup Paek, Shih-Fu Chang
ICIP2
2000 Principles and applications of content-aware video communication
abstract
Most traditional video communication systems consider videos as low-level bit streams, ignoring the underlying visual content. Content-aware video communication is a new framework that explores the strong correlation between video content, resource (bit rate), and utility (quality). Such a framework facilitates new ways of quality modeling and resource allocation in multimedia communication. We demonstrate advantages of the content-aware approaches in two applications. First, content-aware models were developed for predicting video traffic for live video streams. The video traffic models were evaluated in a dynamic network resource allocation system. Our simulations have shown that, compared to existing techniques, significant reduction (55% to 70%) in required network resources can be achieved. Second, we have used the content-aware principle for automatic generation of utility function (subjective quality vs. bit rate) for live video. Our results indicate that high accuracy in estimating utility functions can be achieved. Such utility functions can be applied to optimal transcoding and media scaling in distributed network environments.
Shih-Fu Chang, Paul Bocheck
ISCAS1
2000 Determining computable scenes in films and their structures using audio-visual memory models
abstract
In this paper we present novel algorithms for computing scenes and within-scene structures in films. We begin by mapping insights from film-making rules and experimental results from the psychology of audition into a computational scene model. We define a computable scene to be a chunk of audio-visual data that exhibits long-term consistency with regard to three properties: (a) chromaticity (b) lighting (c) ambient sound. Central to the computational model is the notion of a causal, finite-memory viewer model. We segment the audio and video data separately. In each case we determine the degree of correlation of the most recent data in the memory with the past. The respective scene boundaries are determined using local minima and aligned using a nearest neighbor algorithm. We introduce a periodic analysis transform to automatically determine the structure within a scene. We then use statistical tests on the transform to determine the presence of a dialogue. The algorithms were tested on a difficult data set: five commercial films. We take the first hour of data from each of the five films. The best results: scene detection: 88% recall and 72% precision, dialogue detection: 91% recall and 100% precision.
Hari Sundaram, Shih-Fu Chang
ACM Multimedia2
2000 Error-resilient transcoding for video over wireless channels
abstract
We describe a method to maintain quality for video transported over wireless channels. The method is built on three fundamental blocks. First, we use a transcoder that injects spatial and temporal resilience into an encoded bitstream. The amount of resilience is tailored to the content of the video and the prevailing error conditions, as characterized by bit error rate. Second, we derive analytical models that characterize how corruption propagates in a video that is compressed using motion-compensated encoding and subjected to bit errors. Third, we use rate distortion theory to compute the optimal allocation of bit rate among spatial resilience, temporal resilience, and source rate. Furthermore, we use the analytical models to generate the resilience rate distortion functions that are used to compute the optimal resilience. The transcoder then injects this optimal resilience into the bitstream. Simulation results show that using a transcoder to optimally adjust the resilience improves video quality in the presence of errors while maintaining the same input bit rate.
Gustavo de los Reyes, Amy R. Reibman, Shih-Fu Chang, Justin C.-I. Chuang
IEEE J. Sel. Areas Commun.3
2000 Object-based multimedia content description schemes and applications for MPEG-7
abstract
In this paper, we describe description schemes (DSs) for image, video, multimedia, home media, and archive content proposed to the MPEG-7 standard. MPEG-7 aims to create a multimedia content description standard in order to facilitate various multimedia searching and filtering applications. During the design process, special care was taken to provide simple but powerful structures that represent generic multimedia data. We use the extensible markup language (XML) to illustrate and exemplify the proposed DSs because of its interoperability and flexibility advantages. The main components of the image, video, and multimedia description schemes are object, feature classification, object hierarchy, entity-relation graph, code downloading, multi-abstraction levels, and modality transcoding. The home media description instantiates the former DSs proposing the 6-W semantic features for objects, and 1-P physical and 6-W semantic object hierarchies. The archive description scheme aims to describe collections of multimedia documents, whereas the former DSs only aim at individual multimedia documents. In the archive description scheme, the content of an archive is represented using multiple hierarchies of clusters, which may be related by entity-relation graphs. The hierarchy is a specific case of entity-relation graph using a containment relation. We explicitly include the hierarchy structure in our DSs because it is a natural way of defining composite objects, a more efficient structure for retrieval, and the representation structure used in MPEG-4. We demonstrate the feasibility and the efficiency of our description schemes by presenting applications that already use the proposed structures or will greatly benefit from their use. These applications are the visual apprentice, the AMOS-search system, a multimedia broadcast news browser, a storytelling system, and an image meta-search engine, MetaSEEk.
Ana B. Benitez, Seungyup Paek, Shih-Fu Chang, Atul Puri, John R. Smith, Chung-Sheng Li, Lawrence D. Bergman, Charles N. Judice
Signal Process. Image Commun.3
2000 A practical methodology for guaranteeing quality of service for video-on-demand
abstract
A novel and simple approach for defining end-to-end quality of service (QoS) in video-on-demand (VoD) services is presented. Using this approach, we derive a schedulable region for a video server which guarantees end-to-end QoS, where a specific QoS required in the video client translates into a QoS specification for the video server. Our methodology is based on a generic model for VoD services, which is extendible to any VoD system. In this kind of system, both the network and the video server are potential sources of QoS degradation. Specifically, we examine the effect that impairments in the video server and video client have on the video quality perceived by the end user. The Columbia VoD testbed is presented as an example to validate the model through experimental results. Our model can be connected to network QoS admission control models to create a unified approach for admission control of incoming video requests in the video server and network.
Shih-Fu Chang, Alexandros Eleftheriadis, Dimitris Anastassiou, Stephen Jacobs, Javier Zamora
IEEE Trans. Circuits Syst. Video Technol.1
2000 Video-server retrieval scheduling and resource reservation for variable bit rate scalable video
abstract
State-of-the-art digital video compression produces bursty, variable bit rate video. The bursty nature of compressed video raises challenges in the design of video servers. In this paper, we first present a method for the efficient retrieval of bursty video data from the disk system to the memory of a digital video server. For a single video stream, the proposed retrieval schedule minimizes the buffer requirement for continuous retrieval, given that a fixed disk bandwidth is reserved for the entire duration of retrieval. Secondly, we present an optimal resource-reservation algorithm for multiple video streams based on the proposed retrieval schedule. The resource-reservation algorithm maximizes the number of bursty video streams that ran be supported by a video server given any disk bandwidth and memory resource. Thirdly, we present a progressive display scheme for scalable video that is based on the retrieval schedule and resource-reservation algorithm. Performance evaluations based on simulations using MPEG-2 trace data are presented. For a personal computer with four disks and a memory resource of 120 MB, our approach can support 50%-275% more video streams than previously proposed approaches, depending on the pre-fetch delay that users are willing to tolerate in interactive viewing of videos.
Seungyup Paek, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.2
1999 Region Feature Based Similarity Searching of Semantic Video Objects
abstract
New video representations based on semantic objects (e.g., MPEG-4) provide great potential for content-based video searching. In this paper we present an efficient query model for similarity searching of video objects based on localized region features and spatial-temporal structures. The query model is barred on an existing framework on region-level video query, but addresses several new issues like right spatio-temporal relationships effective performance of the proposed approach.
Di Zhong, Shih-Fu Chang
ICIP (2)2
1999 Multimedia access and retrieval: the state of the art and future directions (panel session)
abstract
Several years have passed since the research topic of content based multimedia retrieval emerged.We have witnessed the burgeoning research activities into a plenitude of new indexing, retrieval, and filtering tools for images, video, audio, music, graphics, and their combinations with text-based information.Exciting research opportunities arise when integrating knowledge from multiple disciplines, such as media content processing, database, information retrieval, and machine user interface.In the commercial domain, we have also witnessed several impressive efforts moving technologies into practical arenas.This panel includes experts from industry, research labs, and academia.The panel will assess the state of the art and articulate the important future directions in the general field of multimedia access and retrieval.
Shih-Fu Chang, Gwendal Auffret, Jonathan Foote, Chung-Sheng Li, Behzad Shahraray, Tanveer F. Syeda-Mahmood, HongJiang Zhang
ACM Multimedia (1)1
1999 GUEST EDITORS' INTRODUCTION: Content-Based Access of Image and Video Libraries
Alberto Del Bimbo, Vittorio Castelli, Shih-Fu Chang, Chung-Sheng Li
Comput. Vis. Image Underst.3
1999 Image Retrieval: Current Techniques, Promising Directions, and Open Issues
Yong Rui, Thomas S. Huang, Shih-Fu Chang
J. Vis. Commun. Image Represent.3
1999 Editorial: Content-Processing for Video Browsing, Retrieval, and Editing
Shih-Fu Chang, HongJiang Zhang
Multim. Syst.1
1999 Searching and Editing MPEG-Compressed Video in a Distributed Online Environment
Horace J. Meng, Di Zhong, Shih-Fu Chang
Multim. Syst.3
1999 Integrated Spatial and Feature Image Query
John R. Smith, Shih-Fu Chang
Multim. Syst.2
1999 Objective and subjective quality of service performance of video-on-demand in ATM-WAN
Javier Zamora, Dimitris Anastassiou, Shih-Fu Chang, Leonard W. Ulbricht
Signal Process. Image Commun.3
1999 Introduction to the special issue on object-based video coding and description
Fernando Pereira 0001, Shih-Fu Chang, Rob Koenen, Atul Puri, Olivier Avaro
IEEE Trans. Circuits Syst. Video Technol.2
1999 An integrated approach for content-based video object segmentation and retrieval
abstract
Object-based video data representations enable unprecedented functionalities of content access and manipulation. We present an integrated approach using region-based analysis for semantic video object segmentation and retrieval. We first present an active system that combines low-level region segmentation with user inputs for defining and tracking semantic video objects. The proposed technique is novel in using an integrated feature fusion framework for tracking and segmentation at both region and object levels. Experimental results and extensive performance evaluation show excellent results compared to existing systems. Building upon the segmentation framework, we then present a unique region-based query system for semantic video object. The model facilitates powerful object search, such as spatio-temporal similarity searching at multiple levels.
Di Zhong, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.2
1998 Digital image/video library and MPEG-7: standardization and research issues
abstract
Much research activity and interest has emerged in two closely related areas: the digital image/video library (DIVL) and MPEG-7. We review the critical research issues in DIVL from a signal processing viewpoint, the objectives and scope of MPEG-7, and the relationships between these two.
Yong Rui, Thomas S. Huang, Shih-Fu Chang
ICASSP3
1998 Semantic Visual Templates: Linking Visual Features to Semantics
Shih-Fu Chang, William Chen 0001, Hari Sundaram
ICIP (3)1
1998 Embedding Visible Video Watermarks in the Compressed Domain
abstract
Digital visible or invisible watermarks are increasingly in demand for protecting or verifying the original image or video ownership. We propose a novel compressed-domain approach to embedding visible watermarks in MPEG-1 and MPEG-2 video streams. Our algorithms operate on the DCT coefficients which are obtained with minimal parsing of input video. The embedded watermarks adapt to the local video features such as brightness and complexity to achieve consistent perceptual visibility. The embedded watermarks are robust against attempts of removal since clear artifacts remain after the possible attacks.
Jianhao Meng, Shih-Fu Chang
ICIP (1)2
1998 Video Transcoding for Resilience in Wireless Channels
abstract
We describe a method to maintain an acceptable quality for video transported over wireless networks under time-varying conditions. We use a transcoder to modify the resilience of the encoded bitstream by using source coding techniques to provide the appropriate level of resilience for the prevailing channel conditions. We develop a statistical model for image loss versus resilience and wireless conditions. Simulation results indicate that using a transcoder to adjust the resilience can improve video quality when errors occur without significantly sacrificing quality when there are no errors. Also, simulation results compare favorably to the analytical model.
Gustavo de los Reyes, Amy R. Reibman, Justin C.-I. Chuang, Shih-Fu Chang
ICIP (1)4
1998 AMOS: An Active System for MPEG-4 Video Object Segmentation
abstract
Object segmentation and tracking is a fundamental step for many digital video applications. In this paper, we present an active system (AMOS) which combines low level automatic region segmentation with an active method for defining and tracking high-level semantic video objects. The system contains two stages: an initial object segmentation stage where user input in the starting frame is used to create a semantic object; and an object tracking stage where underlying regions of the semantic object are tracked and grouped through successive frames. Experiments with different types of videos show very good performance.
Di Zhong, Shih-Fu Chang
ICIP (2)2
1998 VideoQ: a fully automated video retrieval system using motion sketches
abstract
The rapidity with which digital information, particularly video, is being generated, has necessitated the development of tools for efficient search of these media. Content based visual queries have been primarily focused on still image retrieval. In this paper we propose a novel interactive system on the Web, based on the visual paradigm, with spatio-temporal attributes playing a key role in video retrieval. The resulting system VideoQ, is the first on-line video search engine supporting automatic object based indexing and spatio-temporal queries.
Shih-Fu Chang, William Chen 0001, Hari Sundaram
WACV1
1998 Generating Multimedia Briefings: Coordinating Language and Illustration
abstract
Communication can be more effective when several media (such as text, speech, or graphics) are integrated and coordinated to present information. This changes the nature of media-specific generation (e.g., language or graphics generation), which must take into account the multimedia context in which it occurs. This paper presents work on coordinating and integrating speech, text, static and animated three-dimensional graphics, and stored images, as part of several systems we have developed at Columbia University. A particular focus of our work has been on the generation of presentations that brief a user on information of interest
Kathy McKeown, Steven K. Feiner, Mukesh Dalal, Shih-Fu Chang
Artif. Intell.4
1998 Special Issue on Image Technology for World Wide Web Applications: Guest Editors' Comments
Shih-Fu Chang, Ping Wah Wong, HongJiang Zhang
J. Vis. Commun. Image Represent.1
1998 Next-generation content representation, creation, and searching for new-media applications in education
abstract
Content creation, editing, and searching are extremely time-consuming tasks that often require substantial training and experience, especially when high-quality audio and video are involved. New media represents a new paradigm for multimedia information representation and processing, in which the emphasis is placed on the actual content. It thus brings the tasks of content creation and searching much closer to actual users and enables them to be active producers of audio-visual information rather than passive recipients. We discuss the state of the art and present next-generation techniques for content representation, searching, creation and editing. We discuss our experiences in developing a Web-based distributed compressed video editing and searching system (WebClip), a media-representation language (Flavor) and an object-based video authoring system (Zest) based on it, and a large image/video search engine for the World Wide Web (WebSEEk). We also present a case study of new media applications based on specific planned multimedia education experiments with the above systems in several K-12 schools in Manhattan, NY.
Shih-Fu Chang, Alexandros Eleftheriadis, Robert McClintock
Proc. IEEE1
1998 Effective algorithms for video transmission over wireless channels
abstract
New schemes for video transmission over wireless channels are described. Content based approaches for video segmentation and associated resource allocation are proposed. We argue that for transport over wireless channels, different video content requires different form of resource allocation. Joint source/channel coding techniques (particularly content/data dependent FEC/ARQ schemes) are used to do adaptive resource allocation after segmentation. FEC based schemes are used to provide class dependent error robustness and a modified ARQ technique is used to provide constrained delay and loss. The approach is compatible with video coding standards such as H.261 and H.263. We use frametype (extent of intra coding), scene changes, and motion based procedures to provide finer level of control for data segmentation in addition to standard headers and data-type (motion vectors, low and high frequency DCTs) based segmentation. An experimental simulation platform is used to test the objective (SNR) and subjective effectiveness of proposed algorithms. The FEC schemes improve both objective and subjective video quality significantly. An experimental study on the applicability of selective repeat ARQ for one way real time video applications is also presented. We study constraints of using ARQ under display constraints, limited buffering requirements and small initial startups. The proposed ARQ schemes greatly reduce the packet errors, when used along with optimal decoder buffer control and source interleaving. The common theme integrating the study of FEC and ARQ algorithms is the content based resource allocation for wireless video transport.
Pankaj Batra, Shih-Fu Chang
Signal Process. Image Commun.2
1998 A fully automated content-based video search engine supporting spatiotemporal queries
abstract
The rapidity with which digital information, particularly video, is being generated has necessitated the development of tools for efficient search of these media. Content-based visual queries have been primarily focused on still image retrieval. In this paper, we propose a novel, interactive system on the Web, based on the visual paradigm, with spatiotemporal attributes playing a key role in video retrieval. We have developed innovative algorithms for automated video object segmentation and tracking, and use real-time video editing techniques while responding to user queries. The resulting system, called VideoQ , is the first on-line video search engine supporting automatic object-based indexing and spatiotemporal queries. The system performs well, with the user being able to retrieve complex video clips such as those of skiers and baseball players with ease.
Shih-Fu Chang, William Chen 0001, Horace J. Meng, Hari Sundaram, Di Zhong
IEEE Trans. Circuits Syst. Video Technol.1
1997 Exploring Image Functionalities in WWW Applications- Development of Image/Video Search and Editing Engines
abstract
Image technologies provide critical components for achieving various Web-based multimedia applications. We identify key technical challenges in today's image/video applications on the WWW. We present two prototype systems, a video search engine (WebSEEk) and a compressed video editor (WebClip), to demonstrate technical challenges and viable solutions. We also discuss major issues in developing next-generation, scalable solutions for large-scale, distributed on-line environments.
Shih-Fu Chang, John R. Smith, Horace J. Meng
ICIP (3)1
1997 Joint adaptive space and frequency basis selection
abstract
We develop a new method for building a representation of an image from a library of basis elements that is facilitated by a joint adaptive space and frequency (JASF) graph. The JASF graph combines partitionable frequency expansion and spatial segmentation of the image, symmetrically. We demonstrate by using a rate-distortion framework for basis selection that the JASF graph improves compression performance over recent wavelet packet and double-tree methods by offering exponentially more bases in which to represent the images.
John R. Smith, Shih-Fu Chang
ICIP (3)2
1997 Spatio-Temporal Video Search Using the Object-Based Video Representation
abstract
Object-based video representation provides great promises for new search and editing functionalities. Feature regions in video sequences are automatically segmented, tracked, and grouped to form the basis for content-based video search and higher levels of abstraction. We present a new system for video object segmentation and tracking using feature fusion and region grouping. We also present efficient techniques for spatio-temporal video query based on the automatically segmented video objects.
Di Zhong, Shih-Fu Chang
ICIP (1)2
1997 VideoQ: An Automated Content Based Video Search System Using Visual Cues
abstract
The rapidity with which digitat information, particularly video, is being generated, has necessitated the development of tools for efficient search of these media.Content based visual queries have been primarily focussed on still image retrieval.In this papel; we propose a novel, real-time, interactive system on the Web, based on the visual paradigm, with spatio-temporal attributesplaying a key role in video retrieval.We have developed algorithms for automated video object segmentation and tracking and use real-time video editing techniques while responding to user queries.The resulting system pe$orms well, with the user being able to retrieve complex video clips such as those of skiers, baseball players, with ease.Penlli~iollto m&e digitnlhrd copies of ail or pa11 ofthis iilfiterinl for personal or clmsroom use is granted without fee provided tht (112 copiare not made or distributed for profit or commercial ndvrmtnge.the COPYridtt notice, tile title ofthe publicntion and its date appear.and notice is given tl]ntcopyrigllt is by permission ofthe ACM, hC.TO Copy othWh& to republisl), to post on servers or to redistribule to lists, requires specific peniiission nnd/or fee.ACM Multimedia 97 ,~en/f/e I~TfJ.~/lingfO17iJX?I
Shih-Fu Chang, William Chen 0001, Horace J. Meng, Hari Sundaram, Di Zhong
ACM Multimedia1
1997 A distributed system for editing and browsing compressed video over the network
abstract
We present a new framework for distributed video editing and browsing over the network. This framework uses a distributed client-server model including a server engine for content analysis/editing, and clients for interactive controls of video browsing/editing. It utilizes several unique features, including compressed-domain video manipulation, multi-resolution video access, content based video browsing/retrieval, and a distributed network architecture. We have developed a complete functional prototype, CVEPS, which includes a compressed domain video analysis and editing engine, and a JAVA based user interface. It has been incorporated into a World Wide Web application, WebClip, for editing compressed video over the Web.
Horace J. Meng, Di Zhong, Shih-Fu Chang
MMSP3
1997 WebClip: a WWW video editing/browsing system
abstract
WebClip is a complete working prototype for editing/browsing MPEG-1 and MPEG-2 compressed video distributively over the World Wide Web. It uses a general system architecture to store, retrieve, and edit MPEG-1 or MPEG-2 compressed video over the network. It uses an innovative distributed network support architecture. It also uses our unique CVEPS (Compressed video editing, parsing, and search) technologies.
Horace J. Meng, Di Zhong, Shih-Fu Chang
MMSP3
1997 SaFe: a general framework for integrated spatial and feature image search
abstract
We present a system for querying for images by the spatial and feature attributes of regions. The system enables the user to find the images that contain an arrangement of regions similar to that diagrammed in a query image. We propose a general framework which allows for different types of features (e.g., color, texture, shape, motion) to be integrated with spatial information in the query process. We demonstrate that integrated spatial and feature querying improves image search capabilities over previous content-based image retrieval methods.
John R. Smith, Shih-Fu Chang
MMSP2
1997 Enhancing image search engines in visual information environments
abstract
The recent diverse environments and applications for image searching (e.g. Web image search engines) provide an enormous resource of information beyond the image pixels which can be used to improve the image search process. We explore several directions of enhancements that integrate the visual information with other information related to the images in the analysis and query processes. We demonstrate that these methods improve image search functionalities over non-integrated content-based methods.
John R. Smith, Shih-Fu Chang
MMSP2
1997 Guest Editorial
Shih-Fu Chang, Dimitris Anastassiou, Alexandros Eleftheriadis, John V. Pavlik
Multim. Tools Appl.1
1997 Columbia's VoD and Multimedia Research Testbed with Heterogeneous Network Support
Shih-Fu Chang, Alexandros Eleftheriadis, Dimitris Anastassiou, Stephen Jacobs, Hari Kalva, Javier Zamora
Multim. Tools Appl.1
1997 DAVIC and Interoperability Experiments
Hari Kalva, Shih-Fu Chang, Alexandros Eleftheriadis
Multim. Tools Appl.2
1997 Hybrid object-based/block-based coding in video compression at very low bit-rate
Chung-Tao Chu, Dimitris Anastassiou, Shih-Fu Chang
Signal Process. Image Commun.3
1997 A highly efficient system for automatic face region detection in MPEG video
abstract
Human faces provide a useful cue in indexing video content. We present a highly efficient system that can rapidly detect human face regions in MPEG video sequences. The underlying algorithm takes the inverse quantized discrete cosine transform (DCT) coefficients of MPEG video as the input, and outputs the locations of the detected face regions. The algorithm consists of three stages, where chrominance, shape, and frequency information are used, respectively. By detecting faces directly in the compressed domain, there is no need to carry out the inverse DCT transform, so that the algorithm can run faster than the real time. In our experiments, the algorithm detected 85-92% of the faces in three test sets, including both intraframe and interframe coded image frames from news video. The average run time ranges from 13-33 ms per frame. The algorithm can be applied to JPEG unconstrained images or motion JPEG video as well.
Hualu Wang, Shih-Fu Chang
IEEE Trans. Circuits Syst. Video Technol.2
1996 Hybrid Block-Based/Segment-Based Video Compression at very Low Bitrate
abstract
An algorithmic architecture of video compression at very low bit rate is presented. This architecture incorporates block-based as well as segment-based coding algorithms and manages them in a cooperative way in which different approaches are applied to different image frames iteratively: I and P frames are coded by block-based algorithm and O frames are coded by segment-based algorithm. An extra degree of freedom in video coding is introduced in order to achieve higher compression ratios while maintaining good visual quality. Every other frame in the image sequence is coded using block-based coding (ITU-T H.263) with motion (block-matching, 2D translations) and color (DCT) coefficients. These I and P frames are allocated with more bits, generating images with higher spatial accuracy, and providing the coder with generic coding capability. Between every two P frames, the O frame is coded using segment-based algorithm providing the motion (rigid 2D source models, 3D motions) and shape (polygon or spline approximation) information. Since only motion prediction is provided, pixel values are copied or interpolated from neighboring P frames. With good segmentation and affine transform, the prediction can preserve the natural shapes of the objects and give natural motion information which is sufficient for very low bitrate coding. These frames are allocated with less bits, generating images with less temporal accuracy, but providing the coder with higher compression efficiency.
Chung-Tao Chu, Dimitris Anastassiou, Shih-Fu Chang
Data Compression Conference3
1996 Automated binary texture feature sets for image retrieval
abstract
Digital image and video libraries require new algorithms for the automated extraction and indexing of salient image features. Texture features provide one important cue for the visual perception and discrimination of image content. We propose a new approach for automated content extraction that allows for efficient database searching using texture features. The algorithm automatically extracts texture regions from image spatial-frequency data which are represented by binary texture feature vectors. We demonstrate that the binary texture features provide excellent performance in image query response time while providing highly effective texture discriminability, accuracy in spatial localization and capability for extraction from compressed data representations. We present the binary texture feature extraction and indexing technique and examine searching by texture on a database of 500 images.
John R. Smith, Shih-Fu Chang
ICASSP2
1996 A content based video traffic model using camera operations
abstract
We present our work on content based video (CBV) traffic modeling of variable bit rate (VBR) sources. The CBV approach differs from previous works in that it is not based only on matching of various statistics of the original source, but rather on modeling and mapping its visual content into the corresponding bit rate. We show that the CBV model is fully compatible with current and future compression algorithms including those of very low bit rate video coding. We introduce the separation principle between the visual content and encoder dependent bit rate mapping. We construct and verify two experimental CBV models for basic camera operations. The results obtained show that the CBV model can closely match the various statistics of the MPEG-2 VBR stream.
Paul Bocheck, Shih-Fu Chang
ICIP (2)2
1996 A robust content based digital signature for image authentication
abstract
A methodology for designing content based digital signatures which can be used to authenticate images is presented. A continuous measure of authenticity is presented which forms the basis of this methodology. Using this methodology signature systems can be designed which allow certain types of image modification (e.g. lossy compression) but which prevent other types of manipulation. Some experience with content based signatures is also presented. The idea of signature based authentication is extended to video, and a system to generate signatures for video sequences is presented. This signature also allows smaller segments of the secured video to be verified as unmanipulated.
Marc Schneider, Shih-Fu Chang
ICIP (3)2
1996 Local color and texture extraction and spatial query
abstract
In this paper we present a unified system for the extraction, representation and query of spatially localized color and texture regions. The system utilizes a back-projection of binary feature sets to identify and extract prominent regions. The binary feature sets provide an effective and easily indexable representation of color and texture. We also provide a mechanism for integrating features by combining the binary color and texture feature sets. This enables the extraction and representation of joint color and texture regions. Since all extracted regions are spatially localized, in image database queries the user can specify the locations and spatial boundaries of regions. We present the unified color and texture back-projection method and describe its implementation in the VisualSEEk content-based image retrieval system.
John R. Smith, Shih-Fu Chang
ICIP (3)2
1996 CVEPS - A Compressed Video Editing and Parsing System
abstract
Article CVEPS - a compressed video editing and parsing system Share on Authors: Jianhao Meng Department of Electrical Engineering & Center for Image Technology for New Media, Columbia University, New York, NY Department of Electrical Engineering & Center for Image Technology for New Media, Columbia University, New York, NYView Profile , Shih-Fu Chang Department of Electrical Engineering & Center for Image Technology for New Media, Columbia University, New York, NY Department of Electrical Engineering & Center for Image Technology for New Media, Columbia University, New York, NYView Profile Authors Info & Claims MULTIMEDIA '96: Proceedings of the fourth ACM international conference on MultimediaFebruary 1997 Pages 43–53https://doi.org/10.1145/244130.244145Online:01 February 1997Publication History 61citation764DownloadsMetricsTotal Citations61Total Downloads764Last 12 Months6Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jianhao Meng, Shih-Fu Chang
ACM Multimedia2
1996 VisualSEEk: A Fully Automated Content-Based Image Query System
abstract
Article Free Access Share on VisualSEEk: a fully automated content-based image query system Authors: John R. Smith Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y. Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y.View Profile , Shih-Fu Chang Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y. Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y.View Profile Authors Info & Claims MULTIMEDIA '96: Proceedings of the fourth ACM international conference on MultimediaFebruary 1997 Pages 87–98https://doi.org/10.1145/244130.244151Online:01 February 1997Publication History 1,080citation3,873DownloadsMetricsTotal Citations1,080Total Downloads3,873Last 12 Months90Last 6 weeks11 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
John R. Smith, Shih-Fu Chang
ACM Multimedia2
1996 Development of Columbia's video on demand testbed
abstract
This paper describes our progress in developing an advanced video-on-demand (VOD) testbed, which will accommodate various multimedia research and applications such as Electronic News on Demand, Columbia's Video Course Network, and Digital Libraries. Two different prototypes have been completed. The first generation of the testbed was based on a constant-bit-rate (CBR) video server utilizing Ethernet delivery. Contents were encoded and stored as MPEG-2 audio/video elementary streams. Software encoders/decoders were used in content generation and playback. The second generation of the testbed was enhanced with the capability of transmitting true MPEG-2 transport streams over the campus ATM network as well as the wide-area NYNET ATM network. A real-time video pump and a distributed application control protocol (MPEG-2's DSM-CC) have been incorporated. Hardware decoders and set-tops are being incorporated to test wide-area video interoperability. Our VOD testbed also provides an advanced platform for implementing proof-of-concept prototypes of related research. Our current research focus covers video transmission with heterogeneous quality-of-service (QoS) provision, video storage architecture design, content-based video indexing and browsing, multi-resolution (MR) video coding, efficient manipulation of compressed video, and advanced user interfaces. An important aim is to enhance interoperability. Accommodation of practical multimedia applications and interoperability testing with external VOD systems are currently being undertaken.
Shih-Fu Chang, Alexandros Eleftheriadis, Dimitris Anastassiou
Signal Process. Image Commun.1
1995 Frequency and spatially adaptive wavelet packets
abstract
We consider a method for image compression based on frequency and spatially adaptive wavelet packets. We present a new fast directed acyclic graph (DAG) structured decomposition, with both spatial segmentation and orthogonal frequency branching from each node. Whereas traditional wavelet packet decomposition adapts to a global frequency distribution, this technique finds the best joint spatial segmentation and local frequency basis. The algorithm is derived from the fast double tree algorithm proposed by Herley, et. al. (see IEE Transactions on Signal Processing, December 1993), for 1-D signals, with an extension to 2-D and modification to include spatial segmentation of frequency nodes. By collecting redundant nodes in this full adaptive tree, we have derived a directed acyclic graph (DAG) structure which contains the same number of nodes as the double tree, but includes new connections between nodes. We present the adaptive wavelet packet DAG algorithm and examine image compression performance on test images.
John R. Smith, Shih-Fu Chang
ICASSP2
1995 Compressed-domain techniques for image/video indexing and manipulation
abstract
As massive amount of visual materials are captured and stored in visual information systems, effective and efficient image indexing and manipulation techniques are required. Most visual materials in visual information systems are stored in some compressed forms. Therefore, it is desirable to explore image technologies for feature extraction and image manipulation in the compressed domain. In other words, image feature extraction and manipulation are performed on compressed images/video without decoding, or with minimal decoding only. Although the compressed-domain approach imposes many constraints, it provides great potential for reducing computational complexity, because of reduction of the amount of data after compression. This paper provides an overview of our research in this area. Specifically, it describes the results and the future directions of our work on compressed-domain texture feature extraction, image matching, image manipulation, and video indexing.
Shih-Fu Chang
ICIP1
1995 Single color extraction and image query
abstract
We propose a method for automatic color extraction and indexing to support color queries of image and video databases. This approach identifies the regions within images that contain colors from predetermined color sets. By searching over a large number of color sets, a color index for the database is created in a fashion similar to that for file inversion. This allows very fast indexing of the image collection by the color contents of the images. Furthermore, information about the identified regions, such as the color set, size, and location, enables a rich variety of queries that specify both color content and spatial relationships of regions. We present the single color extraction and indexing method and contrast it to other color approaches. We examine single and multiple color extraction and image query on a database of 3000 color images.
John R. Smith, Shih-Fu Chang
ICIP (3)2
1995 Scalable MPEG2 Video Servers with Heterogeneous QoS on Parallel Disk Arrays
Seungyup Paek, Paul Bocheck, Shih-Fu Chang
NOSSDAV3
1995 Manipulation and Compositing of MC-DCT Compressed Video
abstract
Many advanced video applications require manipulations of compressed video signals. Popular video manipulation functions include overlap (opaque or semitransparent), translation, scaling, linear filtering, rotation, and pixel multiplication. We propose algorithms to manipulate compressed video in the compressed domain. Specifically, we focus on compression algorithms using the discrete cosine transform (DCT) with or without motion compensation (MC). Such compression systems include JPEG, motion JPEG, MPEG, and H.261. We derive a complete set of algorithms for all aforementioned manipulation functions in the transform domain, in which video signals are represented by quantized transform coefficients. Due to a much lower data rate and the elimination of decompression/compression conversion, the transform-domain approach has great potential in reducing the computational complexity. The actual computational speedup depends on the specific manipulation functions and the compression characteristics of the input video, such as the compression rate and the nonzero motion vector percentage. The proposed techniques can be applied to general orthogonal transforms, such as the discrete trigonometric transform. For compression systems incorporating MC (such as MPEG), we propose a new decoding algorithm to reconstruct the video in the transform domain and then perform the desired manipulations in the transform domain. The same technique can be applied to efficient video transcoding (e.g., from MPEG to JPEG) with minimal decoding.>
Shih-Fu Chang, David G. Messerschmitt
IEEE J. Sel. Areas Commun.1
1994 Transform Features for Texture Classification and Discrimination in Large Image Databases
abstract
Proposes a method for classification and discrimination of textures based on the energies of image subbands. The authors show that with this relatively simple feature set, effective texture discrimination can be achieved. In the paper, subband-energy feature sets extracted from the following typical image decompositions are compared: wavelet subband, uniform subband, discrete cosine transform (DCT), and spatial partitioning. The authors report that over 90% correct classification was attained using the feature set in classifying the full Brodatz [1965] collection of 112 textures. Furthermore, the subband energy-based feature set can be readily applied to a system for indexing images by texture content in image databases, since the features can be extracted directly from spatial-frequency decomposed image data. The authors also show that to construct a suitable space for discrimination, Fisher discrimination analysis (Dillon and Goldstein, 1984) can be used to compact the original features into a set of uncorrelated linear discriminant functions. This procedure makes it easier to perform texture-based searches in a database by reducing the dimensionality of the discriminant space. The authors also examine the effects of varying training class size, the number of training classes, the dimension of the discriminant space and number of energy measures used for classification. The authors hope that the performance for texture discrimination of these simple energy-based features will allow images in a database to be efficiently and effectively indexed by contents of their textured regions.>
John R. Smith, Shih-Fu Chang
ICIP (3)2
1994 Error Accumulation of Repetitive Image Coding
abstract
Repetitive image coding is encountered in applications such as video transcoding and image/video servers. This paper studies the effect of error accumulation in repetitive image coding. We describe the conditions under which zero-error-accumulation property holds. We study the impact of intermediate operations (like linear filtering, sealing, and shifting) on image quality in repetitive coding. Experiments with DCT, MC, and MPEG coding algorithms are reported.>
Shih-Fu Chang, Alexandros Eleftheriadis
ISCAS1
1994 Quad-Tree Segmentation for Texture-Based Image Query
abstract
In this paper we propose a technique for segmenting images by texture content with application to indexing images in a large image database. Using quad-tree decomposition, texture features are extracted from spatial blocks at a hierarchy of scales in each image. The quad-tree is grown by iteratively testing conditions for splitting parent blocks based on texture content of children blocks. While this approach does not achieve smooth identification of texture region borders, homogeneous blocks of texture are extracted which can be used in a database index. Furthermore, this technique performs the segmentation directly using image spatial-frequency data. In the segmentation reported here, texture features are extracted from the wavelet representation of the image. This method however, can use other subband decompositions including Discrete Cosine Transform (DCT), which has been adopted by the JPEG standard for image coding. This makes our segmentation method extremely applicable to databases containing compressed image data. We show application of the texture segmentation towards providing a new method for searching for images in large image databases using “Query-by-texture.”
Jonathan M. Smith, Shih-Fu Chang
ACM Multimedia2
1993 A new approach to decoding and compositing motion-compensated DCT-based images
Shih-Fu Chang, David G. Messerschmitt
ICASSP (5)1
1993 Transform Coding of Arbitrarily-Shaped Image Segments
abstract
Article Transform coding of arbitrarily-shaped image segments Share on Authors: Shih Fu Chang View Profile , David G. Messerschmitt View Profile Authors Info & Claims MULTIMEDIA '93: Proceedings of the first ACM international conference on MultimediaSeptember 1993 Pages 83–90https://doi.org/10.1145/166266.166275Published:01 September 1993 23citation664DownloadsMetricsTotal Citations23Total Downloads664Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Shih-Fu Chang, David G. Messerschmitt
ACM Multimedia1
1992 Designing high-throughput VLC decoder. I. Concurrent VLSI architectures
abstract
Two classes of architectures-the tree-based and the PLA-based architectures-have been discussed in the literature for the variable length code (VLC) decoder. Pipelined or parallel architectures in these two classes are proposed for high-speed implementation. The pipelined tree-based architectures have the advantages of fully pipelined design, short clock cycle, and partial programmability. They are suitable for concurrent decoding of multiple independent bit streams. The PLA-based architectures have greater flexibility and can take advantages of some high-level optimization techniques. The input/output rate can be fixed or variable to meet the application requirements. As an experiment, the authors have constructed a VLC based on a popular video compression system and compared the architectures. A layout of the major parts and a simulation of the critical path of the pipelined constant-input-rate PLA-based architecture using a high-level synthesis approach estimates that a decoding throughput of 200 Mb/s with a single chip is achievable with CMOS 2.0 mu m technology.>
Shih-Fu Chang, David G. Messerschmitt
IEEE Trans. Circuits Syst. Video Technol.1
1991 VLSI Designs for High-Speed Huffman Decoder
abstract
Many video compression systems require a high-speed implementation of the huffman decoder. The recursive iteration of the decoding process limits the achievable decoding throughput with a given IC technology. Two classes of VLSI architectures for high speed implementation are designed: the tree-based architectures and the programmable logic array (PLA)-based architectures. A variable-length-code based on a popular video compression system is constructed and the pros and cons of each architecture are compared. The major parts of a pipelined constant-input-rate PLA-based architecture using a high-level synthesis approach are simulated. It is claimed that the decoding throughput of 200 Mb/s is achievable with CMOS 2.0 mu m technology.>
Shih-Fu Chang, David G. Messerschmitt
ICCD1