VLDB 2026 Research / reviewers in the wild / expert
Alex Hauptmann 0001
dblp:h/AlexanderGHauptmann · also Alexander G. Hauptmann, Alexander Hauptmann 0001
· DBLP profile ↗
290ranked-venue papers
15as first author
41since 2021 · last 2025
0000-0003-2123-0684ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 201 · 9 first-author · 23 since 2021Artificial intelligence and machine learning · 135 · 5 first-author · 32 since 2021Databases, data management, data science and information retrieval · 37 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorHuman-computer interaction and ubiquitous computing · 7 · 4 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Visual-Semantic Subspace RepresentationsabstractLearning image representations that capture rich semantic relationships remains a significant challenge. Existing approaches are either contrastive, lacking robust theoretical guarantees, or struggle to effectively represent the partial orders inherent to structured visual-semantic data. In this paper, we introduce a nuclear norm-based loss function, grounded in the same information theoretic principles that have proved effective in self-supervised learning. We present a theoretical characterization of this loss, demonstrating that, in addition to promoting class orthogonality, it encodes the spectral geometry of the data within a subspace lattice. This geometric representation allows us to associate logical propositions with subspaces, ensuring that our learned representations adhere to a predefined symbolic structure. Gabriel Moreira, Manuel Marques, João Paulo Costeira, Alex Hauptmann 0001 |
AISTATS | 4 |
| 2025 | Emphasizing Discriminative Features for Dataset Distillation in Complex ScenariosabstractDataset distillation has demonstrated strong performance on simple datasets like CIFAR, MNIST, and TinyImageNet but struggles to achieve similar results in more complex scenarios. In this paper, we propose EDF (emphasizes the discriminative features), a dataset distillation method that enhances key discriminative regions in synthetic images using Grad-CAM activation maps. Our approach is inspired by a key observation: in simple datasets, high-activation areas typically occupy most of the image, whereas in complex scenarios, the size of these areas is much smaller. Unlike previous methods that treat all pixels equally when synthesizing images, EDF uses Grad-CAM activation maps to enhance high-activation areas. From a supervision perspective, we downplay supervision signals produced by lower trajectory-matching losses, as they contain common patterns. Additionally, to help the DD community better explore complex scenarios, we build the Complex Dataset Distillation (Comp-DD) benchmark by meticulously selecting sixteen subsets, eight easy and eight hard, from ImageNet-1K. In particular, EDF consistently outperforms SOTA results in complex scenarios, such as ImageNet-1K subsets. Hopefully, more researchers will be inspired and encouraged to improve the practicality and efficacy of DD. Our code and benchmark have been made public at NUS-HPC-AI-Lab/EDF. Kai Wang 0036, Zhi-Qi Cheng, Samir Khaki, Ahmad Sajedi, Ramakrishna Vedantam, Konstantinos N. Plataniotis, Alex Hauptmann 0001, Yang You 0001 |
CVPR | 8 |
| 2025 | UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal PromptsabstractEmotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module(EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. To facilitate reproducibility and further research, we have made our code publicly available at https://github.com/KTTRCDL/UMETTS. Zhi-Qi Cheng, Jun-Yan He, Junyao Chen, Xiaomao Fan, Xiaojiang Peng, Alex Hauptmann 0001 |
ICASSP | 7 |
| 2025 | MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt SynthesisabstractMetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the creation of customizable WordArt, ranging from semantic enhancements to intricate textural elements. A central feedback mechanism leverages insights from both multimodal models and user evaluations, enabling iterative refinement of design parameters. Through this iterative process, MetaDesigner dynamically adjusts hyperparameters to align with user-defined stylistic and thematic preferences, consistently delivering WordArt that excels in visual quality and contextual resonance. Empirical evaluations underscore the system's versatility and effectiveness across diverse WordArt applications, yielding outputs that are both aesthetically compelling and context-sensitive. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Jingdong Sun, Qi He 0007, Wangmeng Xiang, Jin-Peng Lan, Xianhui Lin, Kang Zhu, Bin Luo 0008, Yifeng Geng, Xuansong Xie, Alex Hauptmann 0001 |
ICLR | 14 |
| 2025 | Direct Preference Optimization of Video Large Multimodal Models from Language Model RewardabstractRuohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G Hauptmann, Yonatan Bisk, Yiming Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alex Hauptmann 0001, Yonatan Bisk, Yiming Yang 0002 |
NAACL (Long Papers) | 9 |
| 2024 | Transitive Consistency Constrained Learning for Entity-to-Entity Stance DetectionabstractEntity-to-entity stance detection identifies the stance between a pair of entities with a directed link that indicates the source, target and polarity.It is a streamlined task without the complex dependency structure for structural sentiment analysis, while it is more informative compared to most previous work assuming that the source is the author.Previous work performs entity-to-entity stance detection training on individual entity pairs.However, stances between inter-connected entity pairs may be correlated.In this paper, we propose transitive consistency constrained learning, which first finds connected entity pairs and their stances, and adds an additional objective to enforce the transitive consistency.We explore consistency training on both classification-based and generation-based models and conduct experiments to compare consistency training with previous work and large language models with in-context learning.Experimental results illustrate that the inter-correlation of stances in political news can be used to improve the entityto-entity stance detection model, while overly strict consistency enforcement may have a negative impact.In addition, we find that large language models struggle with predicting link direction and neutral labels in this task. 1 Haoyang Wen, Eduard H. Hovy, Alex Hauptmann 0001 |
ACL (1) | 3 |
| 2024 | Text Motion Translator: A Bi-directional Model for Enhanced 3D Human Motion Generation from Open-Vocabulary Descriptions
Yijun Qian, Jack Urbanek, Alex Hauptmann 0001, Jungdam Won |
ECCV (63) | 3 |
| 2024 | Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
Hao Zhou 0014, Pengfei Xing, Long Zhao 0003, Junwei Liang 0001, Alex Hauptmann 0001, Ting Liu 0005, Andrew C. Gallagher |
ECCV (29) | 7 |
| 2024 | The Seven Faces of Stress: Understanding Facial Activity Patterns During Cognitive StressabstractStress has been recognized as one of the main contributors to mental health problems, as well as cardiovascular diseases. To reduce the risk of severe diseases, early detection of stress is needed. One of the recent methods studied to detect stress is through facial expression analysis from videos. Although computer vision techniques combined with deep learning have been shown to detect stressful faces, there is a lack of work attempting to define how stressful faces look. One of the main challenges is that the expression of stress is person-dependent and one individual can show stress in various ways. In this work, we present a semi-automatic method that allows to distill from a large quantity of data facial activity patterns that are recognized to show stress. We are the first to combine quantitative and qualitative methods on data from 115 subjects to identify and propose seven facial activity patterns during stress. We support this proposal by analyzing the relationship of the different stress facial expressions with the basic emotions and show how individual components of anger, fear, surprise, and sadness co-occur during our defined stress facial activity patterns. Carla Viegas, Roy A. Maxion, Alex Hauptmann 0001, João Magalhães |
FG | 3 |
| 2024 | PhISANet: Phonetically Informed Speech Animation NetworkabstractRealistic animation is crucial for immersive and seamless human-avatar interactions as digital avatars become more prevalent. This work presents PhISANet, an encoder-decoder model that realistically animates the face and tongue solely from speech. PhISANet leverages neural audio representations trained on vast amounts of speech to map the speech signal into animation parameters that control the lower face and tongue of realistic 3D models. By integrating a novel multi-task learning strategy during the training phase, PhISANet reincorporates the phonetic information from the input speech, improving articulation in the generated animations. A thorough quantitative and qualitative study validates this improvement, and it determines that WavLM and Whisper features are ideal for training a generalizable speech-animation model regardless of gender, age, and language. Sarah L. Taylor, Carsten Stoll, Gareth Edwards, Alex Hauptmann 0001, Shinji Watanabe 0001, Iain A. Matthews |
ICASSP | 5 |
| 2024 | Language Model Beats Diffusion - Tokenizer is key to visual generationabstractWhile Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks. Lijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng 0003, Agrim Gupta, Xiuye Gu, Alex Hauptmann 0001, Boqing Gong, Ming-Hsuan Yang 0001, Irfan A. Essa, David A. Ross, Lu Jiang 0004 |
ICLR | 10 |
| 2024 | VICAN: Very Efficient Calibration Algorithm for Large Camera NetworksabstractThe precise estimation of camera poses within large camera networks is a foundational problem in computer vision and robotics, with broad applications spanning autonomous navigation, surveillance, and augmented reality. In this paper, we introduce a novel methodology that extends state-of-the-art Pose Graph Optimization (PGO) techniques. Departing from the conventional PGO paradigm, which primarily relies on camera-camera edges, our approach centers on the introduction of a dynamic element - any rigid object free to move in the scene - whose pose can be reliably inferred from a single image. Specifically, we consider the bipartite graph encompassing cameras, object poses evolving dynamically, and camera-object relative transformations at each time step. This shift not only offers a solution to the challenges encountered in directly estimating relative poses between cameras, particularly in adverse environments, but also leverages the inclusion of numerous object poses to ameliorate and integrate errors, resulting in accurate camera pose estimates. Though our framework retains compatibility with traditional PGO solvers, its efficacy benefits from a custom-tailored optimization scheme. To this end, we introduce an iterative primal-dual algorithm, capable of handling large graphs. Empirical benchmarks, conducted on a new dataset of simulated indoor environments, substantiate the efficacy and efficiency of our approach. Gabriel Moreira, Manuel Marques, João Paulo Costeira, Alex Hauptmann 0001 |
ICRA | 4 |
| 2024 | Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningabstractAccurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling.
However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset. Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001 |
NeurIPS | 8 |
| 2024 | Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human InteractionsabstractVision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-and-Language Navigation (HA-VLN), extending traditional VLN by incorporating dynamic human activities and relaxing key assumptions. We propose the Human-Aware 3D (HA3D) simulator, which combines dynamic human activities with the Matterport3D dataset, and the Human-Aware Room-to-Room (HA-R2R) dataset, extending R2R with human activity descriptions. To tackle HA-VLN challenges, we present the Expert-Supervised Cross-Modal (VLN-CM) and Non-Expert-Supervised Decision Transformer (VLN-DT) agents, utilizing cross-modal fusion and diverse training strategies for effective navigation in dynamic human environments. A comprehensive evaluation, including metrics considering human activities, and systematic analysis of HA-VLN's unique challenges, underscores the need for further research to enhance HA-VLN agents' real-world robustness and adaptability. Ultimately, this work provides benchmarks and insights for future research on embodied AI and Sim2Real transfer, paving the way for more realistic and applicable VLN systems in human-populated environments. Zhi-Qi Cheng, Yifei Dong 0002, Yuxuan Zhou 0004, Jun-Yan He, Qi Dai 0001, Teruko Mitamura, Alex Hauptmann 0001 |
NeurIPS | 9 |
| 2024 | Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsabstractImproving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation. Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alex Hauptmann 0001, Zhi-Qi Cheng, Kyungwoo Song |
NeurIPS | 7 |
| 2024 | Hyperbolic vs Euclidean Embeddings in Few-Shot Learning: Two Sides of the Same CoinabstractRecent research in representation learning has shown that hierarchical data lends itself to low-dimensional and highly informative representations in hyperbolic space. However, even if hyperbolic embeddings have gathered attention in image recognition, their optimization is prone to numerical hurdles. Further, it remains unclear which applications stand to benefit the most from the implicit bias imposed by hyperbolicity, when compared to traditional Euclidean features. In this paper, we focus on prototypical hyperbolic neural networks. In particular, the tendency of hyperbolic embeddings to converge to the boundary of the Poincaré ball in high dimensions and the effect this has on few-shot classification. We show that the best few-shot results are attained for hyperbolic embeddings at a common hyperbolic radius. In contrast to prior benchmark results, we demonstrate that better performance can be achieved by a fixed-radius encoder equipped with the Euclidean metric, regardless of the embedding dimension. Gabriel Moreira, Manuel Marques, João Paulo Costeira, Alex Hauptmann 0001 |
WACV | 4 |
| 2023 | MAGVIT: Masked Generative Video TransformerabstractWe introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning. We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT. Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600. (ii) MAGVIT outperforms existing methods in inference time by two orders of magnitude against diffusion models and by 60x against autoregressive models. (iii) A single MAGVIT model supports ten diverse generation tasks and generalizes across videos from different visual domains. The source code and trained models will be released to the public at https://magvit.cs.cmu.edu. Lijun Yu, Yong Cheng 0003, Kihyuk Sohn, José Lezama, Han Zhang 0010, Huiwen Chang, Alex Hauptmann 0001, Ming-Hsuan Yang 0001, Yuan Hao, Irfan A. Essa, Lu Jiang 0004 |
CVPR | 7 |
| 2023 | STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action RecognitionabstractWe study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standard-ized skeleton representations as model input, we propose a novel Spatial-Temporal Mesh Transformer (STMT) to directly model the mesh sequences. The model uses a hierarchical transformer with intra-frame off-set attention and inter-frame self-attention. The attention mechanism allows the model to freely attend between any two vertex patches to learn nonlocal relationships in the spatial-temporal domain. Masked vertex modeling and future frame prediction are used as two self-supervised tasks to fully activate the bi-directional and auto-regressive attention in our hierarchical transformer. The proposed method achieves state-of-the-art performance compared to skeleton-based and point-cloud-based models on common MoCap benchmarks. Code is available at https://github.com/zgzxy001/STMT. Po-Yao Huang 0001, Junwei Liang 0001, Celso de Melo, Alex Hauptmann 0001 |
CVPR | 5 |
| 2023 | ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic RulesabstractCharts are a powerful tool for visually conveying complex data, but their comprehension poses a challenge due to the diverse chart types and intricate components. Existing chart comprehension methods suffer from either heuristic rules or an over-reliance on OCR systems, resulting in suboptimal performance. To address these issues, we present ChartReader, a unified framework that seamlessly integrates chart derendering and comprehension tasks. Our approach includes a transformer-based chart component detection module and an extended pre-trained vision-language model for chart-to-X tasks. By learning the rules of charts automatically from annotated datasets, our approach eliminates the need for manual rule-making, reducing effort and enhancing accuracy. We also introduce a data variable replacement technique and extend the input and position embeddings of the pre-trained model for cross-task training. We evaluate ChartReader on Chart-to-Table, ChartQA, and Chart-to-Text tasks, demonstrating its superiority over existing methods. Our proposed framework can significantly reduce the manual effort involved in chart analysis, providing a step towards a universal chart understanding model. Moreover, our approach offers opportunities for plug-and-play integration with mainstream LLMs such as T5 and TaPas, extending their capability to chart comprehension tasks.1 Zhi-Qi Cheng, Qi Dai 0001, Alex Hauptmann 0001 |
ICCV | 3 |
| 2023 | Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative DescriptionsabstractGiven its wide applications, there is increasing focus on generating 3D human motions from textual descriptions. Differing from the majority of previous works, which regard actions as single entities and can only generate short sequences for simple motions, we propose EMS, an elaborative motion synthesis model conditioned on detailed natural language descriptions. It generates natural and smooth motion sequences for long and complicated actions by factorizing them into groups of atomic actions. Meanwhile, it understands atomic-action level attributes (e.g., motion direction, speed, and body parts) and enables users to generate sequences of unseen complex actions from unique sequences of known atomic actions with independent attribute settings and timings applied. We evaluate our method on the KIT Motion-Language and BABEL benchmarks, where it outperforms all previous state-of-the-art with noticeable margins. Yijun Qian, Jack Urbanek, Alex Hauptmann 0001, Jungdam Won |
ICCV | 3 |
| 2023 | SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMsabstractIn this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the rich semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks.
Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%. Lijun Yu, Yong Cheng 0003, Zhiruo Wang 0001, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan A. Essa, Yonatan Bisk, Ming-Hsuan Yang 0001, Kevin Murphy 0002, Alex Hauptmann 0001, Lu Jiang 0004 |
NeurIPS | 12 |
| 2023 | A Comprehensive Survey of Scene Graphs: Generation and ApplicationabstractScene graph is a structured representation of a scene that can clearly express the objects, attributes, and relationships between objects in the scene. As computer vision technology continues to develop, people are no longer satisfied with simply detecting and recognizing objects in images; instead, people look forward to a higher level of understanding and reasoning about visual scenes. For example, given an image, we want to not only detect and recognize objects in the image, but also understand the relationship between objects (visual relationship detection), and generate a text description (image captioning) based on the image content. Alternatively, we might want the machine to tell us what the little girl in the image is doing (Visual Question Answering (VQA)), or even remove the dog from the image and find similar images (image editing and retrieval), etc. These tasks require a higher level of understanding and reasoning for image vision tasks. The scene graph is just such a powerful tool for scene understanding. Therefore, scene graphs have attracted the attention of a large number of researchers, and related research is often cross-modal, complex, and rapidly developing. However, no relatively systematic survey of scene graphs exists at present. To this end, this survey conducts a comprehensive investigation of the current scene graph research. More specifically, we first summarize the general definition of the scene graph, then conducte a comprehensive and systematic discussion on the generation method of the scene graph (SGG) and the SGG with the aid of prior knowledge. We then investigate the main applications of scene graphs and summarize the most commonly used datasets. Finally, we provide some insights into the future development of scene graphs. Xiaojun Chang, Pengzhen Ren, Pengfei Xu 0003, Zhihui Li 0001, Xiaojiang Chen, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Video Pivoting Unsupervised Multi-Modal Machine TranslationabstractThe main challenge in the field of unsupervised machine translation (UMT) is to associate source-target sentences in the latent space. As people who speak different languages share biologically similar visual systems, various unsupervised multi-modal machine translation (UMMT) models have been proposed to improve the performances of UMT by employing visual contents in natural images to facilitate alignment. Commonly, relation information is the important semantic in a sentence. Compared with images, videos can better present the interactions between objects and the ways in which an object transforms over time. However, current state-of-the-art methods only explore scene-level or object-level information from images without explicitly modeling objects relation; thus, they are sensitive to spurious correlations, which poses a new challenge for UMMT models. In this paper, we employ a spatial-temporal graph obtained from videos to exploit object interactions in space and time for disambiguation purposes and to promote latent space alignment in UMMT. Our model employs multi-modal back-translation and features pseudo-visual pivoting, in which we learn a shared multilingual visual-semantic embedding space and incorporate visually pivoted captioning as additional weak supervision. Experimental results on the VATEX Translation 2020 and HowToWorld datasets validate the translation capabilities of our model on both sentence-level and word-level and generalizes well when videos are not available during the testing phase. Mingjie Li 0006, Po-Yao Huang 0001, Xiaojun Chang, Junjie Hu 0001, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | TN-ZSTAD: Transferable Network for Zero-Shot Temporal Activity DetectionabstractAn integral part of video analysis and surveillance is temporal activity detection, which means to simultaneously recognize and localize activities in long untrimmed videos. Currently, the most effective methods of temporal activity detection are based on deep learning, and they typically perform very well with large scale annotated videos for training. However, these methods are limited in real applications due to the unavailable videos about certain activity classes and the time-consuming data annotation. To solve this challenging problem, we propose a novel task setting called zero-shot temporal activity detection (ZSTAD), where activities that have never been seen in training still need to be detected. We design an end-to-end deep transferable network TN-ZSTAD as the architecture for this solution. On the one hand, this network utilizes an activity graph transformer to predict a set of activity instances that appear in the video, rather than produces many activity proposals in advance. On the other hand, this network captures the common semantics of seen and unseen activities from their corresponding label embeddings, and it is optimized with an innovative loss function that considers the classification property on seen activities and the transfer property on unseen activities together. Experiments on the THUMOS'14, Charades, and ActivityNet datasets show promising performance in terms of detecting unseen activities. Lingling Zhang 0005, Xiaojun Chang, Jun Liu 0002, Minnan Luo, Zhihui Li 0001, Lina Yao 0001, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Rethinking Spatial Invariance of Convolutional Networks for Object CountingabstractPrevious work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too strict pixel-level spatial invariance would cause overfit noise in the density map generation. In this paper, we try to use locally connected Gaussian kernels to replace the original convolution filter to estimate the spatial position in the density map. The purpose of this is to allow the feature extraction process to potentially stimulate the density map generation process to overcome the annotation noise. Inspired by previous work, we propose a low-rank approximation accompanied with translation invariance to favorably implement the approximation of massive Gaussian convolution. Our work points a new direction for follow-up research, which should investigate how to properly relax the overly strict pixel-level spatial invariance for object counting. We evaluate our methods on 4 mainstream object counting networks (i.e., MCNN, CSRNet, SANet, and ResNet-50). Extensive experiments were conducted on 7 popular benchmarks for 3 applications (i.e., crowd, vehicle, and plant counting). Experimental results show that our methods significantly outperform other state-of-the-art methods and achieve promising learning of the spatial position of objects11Code is at https://github.com/zhiqic/Rethinking-Counting. Zhi-Qi Cheng, Qi Dai 0001, Jingkuan Song, Xiao Wu 0001, Alex Hauptmann 0001 |
CVPR | 6 |
| 2022 | Speech Driven Tongue AnimationabstractAdvances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not provide realistic animation of the inner mouth. This paper addresses the problem of speech-driven inner mouth animation. Obtaining performance capture data of the tongue and jaw from video alone is difficult because the inner mouth is only partially observable during speech. In this work, we introduce a large-scale speech and mocap dataset that focuses on capturing tongue, jaw, and lip motion. This dataset enables research using data-driven techniques to generate realistic inner mouth animation from speech. We then propose a deep-learning based method for accurate and generalizable speech to tongue and jaw animation, and evaluate several encoder-decoder network architectures and audio feature encoders. We find that recent self-supervised deep learning based audio feature encoders are robust, generalize well to unseen speakers and content, and work best for our task. To demonstrate the practical application of our approach, we show animations on high-quality parametric 3D face models driven by the landmarks generated from our speech-to-tongue animation method. Denis Tomè, Carsten Stoll, Mark K. Tiede, Kevin Munhall, Alex Hauptmann 0001, Iain A. Matthews |
CVPR | 6 |
| 2022 | Rethinking Zero-shot Action Recognition: Learning from Latent Atomic Actions
Yijun Qian, Lijun Yu, Wenhe Liu, Alex Hauptmann 0001 |
ECCV (4) | 4 |
| 2022 | GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementabstractGrounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like'' event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but also predicts all corresponding semantic roles (e.g. agent and goods). Inspired by object detection and image captioning tasks, existing methods typically employ a two-stage framework: 1) detect the activity verb, and then 2) predict semantic roles based on the detected verb. Obviously, this illogical framework constitutes a huge obstacle to semantic understanding. First, pre-detecting verbs solely without semantic roles inevitably fails to distinguish many similar daily activities (e.g., offering and giving, buying and selling). Second, predicting semantic roles in a closed auto-regressive manner can hardly exploit the semantic relations among the verb and roles. To this end, in this paper we propose a novel two-stage framework that focuses on utilizing such bidirectional relations within verbs and roles. In the first stage, instead of pre-detecting the verb, we postpone the detection step and assume a pseudo label, where an intermediate representation for each corresponding semantic role is learned from images. In the second stage, we exploit transformer layers to unearth the potential semantic relations within both verbs and semantic roles. With the help of a set of support images, an alternate learning scheme is designed to simultaneously optimize the results: update the verb using nouns corresponding to the image, and update nouns using verbs from support images. Extensive experimental results on challenging SWiG benchmarks show that our renovated framework outperforms other state-of-the-art methods under various metrics. Zhi-Qi Cheng, Qi Dai 0001, Siyao Li, Teruko Mitamura, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2022 | KAT: A Knowledge Augmented Transformer for Vision-and-LanguageabstractLiangke Gui, Borui Wang, Qiuyuan Huang, Alexander Hauptmann, Yonatan Bisk, Jianfeng Gao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Liangke Gui, Borui Wang, Qiuyuan Huang, Alex Hauptmann 0001, Yonatan Bisk, Jianfeng Gao 0001 |
NAACL-HLT | 4 |
| 2022 | Deep Discrete Cross-Modal Hashing with Multiple Supervision
En Yu, Jiande Sun 0001, Xiaojun Chang, Huaxiang Zhang 0001, Alex Hauptmann 0001 |
Neurocomputing | 6 |
| 2022 | Contrastive Adaptation Network for Single- and Multi-Source Domain AdaptationabstractUnsupervised domain adaptation (UDA) makes predictions for the target domain data while manual annotations are only available in the source domain. Previous methods minimize the domain discrepancy neglecting the class information, which may lead to misalignment and poor generalization performance. To tackle this issue, this paper proposes contrastive adaptation network (CAN) that optimizes a new metric named Contrastive Domain Discrepancy explicitly modeling the intra-class domain discrepancy and the inter-class domain discrepancy. To optimize CAN, two technical issues need to be addressed: 1) the target labels are not available; and 2) the conventional mini-batch sampling is imbalanced. Thus we design an alternating update strategy to optimize both the target label estimations and the feature representations. Moreover, we develop class-aware sampling to enable more efficient and effective training. Our framework can be generally applied to the single-source and multi-source domain adaptation scenarios. In particular, to deal with multiple source domain data, we propose: 1) multi-source clustering ensemble which exploits the complementary knowledge of distinct source domains to make more accurate and robust target label estimations; and 2) boundary-sensitive alignment to make the decision boundary better fitted to the target. Experiments are conducted on three real-world benchmarks (i.e., Office-31 and VisDA-2017 for the single-source scenario, DomainNet for the multi-source scenario). All the results demonstrate that our CAN performs favorably against the state-of-the-art methods. Ablation studies also verify the effectiveness of each key component of our proposed system. Guoliang Kang, Lu Jiang 0004, Yunchao Wei, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Statistical Distance Metric Learning for Image Set RetrievalabstractMeasuring similarity between two image sets is instrumental in many computer vision tasks, such as video face recognition, multi-shot person re-identification and gait recognition. In most of the recent works, it is done by aggregating the embedding features of images as a fixed size vector, and calculating a metric in vector space (i.e. Euclidean distance). The embedding feature function can be learned by deep metric learning (DML) technique. However, methods relying on feature aggregation fail to capture the diversity and uncertainty within image sets. In this paper, we obviate the need of feature aggregation and propose a novel Statistical Distance Metric Learning (SDML) framework, which represents each image set as a probability distribution in embedding feature space and compares two image sets by statistical distance between their distributions. Among all types of statistical distance, we choose Jeffrey’s divergence (JD), which can be obtained from two embedding feature sets by kNN based density estimator. We also design a statistical centroid loss function to enhance the discriminative power of training process. Our SDML framework naturally preserves the diversity within an image set, and the relation between two sets. We evaluate our proposed approach on gait recognition and multi-shot person re-id. The experiment results show that SDML outperforms conventional DML, and also receives competitive/superior performance comparing to the previous state-of-the-arts on the aforementioned tasks. Ting-Yao Hu, Alex Hauptmann 0001 |
ICASSP | 2 |
| 2021 | Learning to Hallucinate Examples from Extrinsic and Intrinsic SupervisionabstractLearning to hallucinate additional examples has recently been shown as a promising direction to address few-shot learning tasks. This work investigates two important yet overlooked natural supervision signals for guiding the hallucination process – (i) extrinsic: classifiers trained on hallucinated examples should be close to strong classifiers that would be learned from a large amount of real examples; and (ii) intrinsic: clusters of hallucinated and real examples belonging to the same class should be pulled together, while simultaneously pushing apart clusters of hallucinated and real examples from different classes. We achieve (i) by introducing an additional mentor model on data-abundant base classes for directing the hallucinator, and achieve (ii) by performing contrastive learning between hallucinated and real examples. As a general, model-agnostic framework, our dual mentor-and self-directed (DMAS) hallucinator significantly improves few-shot learning performance on widely-used benchmarks in various scenarios. Liangke Gui, Adrien Bardes, Ruslan Salakhutdinov, Alex Hauptmann 0001, Martial Hebert, Yu-Xiong Wang |
ICCV | 4 |
| 2021 | Pose Guided Person Image Generation With Hidden P-Norm RegressionabstractIn this paper, we propose a novel approach to solve the pose guided person image generation task. We assume that the relation between pose and appearance information can be described by a simple matrix operation in hidden space. Based on this assumption, our method estimates a pose-invariant feature matrix for each identity, and uses it to predict the target appearance conditioned on the target pose. The estimation process is formulated as a p-norm regression problem in hidden space. By utilizing the differentiation of the solution of this regression problem, the parameters of the whole framework can be trained in an end-to-end manner. While most previous works are only applicable to the supervised training and single-shot generation scenario, our method can be easily adapted to unsupervised training and multi-shot generation. Extensive experiments on the challenging Market-1501 dataset show that our method yields competitive performance in all the aforementioned variant scenarios. Ting-Yao Hu, Alex Hauptmann 0001 |
ICIP | 2 |
| 2021 | Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang 0001, Yuki Markus Asano, Florian Metze, Alex Hauptmann 0001, João F. Henriques, Andrea Vedaldi |
ICLR | 5 |
| 2021 | Person Search Challenges and Solutions: A SurveyabstractPerson search has drawn increasing attention due to its real-world applications and research significance. Person search aims to find a probe person in a gallery of scene images with a wide range of applications, such as criminals search, multicamera tracking, missing person search, etc. Early person search works focused on image-based person search, which uses person image as the search query. Text-based person search is another major person search category that uses free-form natural language as the search query. Person search is challenging, and corresponding solutions are diverse and complex. Therefore, systematic surveys on this topic are essential. This paper surveyed the recent works on image-based and text-based person search from the perspective of challenges and solutions. Specifically, we provide a brief analysis of highly influential person search methods considering the three significant challenges: the discriminative person features, the query-person gap, and the detection-identification inconsistency. We summarise and compare evaluation results. Finally, we discuss open issues and some promising future research directions. Xiangtan Lin, Pengzhen Ren, Xiaojun Chang, Alex Hauptmann 0001 |
IJCAI | 5 |
| 2021 | Importance of Parasagittal Sensor Information in Tongue Motion Capture Through a Diphonic AnalysisabstractOur study examines the information obtained by adding two parasagittal sensors to the standard midsagittal configuration of an Electromagnetic Articulography (EMA) observation of lingual articulation. In this work, we present a large and phonetically balanced corpus obtained from an EMA recording session of a single English native speaker reading 1899 sentences from the Harvard and TIMIT corpora. According to a statistical analysis of the diphones produced during the recording session, the motion captured by the parasagittal sensors has a low correlation to the midsagittal sensors in the mediolateral direction. We perform a geometric analysis of the lateral tongue by the measure of its width and using a proxy of the tongue’s curvature that is computed using the Menger curvature. To provide a better understanding of the tongue sensor motion we present dynamic visualizations of all diphones. Finally, we present a summary of the velocity information computed from the tongue sensor information. Sarah Taylor, Mark K. Tiede, Alex Hauptmann 0001, Iain A. Matthews |
Interspeech | 4 |
| 2021 | MMPT'21: International Joint Workshop on Multi-Modal Pre-Training for Multimedia UnderstandingabstractPre-training has been an emerging topic that provides a way to learn strong representation in many fields (e.g., natural language processing, computing vision). In the last few years, we have witnessed many research works on multi-modal pre-training which have achieved state-of-the-art performances on many multimedia tasks (e.g., image-text retrieval, video localization, speech recognition). In this workshop, we aim to gather peer researchers on related topics for more insightful discussion. We also intend to attract more researchers to explore and investigate more opportunities of designing and using innovative pre-training models for multimedia tasks. Bei Liu 0001, Jianlong Fu, Shizhe Chen, Qin Jin, Alex Hauptmann 0001, Yong Rui |
ICMR | 5 |
| 2021 | MuCAI'21: 2nd ACM Multimedia Workshop on Multimodal Conversational AIabstractThe second edition of the International Workshop on Multimodal Conversational AI puts forward a diverse set of contributions that aim to brainstorm this new field. Conversational agents are now becoming a commodity as this technology is being applied to a wide range of domains. Healthcare, assisting technologies, e-commerce, information seeking, are some of the domains where multimodal conversational AI is being explored. The wide use of multimodal conversational agents exposes the many challenges in achieving more natural, human-like, and engaging conversational agents. The research contributions of the Workshop actively address several of relevant challenges: How to include assistive-technologies in dialog systems? How can agents engage in negotiation in dialogs? How to handle the embodiment of conversational agents? João Magalhães, Alex Hauptmann 0001, Ricardo Gamelas Sousa, Carlos Santiago |
ACM Multimedia | 2 |
| 2021 | Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language ModelsabstractPo-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alexander Hauptmann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Po-Yao Huang 0001, Mandela Patrick, Junjie Hu 0001, Graham Neubig, Florian Metze, Alex Hauptmann 0001 |
NAACL-HLT | 6 |
| 2021 | MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial VideosabstractIn this paper, we study the problem of efficiently assessing building damage after natural disasters like hurricanes, floods or fires, through aerial video analysis. We make two main contributions. The first contribution is a new dataset, consisting of user-generated aerial videos from social media with annotations of instance-level building damage masks. This provides the first benchmark for quantitative evaluation of models to assess building damage using aerial videos. The second contribution is a new model, namely MSNet, which contains novel region proposal network designs and an unsupervised score refinement network for confidence score calibration in both bounding box and mask branches. We show that our model achieves state-of-the-art results compared to previous methods in our dataset. Junwei Liang 0001, Alex Hauptmann 0001 |
WACV | 3 |
| 2020 | Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingabstractUnsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora only.However, it is still challenging to associate source-target sentences in the latent space.As people speak different languages biologically share similar visual systems, the potential of achieving better alignment through visual content is promising yet under-explored in unsupervised multimodal MT (MMT).In this paper, we investigate how to utilize visual content for disambiguation and promoting latent space alignment in unsupervised MMT.Our model employs multimodal back-translation and features pseudo visual pivoting in which we learn a shared multilingual visual-semantic embedding space and incorporate visuallypivoted captioning as additional weak supervision.The experimental results on the widely used Multi30K dataset show that the proposed model significantly improves over the state-ofthe-art methods and generalizes well when images are not available at the testing time. Po-Yao Huang 0001, Junjie Hu 0001, Xiaojun Chang, Alex Hauptmann 0001 |
ACL | 4 |
| 2020 | The Garden of Forking Paths: Towards Multi-Future Trajectory PredictionabstractThis paper studies the problem of predicting the distribution over multiple possible future paths of people as they move through various visual scenes. We make two main contributions. The first contribution is a new dataset, created in a realistic 3D simulator, which is based on real world trajectory data, and then extrapolated by human annotators to achieve different latent goals. This provides the first benchmark for quantitative evaluation of the models to predict multi-future trajectories. The second contribution is a new model to generate multiple plausible future trajectories, which contains novel designs of using multi-scale location encodings and convolutional RNNs over graphs. We refer to our model as Multiverse. We show that our model achieves the best results on our dataset, as well as on the real-world VIRAT/ActEV dataset (which just contains one possible future). Junwei Liang 0001, Lu Jiang 0004, Kevin Murphy 0002, Ting Yu 0003, Alex Hauptmann 0001 |
CVPR | 5 |
| 2020 | ZSTAD: Zero-Shot Temporal Activity DetectionabstractAn integral part of video analysis and surveillance is temporal activity detection, which means to simultaneously recognize and localize activities in long untrimmed videos. Currently, the most effective methods of temporal activity detection are based on deep learning, and they typically perform very well with large scale annotated videos for training. However, these methods are limited in real applications due to the unavailable videos about certain activity classes and the time-consuming data annotation. To solve this challenging problem, we propose a novel task setting called zero-shot temporal activity detection (ZSTAD), where activities that have never been seen in training can still be detected. We design an end-to-end deep network based on R-C3D as the architecture for this solution. The proposed network is optimized with an innovative loss function that considers the embeddings of activity labels and their super-classes while learning the common semantics of seen and unseen activities. Experiments on both the THUMOS’14 and the Charades datasets show promising performance in terms of detecting unseen activities. Lingling Zhang 0005, Xiaojun Chang, Jun Liu 0002, Minnan Luo, Sen Wang 0001, ZongYuan Ge, Alex Hauptmann 0001 |
CVPR | 7 |
| 2020 | SimAug: Learning Robust Representations from Simulation for Trajectory Prediction
Junwei Liang 0001, Lu Jiang 0004, Alex Hauptmann 0001 |
ECCV (13) | 3 |
| 2020 | Stacked Pooling for Boosting Scale Invariance of Crowd CountingabstractIn this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance. Existing deep crowd counting approaches mainly focus on the multi-scale techniques over convolutional layers to capture scale-adaptive features, resulting in high computing costs. In this paper, we propose simple but effective pooling variants, i.e., multi-kernel pooling and stacked pooling, to take place of the vanilla pooling layers in convolutional neural networks (CNNs) for boosting the scale invariance. Our proposed pooling modules do not introduce extra parameters and can be easily implemented in practice. Empirical studies on two benchmark crowd counting datasets show that the proposed pooling modules beat the vanilla pooling layer in most experimental cases. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ICASSP | 5 |
| 2020 | Forward and Backward Multimodal NMT for Improved Monolingual and Multilingual Cross-Modal RetrievalabstractWe explore methods to enrich the diversity of captions associated with pictures for learning improved visual-semantic embeddings (VSE) in cross-modal retrieval. In the spirit of "A picture is worth a thousand words", it would take dozens of sentences to parallel each picture's content adequately. But in fact, real-world multimodal datasets tend to provide only a few (typically, five) descriptions per image. For cross-modal retrieval, the resulting lack of diversity and coverage prevents systems from capturing the fine-grained inter-modal dependencies and intra-modal diversities in the shared VSE space. Using the fact that the encoder-decoder architectures in neural machine translation (NMT) have the capacity to enrich both monolingual and multilingual textual diversity, we propose a novel framework leveraging multimodal neural machine translation (MMT) to perform forward and backward translations based on salient visual objects to generate additional text-image pairs which enables training improved monolingual cross-modal retrieval (English-Image) and multilingual cross-modal retrieval (English-Image and German-Image) models. Experimental results show that the proposed framework can substantially and consistently improve the performance of state-of-the-art models on multiple datasets. The results also suggest that the models with multilingual VSE outperform the models with monolingual VSE. Po-Yao Huang 0001, Xiaojun Chang, Alex Hauptmann 0001, Eduard H. Hovy |
ICMR | 3 |
| 2020 | MuCAI'20: 1st International Workshop on Multimodal Conversational AIabstractRecently, conversational systems have seen a significant rise in demand due to modern commercial applications using systems such as Amazon's Alexa, Apple's Siri, Microsoft's Cortana and Google Assistant. The research on multimodal chatbots is a widely underexplored area, where users and the conversational agent communicate by natural language and visual data. Conversational agents are now becoming a commodity as a number of companies push for this technology. The wide use of these conversational agents exposes the many challenges in achieving more natural, human-like, and engaging conversational agents. The research community is actively addressing several of these challenges: how are visual and text data related in user utterances? How to interpret the user intent? How to encode multimodal dialog status? What are the ethical and legal aspects of conversational AI? The Multimodal Conversational AI workshop will be a forum where researchers and practitioners share their experiences and brainstorm about success and failures in the topic. It will also promote collaboration to strengthen the conversational AI community at ACM Multimedia. Alex Hauptmann 0001, João Magalhães, Ricardo Gamelas Sousa, João Paulo Costeira |
ACM Multimedia | 1 |
| 2020 | Pixel-Level Cycle Association: A New Perspective for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to train a model performing satisfactory pixel-level predictions on the target with only out-of-domain (source) annotations. The conventional solution to this task is to minimize the discrepancy between source and target to enable effective knowledge transfer. Previous domain discrepancy minimization methods are mainly based on the adversarial training. They tend to consider the domain discrepancy globally, which ignore the pixel-wise relationships and are less discriminative. In this paper, we propose to build the pixel-level cycle association between source and target pixel pairs and contrastively strengthen their connections to diminish the domain gap and make the features more discriminative. To the best of our knowledge, this is a new perspective for tackling such a challenging task. Experiment results on two representative domain adaptation benchmarks, i.e. GTAV $\rightarrow$ Cityscapes and SYNTHIA $\rightarrow$ Cityscapes, verify the effectiveness of our proposed method and demonstrate that our method performs favorably against previous state-of-the-arts. Our method can be trained end-to-end in one stage and introduce no additional parameters, which is expected to serve as a general framework and help ease future research in domain adaptive semantic segmentation. Code is available at https://github.com/kgl-prml/Pixel-Level-Cycle-Association. Guoliang Kang, Yunchao Wei, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
NeurIPS | 5 |
| 2020 | Few-shot activity recognition with cross-modal memory network
Lingling Zhang 0005, Xiaojun Chang, Jun Liu 0002, Minnan Luo, Mahesh Prakash, Alex Hauptmann 0001 |
Pattern Recognit. | 6 |
| 2020 | Simultaneous Bearing Fault Recognition and Remaining Useful Life Prediction Using Joint-Loss Convolutional Neural NetworkabstractFault diagnosis and remaining useful life (RUL) prediction are always two major issues in modern industrial systems, which are usually regarded as two separated tasks to make the problem easier but ignore the fact that there are certain information of these two tasks that can be shared to improve the performance. Therefore, to capture common features between different relative problems, a joint-loss convolutional neural network (JL-CNN) architecture is proposed in this paper, which can implement bearing fault recognition and RUL prediction in parallel by sharing the parameters and partial networks, meanwhile keeping the output layers of different tasks. The JL-CNN is constructed based on a CNN, which is a widely used deep learning method because of its powerful feature extraction ability. During optimization phase, a JL function is designed to enable the proposed approach to learn the diagnosis-prognosis features and improve generalization while reducing the overfitting risk and computation cost. Moreover, because the information behind the signals of different problems has been shared and exploited deeper, the generalization and the accuracy of results can also be improved. Finally, the effectiveness of the JL-CNN method is validated by run-to-failure dataset. Compared with support vector regression and traditional CNN, the mean-square-error of the proposed method decreases 82.7% and 24.9%, respectively. Therefore, results and comparisons show that the proposed method can be applied for the intercrossed applications between fault diagnosis and RUL prediction. Boyuan Yang 0002, Alex Hauptmann 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Semantics-Preserving Graph Propagation for Zero-Shot Object DetectionabstractMost existing object detection models are restricted to detecting objects from previously seen categories, an approach that tends to become infeasible for rare or novel concepts. Accordingly, in this paper, we explore object detection in the context of zero-shot learning, i.e., Zero-Shot Object Detection (ZSD), to concurrently recognize and localize objects from novel concepts. Existing ZSD algorithms are typically based on a simple mapping-transfer strategy that is susceptible to the domain shift problem. To resolve this problem, we propose a novel Semantics-Preserving Graph Propagation model for ZSD based on Graph Convolutional Networks (GCN). More specifically, we employ a graph construction module to flexibly build category graphs by incorporating diverse correlations between category nodes; this is followed by two semantics preserving modules that enhance both category and region representations through a multi-step graph propagation process. Compared to existing mapping-transfer based methods, both the semantic description and semantic structural knowledge exhibited in prior category graphs can be effectively leveraged to boost the generalization capability of the learned projection function via knowledge transfer, thereby providing a solution to the domain shift problem. Experiments on existing seen/unseen splits of three popular object detection datasets demonstrate that the proposed approach performs favorably against state-of-the-art ZSD methods. Caixia Yan, Xiaojun Chang, Minnan Luo, Chung-Hsing Yeh, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Pair-based Uncertainty and Diversity Promoting Early Active Learning for Person Re-identificationabstractThe effective training of supervised Person Re-identification (Re-ID) models requires sufficient pairwise labeled data. However, when there is limited annotation resource, it is difficult to collect pairwise labeled data. We consider a challenging and practical problem called Early Active Learning, which is applied to the early stage of experiments when there is no pre-labeled sample available as references for human annotating. Previous early active learning methods suffer from two limitations for Re-ID. First, these instance-based algorithms select instances rather than pairs, which can result in missing optimal pairs for Re-ID. Second, most of these methods only consider the representativeness of instances, which can result in selecting less diverse and less informative pairs. To overcome these limitations, we propose a novel pair-based active learning for Re-ID. Our algorithm selects pairs instead of instances from the entire dataset for annotation. Besides representativeness, we further take into account the uncertainty and the diversity in terms of pairwise relations. Therefore, our algorithm can produce the most representative, informative, and diverse pairs for Re-ID data annotation. Extensive experimental results on five benchmark Re-ID datasets have demonstrated the superiority of the proposed pair-based early active learning algorithm. Wenhe Liu, Xiaojun Chang, Ling Chen 0006, Dinh Q. Phung, Xiaoqin Zhang 0002, Yi Yang 0001, Alex Hauptmann 0001 |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2020 | Learning Distilled Graph for Large-Scale Social Network Data ClusteringabstractSpectral analysis is critical in social network analysis. As a vital step of the spectral analysis, the graph construction in many existing works utilizes content data only. Unfortunately, the content data often consists of noisy, sparse, and redundant features, which makes the resulting graph unstable and unreliable. In practice, besides the content data, social network data also contain link information, which provides additional information for graph construction. Some of previous works utilize the link data. However, the link data is often incomplete, which makes the resulting graph incomplete. To address these issues, we propose a novel Distilled Graph Clustering (DGC) method. It pursuits adistilled graphbased on both the content data and the link data. The proposed algorithm alternates between two steps: in the feature selection step, it finds the most representative feature subset w.r.t. an intermediate graph initialized with link data; in graph distillation step, the proposed method updates and refines the graph based on only the selected features. The final resulting graph, which is referred to as the distilled graph, is then utilized for spectral clustering on the large-scale social network data. Extensive experiments demonstrate the superiority of the proposed method. Wenhe Liu, Dong Gong, Mingkui Tan, Qinfeng Shi, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Deep Top-$k$ Ranking for Image-Sentence MatchingabstractImage-sentence matching is a challenging task for the heterogeneity-gap between different modalities. Ranking-based methods have achieved excellent performance in this task in past decades. Given an image query, these methods typically assume that the correct matched image-sentence pair must rank before all other mismatched ones. However, this assumption may be too strict and prone to the overfitting problem, especially when some sentences in a massive database are similar and confusable with one another. In this paper, we relax the traditional ranking loss and propose a novel deep multi-modal network with a top-k ranking loss to mitigate the data ambiguity problem. With this strategy, query results will not be penalized unless the index of ground truth is outside the range of top-k query results. Considering the non-smoothness and non-convexity of the initial top-k ranking loss, we exploit a tight convex upper bound to approximate the loss and then utilize the traditional back-propagation algorithm to optimize the deep multi-modal network. Finally, we apply the method on three benchmark datasets, namely, Flickr8k, Flickr30k, and MSCOCO. Empirical results on metrics R@K (K = 1, 5, 10) show that our method achieves comparable performance in comparison to state-of-the-art methods. Lingling Zhang 0005, Minnan Luo, Jun Liu 0002, Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Fuzzy Least Squares Support Vector Machine With Adaptive Membership for Object TrackingabstractFuzzy learning has been introduced into tracking and achieved great success. However, the membership in the existing fuzzy learning based tracking algorithm is fixed, which lacks the adaptivity to measure the importance of the samples. To improve the tracking adaptivity and flexibility, in this paper, we propose a novel tracking method based on fuzzy least squares support vector machine with adaptive membership (FLS-SVM-AM). First, we formulate tracking as an adaptive membership based fuzzy learning problem, which addresses the issue of fixed membership in existing methods and can better measure the importance of the training samples. Second, we present the FLS-SVM-AM method to build the appearance model, and develop an iterative optimization process to solve the FLS-SVM-AM problem. Third, we define a new membership based on the PASCAL VOC overlap rate and exponential function, which is used to measure the importance of different samples more accurately. Experimental results in the benchmark datasets demonstrate that the proposed method not only outperforms the existing fuzzy learning based tracking methods, but also is comparable to many state-of-the-art methods. Shunli Zhang 0005, Li Zhang 0023, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Unsupervised Bilingual Lexicon Induction from Mono-Lingual Multimodal DataabstractBilingual lexicon induction, translating words from the source language to the target language, is a long-standing natural language processing task. Recent endeavors prove that it is promising to employ images as pivot to learn the lexicon induction without reliance on parallel corpora. However, these vision-based approaches simply associate words with entire images, which are constrained to translate concrete words and require object-centered images. We humans can understand words better when they are within a sentence with context. Therefore, in this paper, we propose to utilize images and their associated captions to address the limitations of previous approaches. We propose a multi-lingual caption model trained with different mono-lingual multimodal data to map words in different languages into joint spaces. Two types of word representation are induced from the multi-lingual caption model: linguistic features and localized visual features. The linguistic feature is learned from the sentence contexts with visual semantic constraints, which is beneficial to learn translation for words that are less visual-relevant. The localized visual feature is attended to the region in the image that correlates to the word, so that it alleviates the image restriction for salient visual representation. The two types of features are complementary for word translation. Experimental results on multiple language pairs demonstrate the effectiveness of our proposed method, which substantially outperforms previous vision-based approaches without using any parallel sentences or supervision of seed word pairs. Shizhe Chen, Qin Jin, Alex Hauptmann 0001 |
AAAI | 3 |
| 2019 | Training-free Monocular 3D Event Detection System for Traffic SurveillanceabstractWe focus on the problem of detecting traffic events in a surveillance scenario, including the detection of both vehicle actions and traffic collisions. Existing event detection systems are mostly learning-based and have achieved convincing performance when a large amount of training data is available. However, in real-world scenarios, collecting sufficient labeled training data is expensive and sometimes impossible (e.g. for traffic collision detection). Moreover, the conventional 2D representation of surveillance views is easily affected by occlusions and different camera views in nature. To deal with the aforementioned problems, in this paper, we propose a training-free monocular 3D event detection system for traffic surveillance. Our system firstly projects the vehicles into the 3D Euclidean space and estimates their kinematic states. Then we develop multiple simple yet effective ways to identify the events based on the kinematic patterns, which need no further training. Consequently, our system is robust to the occlusions and the viewpoint changes. Exclusive experiments report the superior result of our method on large-scale real-world surveillance datasets, which validates the effectiveness of our proposed system. The demonstration videos of our system are available online1. Lijun Yu, Wenhe Liu, Guoliang Kang, Alex Hauptmann 0001 |
IEEE BigData | 5 |
| 2019 | Shooter Localization Using Videos in the WildabstractNowadays a huge number of user-generated videos are uploaded to social media every second, capturing glimpses of events all over the world. These videos in the wild provide important and useful information for reconstructing events like the Las Vegas Shooting in 2017. In this paper, we describe a system that can localize the shooter location only based on a couple of user-generated videos that capture the gunshot sound. Our system first utilizes established video analysis techniques like video synchronization and automatic gunshot processing to organize the unstructured videos in the wild for users to understand the event effectively. By combining multimodal information from visual, audio and geo-locations, our system can then visualize all possible locations of the shooter in the map. Our system provides a web interface for human-in-the-loop verification to ensure accurate estimations. We present the results of estimating the shooter's location of the Las Vegas Shooting in 2017 and show that our system is able to get accurate location using only the first few gunshots. All relevant source code including the web interface and machine learning models are available. Junwei Liang 0001, Jay D. Aronson, Alex Hauptmann 0001 |
CBMI | 3 |
| 2019 | Contrastive Adaptation Network for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) makes predictions for the target domain data while manual annotations are only available in the source domain. Previous methods minimize the domain discrepancy neglecting the class information, which may lead to misalignment and poor generalization performance. To address this issue, this paper proposes Contrastive Adaptation Network (CAN) optimizing a new metric which explicitly models the intra-class domain discrepancy and the inter-class domain discrepancy. We design an alternating update strategy for training CAN in an end-to-end manner. Experiments on two real-world benchmarks Office-31 and VisDA-2017 demonstrate that CAN performs favorably against the state-of-the-art methods and produces more discriminative features. Guoliang Kang, Lu Jiang 0004, Yi Yang 0001, Alex Hauptmann 0001 |
CVPR | 4 |
| 2019 | Peeking Into the Future: Predicting Future Person Activities and Locations in VideosabstractDeciphering human behaviors to predict their future paths/trajectories and what they would do from videos is important in many applications. Motivated by this idea, this paper studies predicting a pedestrian's future path jointly with future activities. We propose an end-to-end, multi-task learning system utilizing rich visual features about human behavioral information and interaction with their surroundings. To facilitate the training, the network is learned with an auxiliary task of predicting future location in which the activity will happen. Experimental results demonstrate our state-of-the-art performance over two public benchmarks on future trajectory prediction. Moreover, our method is able to produce meaningful future activity prediction in addition to the path. The result provides the first empirical evidence that joint modeling of paths and activities benefits future path prediction. Junwei Liang 0001, Lu Jiang 0004, Juan Carlos Niebles, Alex Hauptmann 0001, Li Fei-Fei 0001 |
CVPR | 4 |
| 2019 | Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal RepresentationsabstractPo-Yao Huang, Xiaojun Chang, Alexander Hauptmann. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Po-Yao Huang 0001, Xiaojun Chang, Alex Hauptmann 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Learning Spatial Awareness to Improve Crowd CountingabstractThe aim of crowd counting is to estimate the number of people in images by leveraging the annotation of center positions for pedestrians' heads. Promising progresses have been made with the prevalence of deep Convolutional Neural Networks. Existing methods widely employ the Euclidean distance (i.e., L2loss) to optimize the model, which, however, has two main drawbacks: (1) the loss has difficulty in learning the spatial awareness (i.e., the position of head) since it struggles to retain the high-frequency variation in the density map, and (2) the loss is highly sensitive to various noises in crowd counting, such as the zeromean noise, head size changes, and occlusions. Although the Maximum Excess over SubArrays (MESA) loss has been previously proposed by [16] to address the above issues by finding the rectangular subregion whose predicted density map has the maximum difference from the ground truth, it cannot be solved by gradient descent, thus can hardly be integrated into the deep learning framework. In this paper, we present a novel architecture called SPatial Awareness Network (SPANet) to incorporate spatial context for crowd counting. The Maximum Excess over Pixels (MEP) loss is proposed to achieve this by finding the pixel-level subregion with high discrepancy to the ground truth. To this end, we devise a weakly supervised learning scheme to generate such region with a multi-branch architecture. The proposed framework can be integrated into existing deep crowd counting methods and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that our method can significantly improve the performance of baselines. More remarkably, our approach outperforms the state-of-the-art methods on all benchmark datasets. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Alex Hauptmann 0001 |
ICCV | 5 |
| 2019 | Learning Sound Events from Webly Labeled DataabstractIn the last couple of years, weakly labeled learning has turned out to be an exciting approach for audio event detection. In this work, we introduce webly labeled learning for sound events which aims to remove human supervision altogether from the learning process. We first develop a method of obtaining labeled audio data from the web (albeit noisy), in which no manual labeling is involved. We then describe methods to efficiently learn from these webly labeled audio recordings. In our proposed system, WeblyNet, two deep neural networks co-teach each other to robustly learn from webly labeled data, leading to around 17% relative improvement over the baseline method. The method also involves transfer learning to obtain efficient representations. Anurag Kumar 0003, Ankit Shah 0001, Alex Hauptmann 0001, Bhiksha Raj |
IJCAI | 3 |
| 2019 | Multi-shot Person Re-identification through Set Distance with Visual Distributional RepresentationabstractPerson re-identification aims to identify a specific person at distinct times and locations. It is challenging because of occlusion, illumination, and viewpoint change in camera views. Recently, multi-shot person re-id task receives more attention since it is closer to real-world application. A key point of a good algorithm for multi-shot person re-id is the temporal aggregation of the person appearance features. While most of the current approaches apply pooling strategies and obtain a fixed-size vector representation, these may lose the matching evidence between examples. In this work, we propose the idea of visual distributional representation, which interprets an image set as samples drawn from an unknown distribution in appearance feature space. Based on the supervision signals from a downstream task of interest, the method reshapes the appearance feature space and further learns the unknown distribution of each image set. In the context of multi-shot person re-id, we apply this novel concept along with Wasserstein distance and jointly learn a distributional set distance function between two image sets. In this way, the proper alignment between two image sets can be discovered naturally in a non-parametric manner. Our experiment results on three public datasets show the advantages of our proposed method compared to other state-of-the-art approaches. Ting-Yao Hu, Alex Hauptmann 0001 |
ICMR | 2 |
| 2019 | Improving What Cross-Modal Retrieval Models Learn through Object-Oriented Inter- and Intra-Modal Attention NetworksabstractAlthough significant progress has been made for cross-modal retrieval models in recent years, few have explored what those models truly learn and what makes one model superior to another. Start by training two state-of-the-art text-to-image retrieval models with adversarial text inputs, we investigate and quantify the importance of syntactic structure and lexical information in learning the joint visual-semantic embedding space for cross-modal retrieval. The results show that the retrieval power mainly comes from localizing and connecting the visual objects and their cross-modal counter-parts, the textual phrases. Inspired by this observation, we propose a novel model which employs object-oriented encoders along with inter- and intra-modal attention networks to improve inter-modal dependencies for cross-modal retrieval. In addition, we develop a new multimodal structure-preserving objective which additionally emphasizes intra-modal hard negative examples to promote intra-modal discrepancies. Extensive experiments show that the proposed approach outperforms the existing best method by a large margin (16.4% and 6.7% relatively with [email protected] in the text-to-image retrieval task on the Flickr30K dataset and the MS-COCO dataset respectively). Po-Yao Huang 0001, Vaibhav, Xiaojun Chang, Alex Hauptmann 0001 |
ICMR | 4 |
| 2019 | PANEL: Challenges for Multimedia/Multimodal Research in the Next DecadeabstractThe multimedia and multi-modal community is witnessing an explosive transformation in the recent years with major societal impact. With the unprecedented deployment of multimedia devices and systems, multimedia research is critical to our abilities and prospects in advancing state-of-the-art technologies and solving real-world challenges facing the society and the nation. To respond to these challenges and further advance the frontiers of the field of multimedia, this panel will discuss the challenges and visions that may guide future research in the next ten years. Shih-Fu Chang, Louis-Philippe Morency, Alex Hauptmann 0001, Alberto Del Bimbo, Cathal Gurrin, Hayley Hung, Heng Ji 0001, Alan F. Smeaton |
ACM Multimedia | 3 |
| 2019 | Improving the Learning of Multi-column Convolutional Neural Network for Crowd CountingabstractTremendous variation in the scale of people/head size is a critical problem for crowd counting. To improve the scale invariance of feature representation, recent works extensively employ Convolutional Neural Networks with multi-column structures to handle different scales and resolutions. However, due to the substantial redundant parameters in columns, existing multi-column networks invariably exhibit almost the same scale features in different columns, which severely affects counting accuracy and leads to overfitting. In this paper, we attack this problem by proposing a novel Multicolumn Mutual Learning (McML) strategy. It has two main innovations: 1) A statistical network is incorporated into the multi-column framework to estimate the mutual information between columns, which can approximately indicate the scale correlation between features from different columns. By minimizing the mutual information, each column is guided to learn features with different image scales. 2) We devise a mutual learning scheme that can alternately optimize each column while keeping the other columns fixed on each mini-batch training data. With such asynchronous parameter update process, each column is inclined to learn different feature representation from others, which can efficiently reduce the parameter redundancy and improve generalization ability. More remarkably, McML can be applied to all existing multi-column networks and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that McML can significantly improve the original multi-column networks and outperform the other state-of-the-art approaches. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He, Alex Hauptmann 0001 |
ACM Multimedia | 6 |
| 2019 | Annotation Efficient Cross-Modal Retrieval with Adversarial Attentive AlignmentabstractVisual-semantic embeddings are central to many multimedia applications such as cross-modal retrieval between visual data and natural language descriptions. Conventionally, learning a joint embedding space relies on large parallel multimodal corpora. Since massive human annotation is expensive to obtain, there is a strong motivation in developing versatile algorithms to learn from large corpora with fewer annotations. In this paper, we propose a novel framework to leverage automatically extracted regional semantics from un-annotated images as additional weak supervision to learn visual-semantic embeddings. The proposed model employs adversarial attentive alignments to close the inherent heterogeneous gaps between annotated and un-annotated portions of visual and textual domains. To demonstrate its superiority, we conduct extensive experiments on sparsely annotated multimodal corpora. The experimental results show that the proposed model outperforms state-of-the-art visual-semantic embedding models by a significant margin for cross-modal retrieval tasks on the sparse Flickr30k and MS-COCO datasets. It is also worth noting that, despite using only 20% of the annotations, the proposed model can achieve competitive performance (Recall at 10 > 80.0% for 1K and > 70.0% for 5K text-to-image retrieval) compared to the benchmarks trained with the complete annotations. Po-Yao Huang 0001, Guoliang Kang, Wenhe Liu, Xiaojun Chang, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2019 | Shooter Localization Using Social Media VideosabstractNowadays a huge number of user-generated videos are uploaded to social media every second, capturing glimpses of events all over the world. These videos provide important and useful information for reconstructing events like the Las Vegas Shooting in 2017. In this paper, we describe a system that can localize the shooter location only based on a couple of user-generated videos that capture the gunshot sound. Our system first utilizes established video analysis techniques like video synchronization and gunshot temporal localization to organize the unstructured social media videos for users to understand the event effectively. By combining multimodal information from visual, audio and geo-locations, our system can then visualize all possible locations of the shooter in the map. Our system provides a web interface for human-in-the-loop verification to ensure accurate estimations. We present the results of estimating the shooter's location of the Las Vegas Shooting in 2017 and show that our system is able to get accurate location using only the first few gunshots. The full technical report, all relevant source code including the web interface and machine learning models are available. Junwei Liang 0001, Jay D. Aronson, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2019 | Focal Visual-Text Attention for Memex Question AnsweringabstractRecent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photo albums, we have to look at whole collections with sequences of photos. This paper proposes a new multimodal MemexQA task: given a sequence of photos from a user, the goal is to automatically answer questions that help users recover their memory about an event captured in these photos. In addition to a text answer, a few grounding photos are also given to justify the answer. The grounding photos are necessary as they help users quickly verifying the answer. Towards solving the task, we 1) present the MemexQA dataset, the first publicly available multimodal question answering dataset consisting of real personal photo albums; 2) propose an end-to-end trainable network that makes use of a hierarchical process to dynamically determine what media and what time to focus on in the sequential data to answer the question. Experimental results on the MemexQA dataset demonstrate that our model outperforms strong baselines and yields the most relevant grounding photos on this challenging task. Junwei Liang 0001, Lu Jiang 0004, Liangliang Cao, Yannis Kalantidis, Li-Jia Li 0001, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Scheduled sampling for one-shot learning via matching network
Lingling Zhang 0005, Jun Liu 0002, Minnan Luo, Xiaojun Chang, Alex Hauptmann 0001 |
Pattern Recognit. | 6 |
| 2019 | Automatic Vacant Parking Places Management System Using Multicamera Vehicle DetectionabstractThis paper presents a multicamera system for vehicles detection and their corresponding mapping into the parking spots of a parking lot. Approaches from the state-of-the-art system, which work properly in controlled scenarios, have been validated using small amount of sequences and without more challenging realistic conditions (illumination changes and different weather conditions). On the other hand, most of them are not complete systems, but provide only parts of them, usually detectors. The proposed system has been designed for realistic scenarios considering different cases of occlusion, illumination changes, and different climatic conditions; a real scenario (the International Pittsburgh Airport parking lot) has been targeted with the condition that existing parking security cameras can be used, avoiding the deployment of new cameras or other sensors infrastructures. For design and validation, a new multicamera data set has been recorded. The system is based on existing object detectors (the results of two of them are shown) and different proposed postprocessing stages. The results clearly show that the proposed system works correctly in challenging scenarios including almost total occlusions, illumination changes, and different weather conditions. Rafael Martin Nieto, Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Generating Video Descriptions With Latent Topic GuidanceabstractAutomatic video description generation (a.k.a video captioning) is one of the ultimate goals for video understanding. Despite the wide range of applications such as video indexing and retrieval etc., the video captioning task remains quite challenging due to the complexity and diversity of video content. First, open-domain videos cover a broad range of topics, which results in highly variable vocabularies and expression styles to describe the video contents. Second, videos naturally contain multiple modalities including image, motion, and acoustic media. The information provided by different modalities differs in different conditions. In this paper, we propose a novel topic-guided video captioning model to address the above-mentioned challenges in video captioning. Our model consists of two joint tasks, namely, latent topic generation and topic-guided caption generation. The topic generation task aims to automatically predict the latent topic of the video. Since there is no groundtruth topic information, we mine multimodal topics in an unsupervised fashion based on video contents and annotated captions, and then distill the topic distribution to a topic prediction model. In the topic-guided generation task, we employ the topic guidance for two purposes. The first is to narrow down the language complexity across topics, where we propose the topic-aware decoder to leverage the latent topics to induce topic-related language models. The decoder is also generic and can be integrated with a temporal attention mechanism. The second is to dynamically attend to important modalities by topics, where we propose a flexible topic-guided multimodal ensemble framework and use the topic gating network to determine the attention weights. The two tasks are correlated with each other, and they collaborate to generate more detailed and accurate video captions. Our extensive experiments on two public benchmark datasets MSR-VTT and Youtube2Text demonstrate the effectiveness of the proposed topic-guided video captioning system, which achieves state-of-the-art performance on both datasets. Shizhe Chen, Qin Jin, Jia Chen 0001, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Adaptive Semi-Supervised Feature Selection for Cross-Modal RetrievalabstractIn order to exploit the abundant potential information of the unlabeled data and contribute to analyzing the correlation among heterogeneous data, we propose the semi-supervised model named adaptive semi-supervised feature selection for cross-modal retrieval. First, we utilize the semantic regression to strengthen the neighboring relationship between the data with the same semantic. And the correlation between heterogeneous data can be optimized via keeping the pairwise closeness when learning the common latent space. Second, we adopt the graph-based constraint to predict accurate labels for unlabeled data, and it can also keep the geometric structure consistency between the label space and the feature space of heterogeneous data in the common latent space. Finally, an efficient joint optimization algorithm is proposed to update the mapping matrices and the label matrix for unlabeled data simultaneously and iteratively. It makes samples from different classes to be far apart, while the samples from same class lie as close as possible. Meanwhile, the l2,1-norm constraint is used for feature selection and outlier reduction when the mapping matrices are learned. In addition, we propose learning different mapping matrices corresponding to different sub-tasks to emphasize the semantic and structural information of query data. Experiment results on three datasets demonstrate that our method performs better than the state-of-the-art methods. En Yu, Jiande Sun 0001, Jing Li 0046, Xiaojun Chang, Xianhua Han, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 6 |
| 2018 | Hidden Two-Stream Convolutional Networks for Action Recognition
Yi Zhu 0001, Zhen-Zhong Lan, Shawn D. Newsam, Alex Hauptmann 0001 |
ACCV (3) | 4 |
| 2018 | CADP: A Novel Dataset for CCTV Traffic Camera based Accident AnalysisabstractThis paper presents a novel dataset for traffic accidents analysis. Our goal is to resolve the lack of public data for research about automatic spatio-temporal annotations for traffic safety in the roads. Through the analysis of the proposed dataset, we observed a significant degradation of object detection in pedestrian category in our dataset, due to the object sizes and complexity of the scenes. To this end, we propose to integrate contextual information into conventional Faster R-CNN using Context Mining (CM) and Augmented Context Mining (ACM) to complement the accuracy for small pedestrian detection. Our experiments indicate a considerable improvement in object detection accuracy: +8.51% for CM and +6.20% for ACM. Finally, we demonstrate the performance of accident forecasting in our dataset using Faster R-CNN and an Accident LSTM architecture. We achieved an average of 1.684 seconds in terms of Time-To-Accident measure with an Average Precision of 47.25%. Our Webpage for the paper is https://goo.gl/cqK2wE. Ankit Parag Shah, Jean-Baptiste Lamare, Tuan Nguyen-Anh, Alex Hauptmann 0001 |
AVSS | 4 |
| 2018 | Traffic Danger Recognition With Surveillance Cameras Without Training DataabstractWe propose a traffic danger recognition model that works with arbitrary traffic surveillance cameras to identify and predict car crashes. There are too many cameras to monitor manually. Therefore, we developed a model to predict and identify car crashes from surveillance cameras based on a 3D reconstruction of the road plane and prediction of trajectories. For normal traffic, it supports real-time proactive safety checks of speeds and distances between vehicles to provide insights about possible high-risk areas. We achieve good prediction and recognition of car crashes without using any labeled training data of crashes. Experiments on the BrnoCompSpeed dataset show that our model can accurately monitor the road, with mean errors of 1.80% for distance measurement, 2.77 km/h for speed measurement, 0.24 m for car position prediction, and 2.53 km/h for speed prediction. Lijun Yu, Xiangqun Chen, Alex Hauptmann 0001 |
AVSS | 4 |
| 2018 | Adaptive Context-aware Reinforced Agent for Handwritten Text Recognition
Liangke Gui, Xiaodan Liang, Xiaojun Chang, Alex Hauptmann 0001 |
BMVC | 4 |
| 2018 | Towards Independent Stress Detection: A Dependent Model Using Facial Action UnitsabstractOur society is increasingly more susceptible to chronic stress. Reasons are daily worries, workload, and the wish to fulfil a myriad of expectations. Unfortunately, long-exposure to stress leads to physical and mental health problems. To avoid the described consequences, mobile applications have been studied to track stress in combination with wearables. However, wearables need to be worn all day long and can be costly. Given that most laptops have inbuilt cameras, using video data for personal tracking of stress levels could be a more affordable alternative. In previous work, videos have been used to detect cognitive stress during driving by measuring the presence of anger or fear through a limited number of facial expressions. In contrast, we propose the use of 17 facial action units (AUs) not solely restricted to those emotions. We used five one-hour long videos from the dataset collected by Lau [1]. The videos show subjects while typing, resting, and exposed to a stressor, being a multitasking exercise combined with social evaluation. We performed binary classification using several simple classifiers on AUs extracted in each video frame and were able to achieve an accuracy of up to 74% in subject independent classification and 91% in subject dependent classification. These preliminary results indicate that the AUs most relevant for stress detection are not consistently the same for all 5 subjects. Also in previous work, using facial cues, a strong person-specific component was found during classification. Carla Viegas, Shing-Hon Lau, Roy A. Maxion, Alex Hauptmann 0001 |
CBMI | 4 |
| 2018 | DecideNet: Counting Varying Density Crowds Through Attention Guided Detection and Density EstimationabstractIn real-world crowd counting applications, the crowd densities vary greatly in spatial and temporal domains. A detection based counting method will estimate crowds accurately in low density scenes, while its reliability in congested areas is downgraded. A regression based approach, on the other hand, captures the general density information in crowded regions. Without knowing the location of each person, it tends to overestimate the count in low density areas. Thus, exclusively using either one of them is not sufficient to handle all kinds of scenes with varying densities. To address this issue, a novel end-to-end crowd counting framework, named DecideNet (DEteCtIon and Density Estimation Network) is proposed. It can adaptively decide the appropriate counting mode for different locations on the image based on its real density conditions. DecideNet starts with estimating the crowd density by generating detection and regression based density maps separately. To capture inevitable variation in densities, it incorporates an attention module, meant to adaptively assess the reliability of the two types of estimations. The final crowd counts are obtained with the guidance of the attention module to adopt suitable estimations from the two kinds of density maps. Experimental results show that our method achieves state-of-the-art performance on three challenging crowd counting datasets. Jiang Liu 0011, Chenqiang Gao, Deyu Meng, Alex Hauptmann 0001 |
CVPR | 4 |
| 2018 | Focal Visual-Text Attention for Visual Question AnsweringabstractRecent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photos, we have to look at whole collections with sequences of photos or videos. When answering questions from a large collection, a natural problem is to identify snippets to support the answer. In this paper, we describe a novel neural network called Focal Visual-Text Attention network (FVTA) for collective reasoning in visual question answering, where both visual and text sequence information such as images and text metadata are presented. FVTA introduces an end-to-end approach that makes use of a hierarchical process to dynamically determine what media and what time to focus on in the sequential data to answer the question. FVTA can not only answer the questions well but also provides the justifications which the system results are based upon to get the answers. FVTA achieves state-of-the-art performance on the MemexQA dataset and competitive results on the MovieQA dataset. Junwei Liang 0001, Lu Jiang 0004, Liangliang Cao, Li-Jia Li 0001, Alex Hauptmann 0001 |
CVPR | 5 |
| 2018 | RCAA: Relational Context-Aware Agents for Person Search
Xiaojun Chang, Po-Yao Huang 0001, Yidong Shen, Xiaodan Liang, Yi Yang 0001, Alex Hauptmann 0001 |
ECCV (9) | 6 |
| 2018 | Class-aware Self-Attention for Audio Event RecognitionabstractAudio event recognition (AER) has been an important research problem with a wide range of applications. However, it is very challenging to develop large scale audio event recognition models. On the one hand, usually there are only "weak" labeled audio training data available, which only contains labels of audio events without temporal boundaries. On the other hand, the distribution of audio events is generally long-tailed, with only a few positive samples for large amounts of audio events. These two issues make it hard to learn discriminative acoustic features to recognize audio events especially for long-tailed events. In this paper, we propose a novel class-aware self-attention mechanism with attention factor sharing to generate discriminative clip-level features for audio event recognition. Since a target audio event only occurs in part of an entire audio clip and its corresponding temporal interval varies, the proposed class-aware self-attention approach learns to highlight relevant temporal intervals and to suppress irrelevant noises at the same time. In order to learn attention patterns effectively for those long-tailed events, we combine both the domain knowledge and data driven strategies to share attention factors in the proposed attention mechanism, which transfers the common knowledge learned from other similar events to the rare events. The proposed attention mechanism is a pluggable component and can be trained end-to-end in the overall AER model. We evaluate our model on a large-scale audio event corpus "Audio Set" with both short-term and long-term acoustic features. The experimental results demonstrate the effectiveness of our model, which improves the overall audio event recognition performance with different acoustic features especially for events with low resources. Moreover, the experiments also show that our proposed model is able to learn new audio events with a few training examples effectively and efficiently without disturbing the previously learned audio events. Shizhe Chen, Jia Chen 0001, Qin Jin, Alex Hauptmann 0001 |
ICMR | 4 |
| 2018 | Multimodal Filtering of Social Media for Temporal Monitoring and Event AnalysisabstractDeveloping an efficient and effective social media monitoring system has become one of the important steps towards improved public safety. With the explosive availability of user-generated content documenting most conflicts and human rights abuses around the world, analysts and first-responders increasingly find themselves overwhelmed with massive amounts of noisy data from social media. In this paper, we construct a large-scale public safety event dataset with retrospective automatic labeling for 4.2 million multimodal tweets from 7 public safety events occurred in 2013~2017. We propose a new multimodal social media filtering system composed of encoding, classification, and correlation networks to jointly learn shared and complementary visual and textual information to filter out the most relevant and useful items among the noisy social media influx. The proposed model is verified and achieves significant improvement over competitive baselines under the retrospective and real-time experimental protocols. Po-Yao Huang 0001, Junwei Liang 0001, Jean-Baptiste Lamare, Alex Hauptmann 0001 |
ICMR | 4 |
| 2018 | Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model SearchabstractAs attribute leaning brings mid-level semantic properties for objects, it can benefit many traditional learning problems in multimedia and computer vision communities. When facing the huge number of attributes, it is extremely challenging to automatically design a generalizable neural network for other attribute learning tasks. Even for a specific attribute domain, the exploration of the neural network architecture is always optimized by a combination of heuristics and grid search, from which there is a large space of possible choices to be searched. In this paper, Generalizable Attribute Learning Model (GALM) is proposed to automatically design the neural networks for generalizable attribute learning. The main novelty of GALM is that it fully exploits the Multi-Task Learning and Reinforcement Learning to speed up the search procedure. With the help of parameter sharing, GALM is able to transfer the pre-searched architecture to different attribute domains. In experiments, we comprehensively evaluate GALM on 251 attributes from three domains: animals, objects, and scenes. Extensive experimental results demonstrate that GALM significantly outperforms the state-of-the-art attribute learning approaches and previous neural architecture search methods on two generalizable attribute learning scenarios. Zhi-Qi Cheng, Xiao Wu 0001, Siyu Huang, Jun-Xiu Li, Alex Hauptmann 0001, Qiang Peng |
ACM Multimedia | 5 |
| 2018 | GNAS: A Greedy Neural Architecture Search Method for Multi-Attribute LearningabstractA key problem in deep multi-attribute learning is to effectively discover the inter-attribute correlation structures. Typically, the conventional deep multi-attribute learning approaches follow the pipeline of manually designing the network architectures based on task-specific expertise prior knowledge and careful network tunings, leading to the inflexibility for various complicated scenarios in practice. Motivated by addressing this problem, we propose an efficient greedy neural architecture search approach (GNAS) to automatically discover the optimal tree-like deep architecture for multi-attribute learning. In a greedy manner, GNAS divides the optimization of global architecture into the optimizations of individual connections step by step. By iteratively updating the local architectures, the global tree-like architecture gets converged where the bottom layers are shared across relevant attributes and the branches in top layers more encode attribute-specific features. Experiments on three benchmark multi-attribute datasets show the effectiveness and compactness of neural architectures derived by GNAS, and also demonstrate the efficiency of GNAS in searching neural architectures. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2018 | A unified framework with a benchmark dataset for surveillance event detection
Zhicheng Zhao 0001, Xuanchong Li, Xingzhong Du, Qi Chen 0014, Yanyun Zhao, Xiaojun Chang, Alex Hauptmann 0001 |
Neurocomputing | 8 |
| 2018 | Deep feature learning via structured graph Laplacian embedding for person re-identification
De Cheng, Yihong Gong, Xiaojun Chang, Weiwei Shi 0003, Alex Hauptmann 0001, Nanning Zheng 0001 |
Pattern Recognit. | 5 |
| 2018 | An Adaptive Semisupervised Feature Analysis for Video Semantic RecognitionabstractVideo semantic recognition usually suffers from the curse of dimensionality and the absence of enough high-quality labeled instances, thus semisupervised feature selection gains increasing attentions for its efficiency and comprehensibility. Most of the previous methods assume that videos with close distance (neighbors) have similar labels and characterize the intrinsic local structure through a predetermined graph of both labeled and unlabeled data. However, besides the parameter tuning problem underlying the construction of the graph, the affinity measurement in the original feature space usually suffers from the curse of dimensionality. Additionally, the predetermined graph separates itself from the procedure of feature selection, which might lead to downgraded performance for video semantic recognition. In this paper, we exploit a novel semisupervised feature selection method from a new perspective. The primary assumption underlying our model is that the instances with similar labels should have a larger probability of being neighbors. Instead of using a predetermined similarity graph, we incorporate the exploration of the local structure into the procedure of joint feature selection so as to learn the optimal graph simultaneously. Moreover, an adaptive loss function is exploited to measure the label fitness, which significantly enhances model's robustness to videos with a small or substantial loss. We propose an efficient alternating optimization algorithm to solve the proposed challenging problem, together with analyses on its convergence and computational complexity in theory. Finally, extensive experimental results on benchmark datasets illustrate the effectiveness and superiority of the proposed approach on video semantic recognition related tasks. Minnan Luo, Xiaojun Chang, Liqiang Nie, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Cybern. | 5 |
| 2018 | Few-Shot Text and Image Classification via Analogical Transfer LearningabstractLearning from very few samples is a challenge for machine learning tasks, such as text and image classification. Performance of such task can be enhanced via transfer of helpful knowledge from related domains, which is referred to as transfer learning. In previous transfer learning works, instance transfer learning algorithms mostly focus on selecting the source domain instances similar to the target domain instances for transfer. However, the selected instances usually do not directly contribute to the learning performance in the target domain. Hypothesis transfer learning algorithms focus on the model/parameter level transfer. They treat the source hypotheses as well-trained and transfer their knowledge in terms of parameters to learn the target hypothesis. Such algorithms directly optimize the target hypothesis by the observable performance improvements. However, they fail to consider the problem that instances that contribute to the source hypotheses may be harmful for the target hypothesis, as instance transfer learning analyzed. To relieve the aforementioned problems, we propose a novel transfer learning algorithm, which follows an analogical strategy. Particularly, the proposed algorithm first learns a revised source hypothesis with only instances contributing to the target hypothesis. Then, the proposed algorithm transfers both the revised source hypothesis and the target hypothesis (only trained with a few samples) to learn an analogical hypothesis. We denote our algorithm as Analogical Transfer Learning. Extensive experiments on one synthetic dataset and three real-world benchmark datasets demonstrate the superior performance of the proposed algorithm. Wenhe Liu, Xiaojun Chang, Yan Yan 0002, Yi Yang 0001, Alex Hauptmann 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2018 | Adaptive Unsupervised Feature Selection With Structure RegularizationabstractFeature selection is one of the most important dimension reduction techniques for its efficiency and interpretation. Since practical data in large scale are usually collected without labels, and labeling these data are dramatically expensive and time-consuming, unsupervised feature selection has become a ubiquitous and challenging problem. Without label information, the fundamental problem of unsupervised feature selection lies in how to characterize the geometry structure of original feature space and produce a faithful feature subset, which preserves the intrinsic structure accurately. In this paper, we characterize the intrinsic local structure by an adaptive reconstruction graph and simultaneously consider its multiconnected-components (multicluster) structure by imposing a rank constraint on the corresponding Laplacian matrix. To achieve a desirable feature subset, we learn the optimal reconstruction graph and selective matrix simultaneously, instead of using a predetermined graph. We exploit an efficient alternative optimization algorithm to solve the proposed challenging problem, together with the theoretical analyses on its convergence and computational complexity. Finally, extensive experiments on clustering task are conducted over several benchmark data sets to verify the effectiveness and superiority of the proposed unsupervised feature selection algorithm. Minnan Luo, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Joint Attributes and Event Analysis for Multimedia Event DetectionabstractSemantic attributes have been increasingly used the past few years for multimedia event detection (MED) with promising results. The motivation is that multimedia events generally consist of lower level components such as objects, scenes, and actions. By characterizing multimedia event videos with semantic attributes, one could exploit more informative cues for improved detection results. Much existing work obtains semantic attributes from images, which may be suboptimal for video analysis since these image-inferred attributes do not carry dynamic information that is essential for videos. To address this issue, we propose to learn semantic attributes from external videos using their semantic labels. We name them video attributes in this paper. In contrast with multimedia event videos, these external videos depict lower level contents such as objects, scenes, and actions. To harness video attributes, we propose an algorithm established on a correlation vector that correlates them to a target event. Consequently, we could incorporate video attributes latently as extra information into the event detector learnt from multimedia event videos in a joint framework. To validate our method, we perform experiments on the real-world large-scale TRECVID MED 2013 and 2014 data sets and compare our method with several state-of-the-art algorithms. The experiments show that our method is advantageous for MED. Zhigang Ma, Xiaojun Chang, Zhongwen Xu, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Visual Memory QA: Your Personal Photo and Video Search AgentabstractThe boom of mobile devices and cloud services has led to an explosion of personal photo and video data. However, due to the missing user-generated metadata such as titles or descriptions, it usually takes a user a lot of swipes to find some video on the cell phone. To solve the problem, we present an innovative idea called Visual Memory QA which allow a user not only to search but also to ask questions about her daily life captured in the personal videos. The proposed system automatically analyzes the content of personal videos without user-generated metadata, and offers a conversational interface to accept and answer questions. To the best of our knowledge, it is the first to answer personal questions discovered in personal photos or videos. The example questions are "what was the lat time we went hiking in the forest near San Francisco?"; "did we have pizza last week?"; "with whom did I have dinner in AAAI 2015?". Lu Jiang 0004, Liangliang Cao, Yannis Kalantidis, Sachin Farfade, Alex Hauptmann 0001 |
AAAI | 5 |
| 2017 | An Event Reconstruction Tool for Conflict Monitoring Using Social MediaabstractWhat happened during the Boston Marathon in 2013? Nowadays, at any major event, lots of people take videos and share them on social media. To fully understand exactly what happened in these major events, researchers and analysts often have to examine thousands of these videos manually. To reduce this manual effort, we present an investigative system that automatically synchronizes these videos to a global timeline and localizes them on a map. In addition to alignment in time and space, our system combines various functions for analysis, including gunshot detection, crowd size estimation, 3D reconstruction and person tracking. To our best knowledge, this is the first time a unified framework has been built for comprehensive event reconstruction for social media videos. Junwei Liang 0001, Desai Fan, Po-Yao Huang 0001, Jia Chen 0001, Lu Jiang 0004, Alex Hauptmann 0001 |
AAAI | 7 |
| 2017 | Webly-Supervised Learning of Multimodal Video DetectorsabstractGiven any complicated or specialized video content search query, e.g. ”Batkid (a kid in batman costume)” or ”destroyed buildings”, existing methods require manually labeled data to build detectors for searching. We present a demonstration of an artificial intelligence application, Webly-labeled Learning (WELL) that enables learning of ad-hoc concept detectors over unlimited Internet videos without any manual an-notations. A considerable number of videos on the web are associated with rich but noisy contextual information, such as the title, which provides a type of weak annotations or la-bels of the video content. To leverage this information, our system employs state-of-the-art webly-supervised learning(WELL) (Liang et al. ). WELL considers multi-modal information including deep learning visual, audio and speech features, to automatically learn accurate video detectors based on the user query. The learned detectors from a large number of web videos allow users to search relevant videos over their personal video archives, not requiring any textual metadata,but as convenient as searching on Youtube. Junwei Liang 0001, Lu Jiang 0004, Alex Hauptmann 0001 |
AAAI | 3 |
| 2017 | Probabilistic Non-Negative Matrix Factorization and Its Robust Extensions for Topic ModelingabstractTraditional topic model with maximum likelihood estimate inevitably suffers from the conditional independence of words given the document’s topic distribution. In this paper, we follow the generative procedure of topic model and learn the topic-word distribution and topics distribution via directly approximating the word-document co-occurrence matrix with matrix decomposition technique. These methods include: (1) Approximating the normalized document-word conditional distribution with the documents probability matrix and words probability matrix based on probabilistic non-negative matrix factorization (NMF); (2) Since the standard NMF is well known to be non-robust to noises and outliers, we extended the probabilistic NMF of the topic model to its robust versions using l21-norm and capped l21-norm based loss functions, respectively. The proposed framework inherits the explicit probabilistic meaning of factors in topic models and simultaneously makes the conditional independence assumption on words unnecessary. Straightforward and efficient algorithms are exploited to solve the corresponding non-smooth and non-convex problems. Experimental results over several benchmark datasets illustrate the effectiveness and superiority of the proposed methods. Minnan Luo, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001 |
AAAI | 5 |
| 2017 | Synchronization for multi-perspective videos in the wildabstractIn the era of social media, a large number of user-generated videos are uploaded to the Internet every day, capturing events all over the world. Reconstructing the event truth based on information mined from these videos has been an emerging challenging task. Temporal alignment of videos “in the wild” which capture different moments at different positions with different perspectives is the critical step. In this paper, we propose a hierarchical approach to synchronize videos. Our system utilizes clustered audio-signatures to align video pairs. Global alignment for all videos is then achieved via forming alignable video groups with self-paced learning. Experiments on the Boston Marathon dataset show that the proposed method achieves excellent precision and robustness. Junwei Liang 0001, Po-Yao Huang 0001, Jia Chen 0001, Alex Hauptmann 0001 |
ICASSP | 4 |
| 2017 | Temporal localization of audio events for conflict monitoring in social mediaabstractWith the explosion in the availability of user-generated videos documenting any conflicts and human rights abuses around the world, analysts and researchers increasingly find themselves overwhelmed with massive amounts of video data to acquire and analyze useful information. In this paper, we develop a temporal localization framework for intense audio events in videos which addresses the problem. The proposed method utilizes Localized Self-Paced Reranking (LSPaR) to refine the localization results. LSPaR utilizes samples from easy to noisier ones so that it can overcome the noisiness of the initial retrieval results from user-generated videos. We show our framework's efficacy on localizing intense audio event like gunshot, and further experiments also indicate that our methods can be generalized to localizing other audio events in noisy videos. Junwei Liang 0001, Lu Jiang 0004, Alex Hauptmann 0001 |
ICASSP | 3 |
| 2017 | Complex Event Detection by Identifying Reliable Shots from Untrimmed VideosabstractThe goal of complex event detection is to automatically detect whether an event of interest happens in temporally untrimmed long videos which usually consist of multiple video shots. Observing some video shots in positive (resp. negative) videos are irrelevant (resp. relevant) to the given event class, we formulate this task as a multi-instance learning (MIL) problem by taking each video as a bag and the video shots in each video as instances. To this end, we propose a new MIL method, which simultaneously learns a linear SVM classifier and infers a binary indicator for each instance in order to select reliable training instances from each positive or negative bag. In our new objective function, we balance the weighted training errors and a l1-l2mixed-norm regularization term which adaptively selects reliable shots as training instances from different videos to have them as diverse as possible. We also develop an alternating optimization approach that can efficiently solve our proposed objective function. Extensive experiments on the challenging real-world Multimedia Event Detection (MED) datasets MEDTest-14, MEDTest-13 and CCV clearly demonstrate the effectiveness of our proposed MIL approach for complex event detection. Hehe Fan, Xiaojun Chang, De Cheng, Yi Yang 0001, Dong Xu 0001, Alex Hauptmann 0001 |
ICCV | 6 |
| 2017 | Rewind to track: Parallelized apprenticeship learning with backward trackletsabstractData association, which could be categorized into offline approaches and the online counterparts, is a crucial part of a multi-object tracker in the tracking-by-detection framework. On the one hand, classical offline data association methods exploit all the video data and have high computation cost, which makes them unscalable to long-term offline video data. On the other hand, online approaches have much lower computation cost, but they suffer from ID-switches and tracklet drifting problem when directly applied to offline data as they are only aware of “past” observations. In this paper, we propose a mixed style tracker, which is not only as efficient as the online tracker but also aware of “future” observations in offline setting. We start from a Markov Decision Process (MDP) online tracker and design a parallelized apprenticeship learning algorithm to learn both the reward function and transition policy in MDP. By proposing a rewind to track strategy to generate backward tracklets, future detections in offline data are efficiently utilized to obtain a more stable similarity measurement for association. Experiment results show that our approach achieves the state-of-the-art performance on challenging datasets. Jiang Liu 0011, Jia Chen 0001, De Cheng, Chenqiang Gao, Alex Hauptmann 0001 |
ICME | 5 |
| 2017 | Discriminative Dictionary Learning With Ranking Metric Embedded for Person Re-IdentificationabstractThe goal of person re-identification (Re-Id) is to match pedestrians captured from multiple non-overlapping cameras. In this paper, we propose a novel dictionary learning based method with the ranking metric embedded, for person Re-Id. A new and essential ranking graph Laplacian term is introduced, which minimizes the intra-personal compactness and maximizes the inter-personal dispersion in the objective. Different from the traditional dictionary learning based approaches and their extensions, which just use the same or not information, our proposed method can explore the ranking relationship among the person images, which is essential for such retrieval related tasks. Simultaneously, one distance measurement has been explicitly learned in the model to further improve the performance. Since we have reformulated these ranking constraints into the graph Laplacian form, the proposed method is easy-to-implement but effective. We conduct extensive experiments on three widely used person Re-Id benchmark datasets, and achieve state-of-the-art performances. De Cheng, Xiaojun Chang, Li Liu 0031, Alex Hauptmann 0001, Yihong Gong, Nanning Zheng 0001 |
IJCAI | 4 |
| 2017 | Leveraging Multi-modal Prior Knowledge for Large-scale Concept Learning in Noisy Web DataabstractLearning video concept detectors automatically from the big but noisy web data with no additional manual annotations is a novel but challenging area in the multimedia and the machine learning community. A considerable amount of videos on the web is associated with rich but noisy contextual information, such as the title and other multi-modal information, which provides weak annotations or labels about the video content. To tackle the problem of large-scale noisy learning, We propose a novel method called Multi-modal WEbly-Labeled Learning (WELL-MM), which is established on the state-of-the-art machine learning algorithm inspired by the learning process of human. WELL-MM introduces a novel multi-modal approach to incorporate meaningful prior knowledge called curriculum from the noisy web videos. We empirically study the curriculum constructed from the multi-modal features of the Internet videos and images. The comprehensive experimental results on FCVID and YFCC100M demonstrate that WELL-MM outperforms state-of-the-art studies by a statically significant margin on learning concepts from noisy web video data. In addition, the results also verify that WELL-MM is robust to the level of noisiness in the video data. Notably, WELL-MM trained on sufficient noisy web labels is able to achieve a better accuracy to supervised learning methods trained on the clean manually labeled data. Junwei Liang 0001, Lu Jiang 0004, Deyu Meng, Alex Hauptmann 0001 |
ICMR | 4 |
| 2017 | Joint Saliency Estimation and Matching using Image Regions for Geo-Localization of Online VideoabstractIn this paper, we study automatic geo-localization of online event videos. Different from general image localization task through matching, the appearance of an environment during significant events varies greatly from its daily appearance, since there are usually crowds, decorations or even destruction when a major event happens. This introduces a major challenge: matching the event environment to the daily environment, e.g. as recorded by Google Street View. We observe that some regions in the image, as part of the environment, still preserve the daily appearance even though the whole image (environment) looks quite different. Based on this observation, we formulate the problem as joint saliency estimation and matching at the image region level, as opposed to the key point or whole-image level. As image-level labels of daily environment are easily generated with GPS information, we treat region based saliency estimation and matching as a weakly labeled learning problem over the training data. Our solution is to iteratively optimize saliency and the region-matching model. For saliency optimization, we derive a closed form solution, which has an intuitive explanation. For region matching model optimization, we use self-paced learning to learn from the pseudo labels generated by (sub-optimal) saliency values. We conduct extensive experiments on two challenging public datasets: Boston Marathon 2013 and Tokyo Time Machine. Experimental results show that our solution significantly improves over matching on whole images and the automatically learned saliency is a strong predictor of distinctive building areas. Freda Shi, Jia Chen 0001, Alex Hauptmann 0001 |
ICMR | 3 |
| 2017 | Video Captioning with Guidance of Multimodal Latent TopicsabstractThe topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified caption framework, M&M TGM, which mines multimodal topics in unsupervised fashion from data and guides the caption decoder with these topics. Compared to pre-defined topics, the mined multimodal topics are more semantically and visually coherent and can reflect the topic distribution of videos better. We formulate the topic-aware caption generation as a multi-task learning problem, in which we add a parallel task, topic prediction, in addition to the caption task. For the topic prediction task, we use the mined topics as the teacher to train a student topic prediction model, which learns to predict the latent topics from multimodal contents of videos. The topic prediction provides intermediate supervision to the learning process. As for the caption task, we propose a novel topic-aware decoder to generate more accurate and detailed video descriptions with the guidance from latent topics. The entire learning procedure is end-to-end and it optimizes both tasks simultaneously. The results from extensive experiments conducted on the MSR-VTT and Youtube2Text datasets demonstrate the effectiveness of our proposed model. M&M TGM not only outperforms prior state-of-the-art methods on multiple evaluation metrics and on both benchmark datasets, but also achieves better generalization ability. Shizhe Chen, Jia Chen 0001, Qin Jin, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2017 | Knowing Yourself: Improving Video Caption via In-depth RecapabstractGenerating natural language descriptions for videos (a.k.a video captioning) has attracted much research attention in recent years, and a lot of models have been proposed to improve the caption performance. However, due to the rapid progress in dataset expansion and feature representation, newly proposed caption models have been evaluated on different settings, which makes it unclear about the contributions from either features or models. Therefore, in this work we aim to gain a deep understanding about "where are we" for the current development of video captioning. First, we carry out extensive experiments to identify the contribution from different components in video captioning task and make fair comparison among several state-of-the-art video caption models. Second, we discover that these state-of-the-art models are complementary so that we could benefit from "wisdom of the crowd" through ensembling and reranking. Finally, we give a preliminary answer to the question "how far are we from the human-level performance in general'' via a series of carefully designed experiments. In summary, our caption models achieve the state-of-the-art performance on the MSR-VTT 2017 challenge, and it is comparable with the average human-level performance on current caption metrics. However, our analysis also shows that we still have a long way to go, such as further improving the generalization ability of current caption models. Qin Jin, Shizhe Chen, Jia Chen 0001, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2017 | MultiEdTech 2017: 1st International Workshop on Multimedia-based Educational and Knowledge Technologies for Personalized and Social Online TrainingabstractEducational and Knowledge Technologies (EdTech), especially in connection to multimedia content and the vision of mobile and personalized learning, is a hot topic in both academia and the business start-ups ecosystem. The driver and enabler of this is on the one side the development and widespread availability of multimedia materials and MOOCs, which represent multimedia content produced specifically for supporting e-learning; and, on the other side, the ever increasing availability of all sorts on information on the Internet and in social media channels (e. g. lectures, research papers, user-generated videos, news items), which, despite not directly targeting e-learning, can prove to be valuable complements to the more targeted learning materials. Although the availability of such content is not a problem these days, finding the right content and associating different relevant pieces of multimedia so as to enable a comprehensive learning experience on a chosen subject is by no means a trivial task. This workshop provides research in areas related to multimedia-based educational and knowledge technologies and particularly on the use of multimedia search and retrieval, analysis and understanding, browsing, summarization, recommendation, and visualization technologies on multimedia content available in specialized learning platforms, the Web, mobile devices and/or social networks for supporting personalized and adaptive e-learning and training. Ansgar Scherp, Vasileios Mezaris, Thomas Köhler 0001, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2017 | Video Search via Ranking Network with Very Few Query Exemplars
De Cheng, Lu Jiang 0004, Yihong Gong, Nanning Zheng 0001, Alex Hauptmann 0001 |
MMM (2) | 5 |
| 2017 | Delving Deep into Personal Photo and Video SearchabstractThe ubiquity of mobile devices and cloud services has led to an unprecedented growth of online personal photo and video collections. Due to the scarcity of personal media search log data, research to date has mainly focused on searching images and videos on the web. However, in order to manage the exploding amount of personal photos and videos, we raise a fundamental question: what are the differences and similarities when users search their own photos versus the photos on the web? To the best of our knowledge, this paper is the first to study personal media search using large-scale real-world search logs. We analyze different types of search sessions mined from Flickr search logs and discover a number of interesting characteristics of personal media search in terms of information needs and click behaviors. The insightful observations will not only be instrumental in guiding future personal media search methods, but also benefit related tasks such as personal photo browsing and recommendation. Our findings suggest there is a significant gap between personal queries and automatically detected concepts, which is responsible for the low accuracy of many personal media search queries. To bridge the gap, we propose the deep query understanding model to learn a mapping from the personal queries to the concepts in the clicked photos. Experimental results verify the efficacy of the proposed method in improving personal media search, where the proposed method consistently outperforms baseline methods. Lu Jiang 0004, Yannis Kalantidis, Liangliang Cao, Sachin Farfade, Jiliang Tang, Alex Hauptmann 0001 |
WSDM | 6 |
| 2017 | Simple to complex cross-modal learning to rank
Minnan Luo, Xiaojun Chang, Zhihui Li 0001, Liqiang Nie, Alex Hauptmann 0001 |
Comput. Vis. Image Underst. | 5 |
| 2017 | Uncovering the Temporal Context for Video Question Answering
Linchao Zhu, Zhongwen Xu, Yi Yang 0001, Alex Hauptmann 0001 |
Int. J. Comput. Vis. | 4 |
| 2017 | Efficient human action recognition using histograms of motion gradients and VLAD with descriptor shape information
I. C. Duta, Jasper R. R. Uijlings, Bogdan Ionescu, Kiyoharu Aizawa, Alex Hauptmann 0001, Nicu Sebe |
Multim. Tools Appl. | 5 |
| 2017 | Avoiding Optimal Mean ℓ2, 1-Norm Maximization-Based Robust PCA for ReconstructionabstractRobust principal component analysis (PCA) is one of the most important dimension-reduction techniques for handling high-dimensional data with outliers. However, most of the existing robust PCA presupposes that the mean of the data is zero and incorrectly utilizes the average of data as the optimal mean of robust PCA. In fact, this assumption holds only for the squared [Formula: see text]-norm-based traditional PCA. In this letter, we equivalently reformulate the objective of conventional PCA and learn the optimal projection directions by maximizing the sum of projected difference between each pair of instances based on [Formula: see text]-norm. The proposed method is robust to outliers and also invariant to rotation. More important, the reformulated objective not only automatically avoids the calculation of optimal mean and makes the assumption of centered data unnecessary, but also theoretically connects to the minimization of reconstruction error. To solve the proposed nonsmooth problem, we exploit an efficient optimization algorithm to soften the contributions from outliers by reweighting each data point iteratively. We theoretically analyze the convergence and computational complexity of the proposed algorithm. Extensive experimental results on several benchmark data sets illustrate the effectiveness and superiority of the proposed method. Minnan Luo, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001 |
Neural Comput. | 5 |
| 2017 | Bi-Level Semantic Representation Analysis for Multimedia Event DetectionabstractMultimedia event detection has been one of the major endeavors in video event analysis. A variety of approaches have been proposed recently to tackle this problem. Among others, using semantic representation has been accredited for its promising performance and desirable ability for human-understandable reasoning. To generate semantic representation, we usually utilize several external image/video archives and apply the concept detectors trained on them to the event videos. Due to the intrinsic difference of these archives, the resulted representation is presumable to have different predicting capabilities for a certain event. Notwithstanding, not much work is available for assessing the efficacy of semantic representation from the source-level. On the other hand, it is plausible to perceive that some concepts are noisy for detecting a specific event. Motivated by these two shortcomings, we propose a bi-level semantic representation analyzing method. Regarding source-level, our method learns weights of semantic representation attained from different multimedia archives. Meanwhile, it restrains the negative influence of noisy or irrelevant concepts in the overall concept-level. In addition, we particularly focus on efficient multimedia event detection with few positive examples, which is highly appreciated in the real-world scenario. We perform extensive experiments on the challenging TRECVID MED 2013 and 2014 datasets with encouraging results that validate the efficacy of our proposed approach. Xiaojun Chang, Zhigang Ma, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Cybern. | 5 |
| 2017 | Feature Interaction Augmented Sparse Learning for Fast Kinect Motion DetectionabstractThe Kinect sensing devices have been widely used in current Human-Computer Interaction entertainment. A fundamental issue involved is to detect users' motions accurately and quickly. In this paper, we tackle it by proposing a linear algorithm, which is augmented by feature interaction. The linear property guarantees its speed whereas feature interaction captures the higher order effect from the data to enhance its accuracy. The Schatten-p norm is leveraged to integrate the main linear effect and the higher order nonlinear effect by mining the correlation between them. The resulted classification model is a desirable combination of speed and accuracy. We propose a novel solution to solve our objective function. Experiments are performed on three public Kinect-based entertainment data sets related to fitness and gaming. The results show that our method has its advantage for motion detection in a real-time Kinect entertaining environment. Xiaojun Chang, Zhigang Ma, Ming Lin 0002, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | The Many Shades of NegativityabstractComplex event detection has been progressively researched in recent years for the broad interest of video indexing and retrieval. To fulfill the purpose of event detection, one needs to train a classifier using both positive and negative examples. Current classifier training treats the negative videos as equally negative. However, we notice that many negative videos resemble the positive videos in different degrees. Intuitively, we may capture more informative cues from the negative videos if we assign them fine-grained labels, thus benefiting the classifier learning. Aiming for this, we use a statistical method on both the positive and negative examples to get the decisive attributes of a specific event. Based on these decisive attributes, we assign the fine-grained labels to negative examples to treat them differently for more effective exploitation. The resulting fine-grained labels may be not optimal to capture the discriminative cues from the negative videos. Hence, we propose to jointly optimize the fine-grained labels with the classifier learning, which brings mutual reciprocality. Meanwhile, the labels of positive examples are supposed to remain unchanged. We thus additionally introduce a constraint for this purpose. On the other hand, the state-of-the-art deep convolutional neural network features are leveraged in our approach for event detection to further boost the performance. Extensive experiments on the challenging TRECVID MED 2014 dataset have validated the efficacy of our proposed approach. Zhigang Ma, Xiaojun Chang, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Dynamic Concept Composition for Zero-Example Event DetectionabstractIn this paper, we focus on automatically detecting events in unconstrained videos without the use of any visual training exemplars. In principle, zero-shot learning makes it possible to train an event detection model based on the assumption that events (e.g. birthday party) can be described by multiple mid-level semantic concepts (e.g. ``blowing candle'', ``birthday cake''). Towards this goal, we first pre-train a bundle of concept classifiers using data from other sources. Then we evaluate the semantic correlation of each concept w.r.t. the event of interest and pick up the relevant concept classifiers, which are applied on all test videos to get multiple prediction score vectors. While most existing systems combine the predictions of the concept classifiers with fixed weights, we propose to learn the optimal weights of the concept classifiers for each testing video by exploring a set of online available videos with free-form text descriptions of their content. To validate the effectiveness of the proposed approach, we have conducted extensive experiments on the latest TRECVID MEDTest 2014, MEDTest 2013 and CCV dataset. The experimental results confirm the superiority of the proposed approach. Xiaojun Chang, Yi Yang 0001, Guodong Long, Chengqi Zhang, Alex Hauptmann 0001 |
AAAI | 5 |
| 2016 | Concepts Not Alone: Exploring Pairwise Relationships for Zero-Shot Video Activity RecognitionabstractVast quantities of videos are now being captured at astonishing rates, but the majority of these are not labelled. To cope with such data, we consider the task of content-based activity recognition in videos without any manually labelled examples, also known as zero-shot video recognition. To achieve this, videos are represented in terms of detected visual concepts, which are then scored as relevant or irrelevant according to their similarity with a given textual query. In this paper, we propose a more robust approach for scoring concepts in order to alleviate many of the brittleness and low precision problems of previous work. Not only do we jointly consider semantic relatedness, visual reliability, and discriminative power. To handle noise and non-linearities in the ranking scores of the selected concepts, we propose a novel pairwise order matrix approach for score aggregation. Extensive experiments on the large-scale TRECVID Multimedia Event Detection data show the superiority of our approach. Chuang Gan 0001, Ming Lin 0002, Yi Yang 0001, Gerard de Melo, Alex Hauptmann 0001 |
AAAI | 5 |
| 2016 | The Solution Path Algorithm for Identity-Aware Multi-object TrackingabstractWe propose an identity-aware multi-object tracker based on the solution path algorithm. Our tracker not only produces identity-coherent trajectories based on cues such as face recognition, but also has the ability to pinpoint potential tracking errors. The tracker is formulated as a quadratic optimization problem with ℓ0norm constraints, which we propose to solve with the solution path algorithm. The algorithm successively solves the same optimization problem but under different ℓpnorm constraints, where p gradually decreases from 1 to 0. Inspired by the success of the solution path algorithm in various machine learning tasks, this strategy is expected to converge to a better local minimum than directly minimizing the hardly solvable ℓ0norm or the roughly approximated ℓ1norm constraints. Furthermore, the acquired solution path complies with the "decision making process" of the tracker, which provides more insight to locating potential tracking errors. Experiments show that not only is our proposed tracker effective, but also the solution path enables automatic pinpointing of potential tracking failures, which can be readily utilized in an active learning framework to improve identity-aware multi-object tracking. Shoou-I Yu, Deyu Meng, Wangmeng Zuo, Alex Hauptmann 0001 |
CVPR | 4 |
| 2016 | Learning to Detect Concepts from Webly-Labeled Video Data
Junwei Liang 0001, Lu Jiang 0004, Deyu Meng, Alex Hauptmann 0001 |
IJCAI | 4 |
| 2016 | Avoiding Optimal Mean Robust PCA/2DPCA with Non-greedy ℓ1-Norm Maximization
Minnan Luo, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001 |
IJCAI | 5 |
| 2016 | Describing Videos using Multi-modal FusionabstractDescribing videos with natural language is one of the ultimate goals of video understanding. Video records multi-modal information including image, motion, aural, speech and so on. MSR Video to Language Challenge provides a good chance to study multi-modality fusion in caption task. In this paper, we propose the multi-modal fusion encoder and integrate it with text sequence decoder into an end-to-end video caption framework. Features from visual, aural, speech and meta modalities are fused together to represent the video contents. Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) are then used as the decoder to generate natural language sentences. Experimental results show the effectiveness of multi-modal fusion encoder trained in the end-to-end framework, which achieved top performance in both common metrics evaluation and human evaluation. Qin Jin, Jia Chen 0001, Shizhe Chen, Yifan Xiong 0001, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2016 | Which Information Sources are More Effective and Reliable in Video SearchabstractIt is common that users are interested in finding video segments, which contain further information about the video contents in a segment of interest. To facilitate users to find and browse related video contents, video hyperlinking aims at constructing links among video segments with relevant information in a large video collection. In this study, we explore the effectiveness of various video features on the performance of video hyperlinking, including subtitle, metadata, content features (i.e., audio and visual), surrounding context, as well as the combinations of those features. Besides, we also test different search strategies over different types of queries, which are categorized according to their video contents. Comprehensive experimental studies have been conducted on the dataset of TRECVID 2015 video hyperlinking task. Results show that (1) text features play a crucial role in search performance, and the combination of audio and visual features cannot provide improvements; (2) the consideration of contexts cannot obtain better results; and (3) due to the lack of training examples, machine learning techniques cannot improve the performance. Zhiyong Cheng 0001, Xuanchong Li, Jialie Shen 0001, Alex Hauptmann 0001 |
SIGIR | 4 |
| 2016 | InfAR dataset: Infrared action recognition at different times
Chenqiang Gao, Yinhe Du, Jiang Liu 0011, Jing Lv, Luyu Yang, Deyu Meng, Alex Hauptmann 0001 |
Neurocomputing | 7 |
| 2016 | Smart computing for large scale visual data sensing and processing
Lei Zhang 0006, Pinar Duygulu, Wangmeng Zuo, Shiguang Shan, Alex Hauptmann 0001 |
Neurocomputing | 5 |
| 2016 | Dictionary pruning with visual word significance for medical image retrieval
Fan Zhang 0013, Yang Song 0001, Tom Weidong Cai, Alex Hauptmann 0001, Sidong Liu, Sonia Pujol, Ron Kikinis, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 4 |
| 2015 | Exploring Semantic Inter-Class Relationships (SIR) for Zero-Shot Action RecognitionabstractAutomatically recognizing a large number of action categories from videos is of significant importance for video understanding. Most existing works focused on the design of more discriminative feature representation, and have achieved promising results when the positive samples are enough. However, very limited efforts were spent on recognizing a novel action without any positive exemplars, which is often the case in the real settings due to the large amount of action classes and the users' queries dramatic variations. To address this issue, we propose to perform action recognition when no positive exemplars of that class are provided, which is often known as the zero-shot learning. Different from other zero-shot learning approaches, which exploit attributes as the intermediate layer for the knowledge transfer, our main contribution is SIR, which directly leverages the semantic inter-class relationships between the known and unknown actions followed by label transfer learning. The inter-class semantic relationships are automatically measured by continuous word vectors, which learned by the skip-gram model using the large-scale text corpus. Extensive experiments on the UCF101 dataset validate the superiority of our method over fully-supervised approaches using few positive exemplars. Chuang Gan 0001, Ming Lin 0002, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
AAAI | 5 |
| 2015 | Self-Paced Curriculum LearningabstractCurriculum learning (CL) or self-paced learning (SPL) represents a recently proposed learning regime inspired by the learning process of humans and animals that gradually proceeds from easy to more complex samples in training. The two methods share a similar conceptual learning paradigm, but differ in specific learning schemes. In CL, the curriculum is predetermined by prior knowledge, and remain fixed thereafter. Therefore, this type of method heavily relies on the quality of prior knowledge while ignoring feedback about the learner. In SPL, the curriculum is dynamically determined to adjust to the learning pace of the leaner. However, SPL is unable to deal with prior knowledge, rendering it prone to overfitting. In this paper, we discover the missing link between CL and SPL, and propose a unified framework named self-paced curriculum leaning (SPCL). SPCL is formulated as a concise optimization problem that takes into account both prior knowledge known before training and the learning progress during training. In comparison to human education, SPCL is analogous to "instructor-student-collaborative" learning mode, as opposed to "instructor-driven" in CL or "student-driven" in SPL. Empirically, we show that the advantage of SPCL on two tasks. Lu Jiang 0004, Deyu Meng, Qian Zhao 0002, Shiguang Shan, Alex Hauptmann 0001 |
AAAI | 5 |
| 2015 | Complex Event Detection via Event Oriented Dictionary LearningabstractComplex event detection is a retrieval task with the goal of finding videos of a particular event in a large-scale unconstrained internet video archive, given example videos and text descriptions. Nowadays, different multimodal fusion schemes of low-level and high-level features are extensively investigated and evaluated for the complex event detection task. However, how to effectively select the high-level semantic meaningful concepts from a large pool to assist complex event detection is rarely studied in the literature. In this paper, we propose two novel strategies to automatically select semantic meaningful concepts for the event detection task based on both the events-kit text descriptions and the concepts high-level feature descriptions. Moreover, we introduce a novel event oriented dictionary representation based on the selected semantic concepts. Towards this goal, we leverage training samples of selected concepts from the Semantic Indexing (SIN) dataset with a pool of 346 concepts, into a novel supervised multi-task dictionary learning framework. Extensive experimental results on TRECVID Multimedia Event Detection (MED) dataset demonstrate the efficacy of our proposed method. Yan Yan 0002, Yi Yang 0001, Haoquan Shen, Deyu Meng, Gaowen Liu, Alex Hauptmann 0001, Nicu Sebe |
AAAI | 6 |
| 2015 | Self-Paced Learning for Matrix FactorizationabstractMatrix factorization (MF) has been attracting much attention due to its wide applications. However, since MF models are generally non-convex, most of the existing methods are easily stuck into bad local minima, especially in the presence of outliers and missing data. To alleviate this deficiency, in this study we present a new MF learning methodology by gradually including matrix elements into MF training from easy to complex. This corresponds to a recently proposed learning fashion called self-paced learning (SPL), which has been demonstrated to be beneficial in avoiding bad local minima. We also generalize the conventional binary (hard) weighting scheme for SPL to a more effective real-valued (soft) weighting manner. The effectiveness of the proposed self-paced MF method is substantiated by a series of experiments on synthetic, structure from motion and background subtraction data. Qian Zhao 0002, Deyu Meng, Lu Jiang 0004, Qi Xie 0002, Zongben Xu, Alex Hauptmann 0001 |
AAAI | 6 |
| 2015 | Massive Open Online Proctor: Protecting the Credibility of MOOCs certificatesabstractMassive Open Online Courses (MOOCs) enable everyone to receive high-quality education. However, current MOOC creators cannot provide an effective, economical, and scalable method to detect cheating on tests, which would be required for any certification. In this paper, we propose a Massive Open Online Proctoring (MOOP) framework, which combines both automatic and collaborative approaches to detect cheating behaviors in online tests. The MOOP framework consists of three major components: Automatic Cheating Detector (ACD), Peer Cheating Detector (PCD), and Final Review Committee (FRC). ACD uses webcam video or other sensors to monitor students and automatically flag suspected cheating behavior. Ambiguous cases are then sent to the PCD, where students peer-review flagged webcam video to confirm suspicious cheating behaviors. Finally, the list of suspicious cheating behaviors is sent to the FRC to make the final punishing decision. Our experiment show that ACD and PCD can detect usage of a cheat sheet with good accuracy and can reduce the overall human resources required to monitor MOOCs for cheating. Xuanchong Li, Kai-min Chang, Yueran Yuan, Alex Hauptmann 0001 |
CSCW | 4 |
| 2015 | DevNet: A Deep Event Network for multimedia event detection and evidence recountingabstractIn this paper, we focus on complex event detection in internet videos while also providing the key evidences of the detection results. Convolutional Neural Networks (CNNs) have achieved promising performance in image classification and action recognition tasks. However, it remains an open problem how to use CNNs for video event detection and recounting, mainly due to the complexity and diversity of video events. In this work, we propose a flexible deep CNN infrastructure, namely Deep Event Network (DevNet), that simultaneously detects pre-defined events and provides key spatial-temporal evidences. Taking key frames of videos as input, we first detect the event of interest at the video level by aggregating the CNN features of the key frames. The pieces of evidences which recount the detection results, are also automatically localized, both temporally and spatially. The challenge is that we only have video level labels, while the key evidences usually take place at the frame levels. Based on the intrinsic property of CNNs, we first generate a spatial-temporal saliency map by back passing through DevNet, which then can be used to find the key frames which are most indicative to the event, as well as to localize the specific spatial position, usually an object, in the frame of the highly indicative area. Experiments on the large scale TRECVID 2014 MEDTest dataset demonstrate the promising performance of our method, both for event detection and evidence recounting. Chuang Gan 0001, Naiyan Wang, Yi Yang 0001, Dit-Yan Yeung, Alex Hauptmann 0001 |
CVPR | 5 |
| 2015 | Beyond Gaussian Pyramid: Multi-skip Feature Stacking for action recognitionabstractMost state-of-the-art action feature extractors involve differential operators, which act as highpass filters and tend to attenuate low frequency action information. This attenuation introduces bias to the resulting features and generates ill-conditioned feature matrices. The Gaussian Pyramid has been used as a feature enhancing technique that encodes scale-invariant characteristics into the feature space in an attempt to deal with this attenuation. However, at the core of the Gaussian Pyramid is a convolutional smoothing operation, which makes it incapable of generating new features at coarse scales. In order to address this problem, we propose a novel feature enhancing technique called Multi-skIp Feature Stacking (MIFS), which stacks features extracted using a family of differential filters parameterized with multiple time skips and encodes shift-invariance into the frequency space. MIFS compensates for information lost from using differential operators by recapturing information at coarse scales. This recaptured information allows us to match actions at different speeds and ranges of motion. We prove that MIFS enhances the learnability of differential-based features exponentially. The resulting feature matrices from MIFS have much smaller conditional numbers and variances than those from conventional methods. Experimental results show significantly improved performance on challenging action recognition and event detection tasks. Specifically, our method exceeds the state-of-the-arts on Hollywood2, UCF101 and UCF50 datasets and is comparable to state-of-the-arts on HMDB51 and Olympics Sports datasets. MIFS can also be used as a speedup strategy for feature extraction with minimal or no accuracy cost. Zhen-Zhong Lan, Ming Lin 0002, Xuanchong Li, Alex Hauptmann 0001, Bhiksha Raj |
CVPR | 4 |
| 2015 | A discriminative CNN video representation for event detectionabstractIn this paper, we propose a discriminative video representation for event detection over a large scale video dataset when only limited hardware resources are available. The focus of this paper is to effectively leverage deep Convolutional Neural Networks (CNNs) to advance event detection, where only frame level static descriptors can be extracted by the existing CNN toolkits. This paper makes two contributions to the inference of CNN video representation. First, while average pooling and max pooling have long been the standard approaches to aggregating frame level static features, we show that performance can be significantly improved by taking advantage of an appropriate encoding method. Second, we propose using a set of latent concept descriptors as the frame descriptor, which enriches visual information while keeping it computationally affordable. The integration of the two contributions results in a new state-of-the-art performance in event detection over the largest video datasets. Compared to improved Dense Trajectories, which has been recognized as the best video representation for event detection, our new representation improves the Mean Average Precision (mAP) from 27.6% to 36.8% for the TRECVID MEDTest 14 dataset and from 34.0% to 44.6% for the TRECVID MEDTest 13 dataset. Zhongwen Xu, Yi Yang 0001, Alex Hauptmann 0001 |
CVPR | 3 |
| 2015 | Semantic Concept Discovery for Large-Scale Zero-Shot Event Detection
Xiaojun Chang, Yi Yang 0001, Alex Hauptmann 0001, Eric P. Xing, Yaoliang Yu |
IJCAI | 3 |
| 2015 | Density Corrected Sparse Recovery when R.I.P. Condition Is Broken
Ming Lin 0002, Zhen-Zhong Lan, Alex Hauptmann 0001 |
IJCAI | 3 |
| 2015 | Bridging the Ultimate Semantic Gap: A Semantic Search Engine for Internet VideosabstractSemantic search in video is a novel and challenging problem in information and multimedia retrieval. Existing solutions are mainly limited to text matching, in which the query words are matched against the textual metadata generated by users. This paper presents a state-of-the-art system for event search without any textual metadata or example videos. The system relies on substantial video content understanding and allows for semantic search over a large collection of videos. The novelty and practicality is demonstrated by the evaluation in NIST TRECVID 2014, where the proposed system achieves the best performance. We share our observations and lessons in building such a state-of-the-art system, which may be instrumental in guiding the design of the future system for semantic search in video. Lu Jiang 0004, Shoou-I Yu, Deyu Meng, Teruko Mitamura, Alex Hauptmann 0001 |
ICMR | 5 |
| 2015 | Incremental Multimodal Query Construction for Video SearchabstractRecent improvements in content-based video search have led to systems with promising accuracy, thus opening up the possibility for interactive content-based video search to the general public. We present an interactive system based on a state-of-the-art content-based video search pipeline which enables users to do multimodal text-to-video and video-to-video search in large video collections, and to incrementally refine queries through relevance feedback and model visualization. Also, the comprehensive functionalities enhance a flexible formulation of multimodal queries with different characteristics. Quantitative and qualitative analysis shows that our system is capable of assisting users to incrementally build effective queries over complex event topics. Xiaojun Chang, Shoou-I Yu, Xingzhong Du, Xuanchong Li, Lu Jiang 0004, Zexi Mao, Zhen-Zhong Lan, Susanne Burger, Alex Hauptmann 0001 |
ICMR | 11 |
| 2015 | Content-Based Video Search over 1 Million Videos with 1 Core in 1 SecondabstractMany content-based video search (CBVS) systems have been proposed to analyze the rapidly-increasing amount of user-generated videos on the Internet. Though the accuracy of CBVS systems have drastically improved, these high accuracy systems tend to be too inefficient for interactive search. Therefore, to strive for real-time web-scale CBVS, we perform a comprehensive study on the different components in a CBVS system to understand the trade-offs between accuracy and speed of each component. Directions investigated include exploring different low-level and semantics-based features, testing different compression factors and approximations during video search, and understanding the time v.s. accuracy trade-off of reranking. Extensive experiments on data sets consisting of more than 1,000 hours of video showed that through a combination of effective features, highly compressed representations, and one iteration of reranking, our proposed system can achieve an 10,000-fold speedup while retaining 80% accuracy of a state-of-the-art CBVS system. We further performed search over 1 million videos and demonstrated that our system can complete the search in 0.975 seconds with a single core, which potentially opens the door to interactive web-scale CBVS for the general public. Shoou-I Yu, Lu Jiang 0004, Zhongwen Xu, Yi Yang 0001, Alex Hauptmann 0001 |
ICMR | 5 |
| 2015 | Searching Persuasively: Joint Event Detection and Evidence Recounting with Limited SupervisionabstractMultimedia event detection (MED) and multimedia event recounting (MER) are fundamental tasks in managing large amounts of unconstrained web videos, and have attracted a lot of attention in recent years. Most existing systems perform MER as a post-processing step on top of the MED results. In order to leverage the mutual benefits of the two tasks, we propose a joint framework that simultaneously detects high-level events and localizes the indicative concepts of the events. Our premise is that a good recounting algorithm should not only explain the detection result, but should also be able to assist detection in the first place. Coupled in a joint optimization framework, recounting improves detection by pruning irrelevant noisy concepts while detection directs recounting to the most discriminative evidences. To better utilize the powerful and interpretable semantic video representation, we segment each video into several shots and exploit the rich temporal structures at shot level. The consequent computational challenge is carefully addressed through a significant improvement of the current ADMM algorithm, which, after eliminating all inner loops and equipping novel closed-form solutions for all intermediate steps, enables us to efficiently process extremely large video corpora. We test the proposed method on the large scale TRECVID MEDTest 2014 and MEDTest 2013 datasets, and obtain very promising results for both MED and MER. Xiaojun Chang, Yaoliang Yu, Yi Yang 0001, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2015 | Image Profiling for History Events on the FlyabstractHistory event related knowledge is precious and imagery is a powerful medium that records diverse information about the event. In this paper, we propose to automatically construct an image profile given a one sentence description of the historic event which contains where, when, who and what elements. Such a simple input requirement makes our solution easy to scale up and support a wide range of culture preservation and curation related applications ranging from wikipedia enrichment to history education. However, history relevant information on the web is available as "wild and dirty" data, which is quite different from clean, manually curated and structured information sources. There are two major challenges to build our proposed image profiles: 1) unconstrained image genre diversity. We categorize images into genres of documents/maps, paintings or photos. Image genre classification involves a full-spectrum of features from low-level color to high-level semantic concepts. 2) image content diversity. It can include faces, objects and scenes. Furthermore, even within the same event, the views and subjects of images are diverse and correspond to different facets of the event. To solve this challenge, we group images at two levels of granularity: iconic image grouping and facet image grouping. These require different types of features and analysis from near exact matching to soft semantic similarity. We develop a full-range feature analysis module which is composed of several levels, each suitable for different types of image analysis tasks. The wide range of features are based on both classical hand-crafted features and different layers of a convolutional neural network. We compare and study the performance of the different levels in the full-range features and show their effectiveness on handling such a wild, unconstrained dataset. Jia Chen 0001, Qin Jin, Yong Yu 0001, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2015 | Fast and Accurate Content-based Semantic Search in 100M Internet VideosabstractLarge-scale content-based semantic search in video is an interesting and fundamental problem in multimedia analysis and retrieval. Existing methods index a video by the raw concept detection score that is dense and inconsistent, and thus cannot scale to "big data" that are readily available on the Internet. This paper proposes a scalable solution. The key is a novel step called concept adjustment that represents a video by a few salient and consistent concepts that can be efficiently indexed by the modified inverted index. The proposed adjustment model relies on a concise optimization framework with interpretations. The proposed index leverages the text-based inverted index for video retrieval. Experimental results validate the efficacy and the efficiency of the proposed method. The results show that our method can scale up the semantic search while maintaining state-of-the-art search performance. Specifically, the proposed method (with reranking) achieves the best result on the challenging TRECVID Multimedia Event Detection (MED) zero-example task. It only takes 0.2 second on a single CPU core to search a collection of 100 million Internet videos. Lu Jiang 0004, Shoou-I Yu, Deyu Meng, Yi Yang 0001, Teruko Mitamura, Alex Hauptmann 0001 |
ACM Multimedia | 6 |
| 2015 | Multi-Class Active Learning by Uncertainty Sampling with Diversity Maximization
Yi Yang 0001, Zhigang Ma, Feiping Nie 0001, Xiaojun Chang, Alex Hauptmann 0001 |
Int. J. Comput. Vis. | 5 |
| 2015 | Multi-view discriminative and structured dictionary learning with group sparsity for human action recognition
Zan Gao 0002, Hua Zhang 0003, Guangping Xu, Yanbin Xue, Alex Hauptmann 0001 |
Signal Process. | 5 |
| 2015 | Event Oriented Dictionary Learning for Complex Event DetectionabstractComplex event detection is a retrieval task with the goal of finding videos of a particular event in a large-scale unconstrained Internet video archive, given example videos and text descriptions. Nowadays, different multimodal fusion schemes of low-level and high-level features are extensively investigated and evaluated for the complex event detection task. However, how to effectively select the high-level semantic meaningful concepts from a large pool to assist complex event detection is rarely studied in the literature. In this paper, we propose a novel strategy to automatically select semantic meaningful concepts for the event detection task based on both the events-kit text descriptions and the concepts high-level feature descriptions. Moreover, we introduce a novel event oriented dictionary representation based on the selected semantic concepts. Toward this goal, we leverage training images (frames) of selected concepts from the semantic indexing dataset with a pool of 346 concepts, into a novel supervised multitask lp -norm dictionary learning framework. Extensive experimental results on TRECVID multimedia event detection dataset demonstrate the efficacy of our proposed method. Yan Yan 0002, Yi Yang 0001, Deyu Meng, Gaowen Liu, Alex Hauptmann 0001, Nicu Sebe |
IEEE Trans. Image Process. | 6 |
| 2014 | Event Detection Using Multi-level Relevance Labels and Multiple FeaturesabstractWe address the challenging problem of utilizing related exemplars for complex event detection while multiple features are available. Related exemplars share certain positive elements of the event, but have no uniform pattern due to the huge variance of relevance levels among different related exemplars. None of the existing multiple feature fusion methods can deal with the related exemplars. In this paper, we propose an algorithm which adaptively utilizes the related exemplars by cross-feature learning. Ordinal labels are used to represent the multiple relevance levels of the related videos. Label candidates of related exemplars are generated by exploring the possible relevance levels of each related exemplar via a cross-feature voting strategy. Maximum margin criterion is then applied in our framework to discriminate the positive and negative exemplars, as well as the related exemplars from different relevance levels. We test our algorithm using the large scale TRECVID 2011 dataset and it gains promising performance. Zhongwen Xu, Ivor W. Tsang, Yi Yang 0001, Zhigang Ma, Alex Hauptmann 0001 |
CVPR | 5 |
| 2014 | Unsupervised Video Adaptation for Parsing Human Motion
Haoquan Shen, Shoou-I Yu, Yi Yang 0001, Deyu Meng, Alex Hauptmann 0001 |
ECCV (5) | 5 |
| 2014 | Interactive Surveillance Event Detection through Mid-level Discriminative RepresentationabstractEvent detection from real surveillance videos with complicated background environment is always a very hard task. Different from the traditional retrospective and interactive systems designed on this task, which are mainly executed on video fragments located within the event-occurrence time, in this paper we propose a new interactive system constructed on the mid-level discriminative representations (patches/shots) which are closely related to the event (might occur beyond the event-occurrence period) and are easier to be detected than video fragments. By virtue of such easily-distinguished mid-level patterns, our framework realizes an effective labor division between computers and human participants. The task of computers is to train classifiers on a bunch of mid-level discriminative representations, and to sort all the possible mid-level representations in the evaluation sets based on the classifier scores. The task of human participants is then to readily search the events based on the clues offered by these sorted mid-level representations. For computers, such mid-level representations, with more concise and consistent patterns, can be more accurately detected than video fragments utilized in the conventional framework, and on the other hand, a human participant can always much more easily search the events of interest implicated by these location-anchored mid-level representations than conventional video fragments containing entire scenes. Both of these two properties facilitate the availability of our framework in real surveillance event detection applications. Chenqiang Gao, Deyu Meng, Yi Yang 0001, Yang Cai 0002, Haoquan Shen, Gaowen Liu, Alex Hauptmann 0001 |
ICMR | 9 |
| 2014 | Zero-Example Event Search using MultiModal Pseudo Relevance FeedbackabstractWe propose a novel method MultiModal Pseudo Relevance Feedback (MMPRF) for event search in video, which requires no search examples from the user. Pseudo Relevance Feedback has shown great potential in retrieval tasks, but previous works are limited to unimodal tasks with only a single ranked list. To tackle the event search task which is inherently multimodal, our proposed MMPRF takes advantage of multiple modalities and multiple ranked lists to enhance event search performance in a principled way. The approach is unique in that it leverages not only semantic features, but also non-semantic low-level features for event search in the absence of training data. Evaluated on the TRECVID MEDTest dataset, the approach improves the baseline by up to 158% in terms of the mean average precision. It also significantly contributes to CMU Team's final submission in TRECVID-13 Multimedia Event Detection. Lu Jiang 0004, Teruko Mitamura, Shoou-I Yu, Alex Hauptmann 0001 |
ICMR | 4 |
| 2014 | Viral Video Style: A Closer Look at Viral Videos on YouTubeabstractViral videos that gain popularity through the process of Internet sharing are having a profound impact on society. Existing studies on viral videos have only been on small or confidential datasets. We collect by far the largest open benchmark for viral video study called CMU Viral Video Dataset, and share it with researchers from both academia and industry. Having verified existing observations on the dataset, we discover some interesting characteristics of viral videos. Based on our analysis, in the second half of the paper, we propose a model to forecast the future peak day of viral videos. The application of our work is not only important for advertising agencies to plan advertising campaigns and estimate costs, but also for companies to be able to quickly respond to rivals in viral marketing campaigns. The proposed method is unique in that it is the first attempt to incorporate video metadata into the peak day prediction. The empirical results demonstrate that the proposed method outperforms the state-of-the-art methods, with statistically significant differences. Lu Jiang 0004, Yajie Miao, Yi Yang 0001, Zhen-Zhong Lan, Alex Hauptmann 0001 |
ICMR | 5 |
| 2014 | Towards Efficient Learning of Optimal Spatial Bag-of-Words RepresentationsabstractSpatial Pyramid Matching (SPM) assumes that the spatial Bag-of-Words (BoW) representation is independent of data. However, evidence has shown that the assumption usually leads to a suboptimal representation. In this paper, we propose a novel method called Jensen-Shannon (JS) Tiling to learn the BoW representation from data directly at the BoW level. The proposed JS Tiling is especially appropriate for large-scale datasets as it is orders of magnitude faster than existing methods, but with comparable or even better classification precision. Experimental results on four benchmarks including two TRECVID12 datasets validate that JS Tiling outperforms the SPM and the state-of-the-art methods. The runtime comparison demonstrates that selecting BoW representations by JS Tiling is more than 1,000 times faster than running classifiers. Besides, JS Tiling is an important component contributing to CMU Teams' final submission in TRECVID 2012 Multimedia Event Detection. Lu Jiang 0004, Deyu Meng, Alex Hauptmann 0001 |
ICMR | 4 |
| 2014 | The Mystery of Faces: Investigating Face Contribution for Multimedia Event DetectionabstractMultimedia event detection (MED) is a retrieval task with the goal of finding videos of a particular event in a large scale internet video archive, given example videos and text descriptions. Nowadays, different multimodal fusion schemes of low-level and high-level features are extensively investigated and evaluated for MED. For most of events in MED, people are usually the central subjects in videos. The face of a person can be considered as the most important factor which brings a lot of information describing the video events. However, face information has not been systematically investigated in the previous research for MED. In this paper, we investigate the possibility of using the high-level face information to assist multimedia event detection. Moreover, since the labeled data in TRECVID MED dataset are limited, we propose a semi-supervised kernel ridge regression which works well in practice to explore the useful information from unlabeled data to assist the event detection. Extensive experimental results on TRECVID MED dataset show that our proposed method outperforms the state-of-the-art methods by up to 4%. Gaowen Liu, Yan Yan 0002, Chenqiang Gao, Alex Hauptmann 0001, Nicu Sebe |
ICMR | 5 |
| 2014 | Easy Samples First: Self-paced Reranking for Zero-Example Multimedia SearchabstractReranking has been a focal technique in multimedia retrieval due to its efficacy in improving initial retrieval results. Current reranking methods, however, mainly rely on the heuristic weighting. In this paper, we propose a novel reranking approach called Self-Paced Reranking (SPaR) for multimodal data. As its name suggests, SPaR utilizes samples from easy to more complex ones in a self-paced fashion. SPaR is special in that it has a concise mathematical objective to optimize and useful properties that can be theoretically verified. It on one hand offers a unified framework providing theoretical justifications for current reranking methods, and on the other hand generates a spectrum of new reranking schemes. This paper also advances the state-of-the-art self-paced learning research which potentially benefits applications in other fields. Experimental results validate the efficacy and the efficiency of the proposed method on both image and video search tasks. Notably, SPaR achieves by far the best result on the challenging TRECVID multimedia event search task. Lu Jiang 0004, Deyu Meng, Teruko Mitamura, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2014 | Multiple Features But Few Labels?: A Symbiotic Solution Exemplified for Video AnalysisabstractVideo analysis has been attracting increasing research due to the proliferation of internet videos. In this paper, we investigate how to improve the performance on internet quality video analysis. Particularly, we work on the scenario of few labeled training videos being provided, which is less focused in multimedia. To being with, we consider how to more effectively harness the evidences from the low-level features. Researchers have developed several promising features to represent videos to capture the semantic information. However, as videos usually characterize rich semantic contents, the analysis performance by using one single feature is potentially limited. Simply combining multiple features through early fusion or late fusion to incorporate more informative cues is doable but not optimal due to the heterogeneity and different predicting capability of these features. For better exploitation of multiple features, we propose to mine the importance of different features and cast it into the learning of the classification model. Our method is based on multiple graphs from different features and uses the Riemannian metric to evaluate the feature importance. On the other hand, to be able to use limited labeled training videos for a respectable accuracy we formulate our method in a semi-supervised way. The main contribution of this paper is a novel scheme of evaluating the feature importance that is further casted into a unified framework of harnessing multiple weighted features with limited labeled training videos. We perform extensive experiments on video action recognition and multimedia event recognition and the comparison to other state-of-the-art multi-feature learning algorithms has validated the efficacy of our framework. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2014 | Instructional Videos for Unsupervised Harvesting and Learning of Action ExamplesabstractOnline instructional videos have become a popular way for people to learn new skills encompassing art, cooking and sports. As watching instructional videos is a natural way for humans to learn, analogously, machines can also gain knowledge from these videos. We propose to utilize the large amount of instructional videos available online to harvest examples of various actions in an unsupervised fashion. The key observation is that in instructional videos, the instructor's action is highly correlated with the instructor's narration. By leveraging this correlation, we can exploit the timing of action corresponding terms in the speech transcript to temporally localize actions in the video and harvest action examples. The proposed method is scalable as it requires no human intervention. Experiments show that the examples harvested are of reasonably good quality, and action detectors trained on data collected by our unsupervised method yields comparable performance with detectors trained with manually collected data on the TRECVID Multimedia Event Detection task. Shoou-I Yu, Lu Jiang 0004, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2014 | Resource Constrained Multimedia Event Detection
Zhen-Zhong Lan, Yi Yang 0001, Nicolas Ballas, Shoou-I Yu, Alex Hauptmann 0001 |
MMM (1) | 5 |
| 2014 | Self-Paced Learning with Diversity
Lu Jiang 0004, Deyu Meng, Shoou-I Yu, Zhen-Zhong Lan, Shiguang Shan, Alex Hauptmann 0001 |
NIPS | 6 |
| 2014 | Harnessing Lab Knowledge for Real-World Action Recognition
Zhigang Ma, Yi Yang 0001, Feiping Nie 0001, Nicu Sebe, Shuicheng Yan, Alex Hauptmann 0001 |
Int. J. Comput. Vis. | 6 |
| 2014 | Enhanced and hierarchical structure algorithm for data imbalance problem in semantic extraction under massive video dataset
Zan Gao 0002, Ming-yu Chen 0001, Alex Hauptmann 0001, Hua Zhang 0003, Anni Cai |
Multim. Tools Appl. | 4 |
| 2014 | Multimedia classification and event detection using double fusion
Zhen-Zhong Lan, Shoou-I Yu, Wei Liu 0015, Alex Hauptmann 0001 |
Multim. Tools Appl. | 5 |
| 2014 | E-LAMP: integration of innovative ideas for multimedia event detection
Yi Yang 0001, Lu Jiang 0004, Shoou-I Yu, Zhen-Zhong Lan, Zhigang Ma, Waito Sze, Ehsan Younessian, Alex Hauptmann 0001 |
Mach. Vis. Appl. | 9 |
| 2014 | Knowledge Adaptation with PartiallyShared Features for Event DetectionUsing Few ExemplarsabstractMultimedia event detection (MED) is an emerging area of research. Previous work mainly focuses on simple event detection in sports and news videos, or abnormality detection in surveillance videos. In contrast, we focus on detecting more complicated and generic events that gain more users' interest, and we explore an effective solution for MED. Moreover, our solution only uses few positive examples since precisely labeled multimedia content is scarce in the real world. As the information from these few positive examples is limited, we propose using knowledge adaptation to facilitate event detection. Different from the state of the art, our algorithm is able to adapt knowledge from another source for MED even if the features of the source and the target are partially different, but overlapping. Avoiding the requirement that the two domains are consistent in feature types is desirable as data collection platforms change or augment their capabilities and we should be able to respond to this with little or no effort. We perform extensive experiments on real-world multimedia archives consisting of several challenging events. The results show that our approach outperforms several other state-of-the-art detection algorithms. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Symbiotic Tracker Ensemble Toward A Unified Tracking FrameworkabstractTracking people and objects is a fundamental stage toward many video surveillance systems, for which various trackers have been specifically designed in the past decade. However, it comes to a consensus that there is not any specific tracker that works sufficiently well under all circumstances. Therefore, one potential solution is to deploy multiple trackers, with a tracker output fusion step to boost the overall performance. Subsequently, an intelligent fusion design, yet general and orthogonal to any specific tracker, plays a key role in successful tracking. In this paper, we propose a symbiotic tracker ensemble toward a unified tracking framework, which is based on only the output of each individual tracker, without knowing its specific mechanism. In our approach, all trackers run in parallel, without requiring any details for tracker running, which means that all trackers are treated as black boxes. The proposed symbiotic tracker ensemble framework aims at learning an optimal combination of these tracking results. Our method captures the relation among individual trackers robustly from two aspects. First, the consistency between two successive frames is calculated for each tracker. Then, the pair-wise correlation among different trackers is estimated in the new coming frame by a graph-propagation process. Experimental results on the Caremedia dataset and the Caviar dataset demonstrate the effectiveness of the proposed method, with comparisons to several state-of-the-art methods. Yue Gao 0002, Rongrong Ji, Alex Hauptmann 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Semi-Supervised Multiple Feature Analysis for Action RecognitionabstractThis paper presents a semi-supervised method for categorizing human actions using multiple visual features. The proposed algorithm simultaneously learns multiple features from a small number of labeled videos, and automatically utilizes data distributions between labeled and unlabeled data to boost the recognition performance. Shared structural analysis is applied in our approach to discover a common subspace shared by each type of feature. In the subspace, the proposed algorithm is able to characterize more discriminative information of each feature type. Additionally, data distribution information of each type of feature has been preserved. The aforementioned attributes make our algorithm robust for action recognition, especially when only limited labeled training samples are provided. Extensive experiments have been conducted on both the choreographed and the realistic video datasets, including KTH, Youtube action and UCF50. Experimental results show that our method outperforms several state-of-the-art algorithms. Most notably, much better performances have been achieved when there are only a few labeled training samples. Sen Wang 0001, Zhigang Ma, Yi Yang 0001, Xue Li 0001, Chaoyi Pang, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 6 |
| 2013 | Complex Event Detection via Multi-source Video AttributesabstractComplex events essentially include human, scenes, objects and actions that can be summarized by visual attributes, so leveraging relevant attributes properly could be helpful for event detection. Many works have exploited attributes at image level for various applications. However, attributes at image level are possibly insufficient for complex event detection in videos due to their limited capability in characterizing the dynamic properties of video data. Hence, we propose to leverage attributes at video level (named as video attributes in this work), i.e., the semantic labels of external videos are used as attributes. Compared to complex event videos, these external videos contain simple contents such as objects, scenes and actions which are the basic elements of complex events. Specifically, building upon a correlation vector which correlates the attributes and the complex event, we incorporate video attributes latently as extra informative cues into the event detector learnt from complex event videos. Extensive experiments on a real-world large-scale dataset validate the efficacy of the proposed approach. Zhigang Ma, Yi Yang 0001, Zhongwen Xu, Shuicheng Yan, Nicu Sebe, Alex Hauptmann 0001 |
CVPR | 6 |
| 2013 | Harry Potter's Marauder's Map: Localizing and Tracking Multiple Persons-of-Interest by Nonnegative DiscretizationabstractA device just like Harry Potter's Marauder's Map, which pinpoints the location of each person-of-interest at all times, provides invaluable information for analysis of surveillance videos. To make this device real, a system would be required to perform robust person localization and tracking in real world surveillance scenarios, especially for complex indoor environments with many walls causing occlusion and long corridors with sparse surveillance camera coverage. We propose a tracking-by-detection approach with nonnegative discretization to tackle this problem. Given a set of person detection outputs, our framework takes advantage of all important cues such as color, person detection, face recognition and non-background information to perform tracking. Local learning approaches are used to uncover the manifold structure in the appearance space with spatio-temporal constraints. Nonnegative discretization is used to enforce the mutual exclusion constraint, which guarantees a person detection output to only belong to exactly one individual. Experiments show that our algorithm performs robust localization and tracking of persons-of-interest not only in outdoor scenes, but also in a complex indoor real-world nursing home environment. Shoou-I Yu, Yi Yang 0001, Alex Hauptmann 0001 |
CVPR | 3 |
| 2013 | Space-Time Robust Representation for Action RecognitionabstractWe address the problem of action recognition in unconstrained videos. We propose a novel content driven pooling that leverages space-time context while being robust toward global space-time transformations. Being robust to such transformations is of primary importance in unconstrained videos where the action localizations can drastically shift between frames. Our pooling identifies regions of interest using video structural cues estimated by differ ent saliency functions. To combine the different structural information, we introduce an iterative structure learning algorithm, WSVM (weighted SVM), that determines the optimal saliency layout of an action model through a sparse regularizer. A new optimization method is proposed to solve the WSVM' highly non-smooth objective function. We evaluate our approach on standard action datasets (KTH, UCF50 and HMDB). Most noticeably, the accuracy of our algorithm reaches 51.8% on the challenging HMDB dataset which outperforms the state-of-the-art of 7.3% relatively. Nicolas Ballas, Yi Yang 0001, Zhen-Zhong Lan, Bertrand Delezoide, Françoise J. Prêteux, Alex Hauptmann 0001 |
ICCV | 6 |
| 2013 | Feature Weighting via Optimal Thresholding for Video AnalysisabstractFusion of multiple features can boost the performance of large-scale visual classification and detection tasks like TRECVID Multimedia Event Detection (MED) competition [1]. In this paper, we propose a novel feature fusion approach, namely Feature Weighting via Optimal Thresholding (FWOT) to effectively fuse various features. FWOT learns the weights, thresholding and smoothing parameters in a joint framework to combine the decision values obtained from all the individual features and the early fusion. To the best of our knowledge, this is the first work to consider the weight and threshold factors of fusion problem simultaneously. Compared to state-of-the-art fusion algorithms, our approach achieves promising improvements on HMDB [8] action recognition dataset and CCV [5] video classification dataset. In addition, experiments on two TRECVID MED 2011 collections show that our approach outperforms the state-of-the-art fusion methods for complex event detection. Zhongwen Xu, Yi Yang 0001, Ivor W. Tsang, Nicu Sebe, Alex Hauptmann 0001 |
ICCV | 5 |
| 2013 | How Related Exemplars Help Complex Event Detection in Web Videos?abstractCompared to visual concepts such as actions, scenes and objects, complex event is a higher level abstraction of longer video sequences. For example, a "marriage proposal" event is described by multiple objects (e.g., ring, faces), scenes (e.g., in a restaurant, outdoor) and actions (e.g., kneeling down). The positive exemplars which exactly convey the precise semantic of an event are hard to obtain. It would be beneficial to utilize the related exemplars for complex event detection. However, the semantic correlations between related exemplars and the target event vary substantially as relatedness assessment is subjective. Two related exemplars can be about completely different events, e.g., in the TRECVID MED dataset, both bicycle riding and equestrianism are labeled as related to "attempting a bike trick" event. To tackle the subjectiveness of human assessment, our algorithm automatically evaluates how positive the related exemplars are for the detection of an event and uses them on an exemplar-specific basis. Experiments demonstrate that our algorithm is able to utilize related exemplars adaptively, and the algorithm gains good performance for complex event detection. Yi Yang 0001, Zhigang Ma, Zhongwen Xu, Shuicheng Yan, Alex Hauptmann 0001 |
ICCV | 5 |
| 2013 | ACM MM MIIRH 2013: workshop on multimedia indexing and information retrieval for healthcareabstractHealthcare systems are depending on increasingly sophisticated and ubiquitous technology, while telehealth is rapidly gaining importance with the advent of low-cost and effective technological solutions in medicine. The increase in the worldwide elderly population and the burden this is inflicting upon the workforce, societies and economies are making remote care and independent living at home a necessity. MIIRH is the first workshop on multimedia analysis for remote care of and assisted living solutions which enable people that are incapacitated in some regard to continue living independently at home and remain active members of society. The topics addressed in MIIRH are extremely timely, as multitudes of cost-effective and high quality care solutions are already being developed and used, rendering the examination of new medical, healthcare paradigms an absolute necessity. Jenny Benois-Pineau, Alexia Briassouli, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2013 | Spatio-temporal fisher vector coding for surveillance event detectionabstractWe present a generic event detection system evaluated in the Surveillance Event Detection (SED) task of TRECVID 2012. We investigate a statistical approach with spatio-temporal features applied to seven event classes, which were defined by the SED task. This approach is based on local spatio-temporal descriptors, called MoSIFT and generated by pair-wise video frames. A Gaussian Mixture Model(GMM) is learned to model the distribution of the low level features. Then for each sliding window, the Fisher vector encoding [improvedFV] is used to generate the sample representation. The model is learnt using a Linear SVM for each event. The main novelty of our system is the introduction of Fisher vector encoding into video event detection. Fisher vector encoding has demonstrated great success in image classification. The key idea is to model the low level visual features as a Gaussian Mixture Model and to generate an intermediate vector representation for bag of features. FV encoding uses higher order statistics in place of histograms in the standard BoW. FV has several good properties: (a) it can naturally separate the video specific information from the noisy local features and (b) we can use a linear model for this representation. We build an efficient implementation for FV encoding which can attain a 10 times speed-up over real-time. We also take advantage of non-trivial object localization techniques to feed into the video event detection, e.g. multi-scale detection and non-maximum suppression. This approach outperformed the results of all other teams submissions in TRECVID SED 2012 on four of the seven event types. Qiang Chen 0007, Yang Cai 0002, Lisa M. Brown, Ankur Datta, Quanfu Fan, Rogério Feris, Shuicheng Yan, Alex Hauptmann 0001, Sharath Pankanti |
ACM Multimedia | 8 |
| 2013 | We are not equally negative: fine-grained labeling for multimedia event detectionabstractMultimedia event detection (MED) is an effective technique for video indexing and retrieval. Current classifier training for MED treats the negative videos equally. However, many negative videos may resemble the positive videos in different degrees. Intuitively, we may capture more informative cues from the negative videos if we assign them fine-grained labels, thus benefiting the classifier learning. Aiming for this, we use a statistical method on both the positive and negative examples to get the decisive attributes of a specific event. Based on these decisive attributes, we assign the fine-grained labels to negative examples to treat them differently for more effective exploitation. The resulting fine-grained labels may be not accurate enough to characterize the negative videos. Hence, we propose to jointly optimize the fine-grained labels with the knowledge from the visual features and the attributes representations, which brings mutual reciprocality. Our model obtains two kinds of classifiers, one from the attributes and one from the features, which incorporate the informative cues from the fine-grained labels. The outputs of both classifiers on the testing videos are fused for detection. Extensive experiments on the challenging TRECVID MED 2012 development set have validated the efficacy of our proposed approach. Zhigang Ma, Yi Yang 0001, Zhongwen Xu, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2013 | Multi-camera Egocentric Activity Detection for Personal Assistant
Yue Gao 0002, Alex Hauptmann 0001 |
MMM (2) | 5 |
| 2013 | Unified Dictionary Learning and Region Tagging with Hierarchical Sparse Representation
Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
Comput. Vis. Image Underst. | 6 |
| 2013 | Corrigendum to cross-domain video concept detection: A joint discriminative and generative active learning approach [Expert Systems with Applications 39 (15) (2012) 12220-12228]
Yang Liu 0007, Alex Hauptmann 0001, Zhang Xiong 0001 |
Expert Syst. Appl. | 4 |
| 2013 | The co-attention model for tiny activity analysis
Ziyu Guan, Alex Hauptmann 0001 |
Neurocomputing | 3 |
| 2013 | Infrared Patch-Image Model for Small Target Detection in a Single ImageabstractThe robust detection of small targets is one of the key techniques in infrared search and tracking applications. A novel small target detection method in a single infrared image is proposed in this paper. Initially, the traditional infrared image model is generalized to a new infrared patch-image model using local patch construction. Then, because of the non-local self-correlation property of the infrared background image, based on the new model small target detection is formulated as an optimization problem of recovering low-rank and sparse matrices, which is effectively solved using stable principle component pursuit. Finally, a simple adaptive segmentation method is used to segment the target image and the segmentation result can be refined by post-processing. Extensive synthetic and real data experiments show that under different clutter backgrounds the proposed method not only works more stably for different target sizes and signal-to-clutter ratio values, but also has better detection performance compared with conventional baseline methods. Chenqiang Gao, Deyu Meng, Yi Yang 0001, Yongtao Wang, Xiaofang Zhou 0001, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 6 |
| 2013 | Multimedia Event Detection Using A Classifier-Specific Intermediate RepresentationabstractMultimedia event detection (MED) plays an important role in many applications such as video indexing and retrieval. Current event detection works mainly focus on sports and news event detection or abnormality detection in surveillance videos. Differently, our research aims to detect more complicated and generic events within a longer video sequence. In the past, researchers have proposed using intermediate concept classifiers with concept lexica to help understand the videos. Yet it is difficult to judge how many and what concepts would be sufficient for the particular video analysis task. Additionally, obtaining robust semantic concept classifiers requires a large number of positive training examples, which in turn has high human annotation cost. In this paper, we propose an approach that exploits the external concepts-based videos and event-based videos simultaneously to learn an intermediate representation from video features. Our algorithm integrates the classifier inference and latent intermediate representation into a joint framework. The joint optimization of the intermediate representation and the classifier makes them mutually beneficial and reciprocal. Effectively, the intermediate representation and the classifier are tightly correlated. The classifier dependent intermediate representation not only accurately reflects the task semantics but is also more suitable for the specific classifier. Thus we have created a discriminative semantic analysis framework based on a tightly coupled intermediate representation. Extensive experiments on multimedia event detection using real-world videos demonstrate the effectiveness of the proposed approach. Zhigang Ma, Yi Yang 0001, Nicu Sebe, Kai Zheng 0001, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 5 |
| 2013 | Feature Selection for Multimedia Analysis by Sharing Information Among Multiple TasksabstractWhile much progress has been made to multi-task classification and subspace learning, multi-task feature selection has long been largely unaddressed. In this paper, we propose a new multi-task feature selection algorithm and apply it to multimedia (e.g., video and image) analysis. Instead of evaluating the importance of each feature individually, our algorithm selects features in a batch mode, by which the feature correlation is considered. While feature selection has received much research attention, less effort has been made on improving the performance of feature selection by leveraging the shared knowledge from multiple related tasks. Our algorithm builds upon the assumption that different related tasks have common structures. Multiple feature selection functions of different tasks are simultaneously learned in a joint framework, which enables our algorithm to utilize the common knowledge of multiple tasks as supplementary information to facilitate decision making. An efficient iterative algorithm is proposed to optimize it, whose convergence is guaranteed. Experiments on different databases have demonstrated the effectiveness of the proposed algorithm. Yi Yang 0001, Zhigang Ma, Alex Hauptmann 0001, Nicu Sebe |
IEEE Trans. Multim. | 3 |
| 2013 | Multi-Feature Fusion via Hierarchical Regression for Multimedia AnalysisabstractMultimedia data are usually represented by multiple features. In this paper, we propose a new algorithm, namely Multi-feature Learning via Hierarchical Regression for multimedia semantics understanding, where two issues are considered. First, labeling large amount of training data is labor-intensive. It is meaningful to effectively leverage unlabeled data to facilitate multimedia semantics understanding. Second, given that multimedia data can be represented by multiple features, it is advantageous to develop an algorithm which combines evidence obtained from different features to infer reliable multimedia semantic concept classifiers. We design a hierarchical regression model to exploit the information derived from each type of feature, which is then collaboratively fused to obtain a multimedia semantic concept classifier. Both label information and data distribution of different features representing multimedia data are considered. The algorithm can be applied to a wide range of multimedia applications and experiments are conducted on video data for video concept annotation and action recognition. Using Trecvid and CareMedia video datasets, the experimental results show that it is beneficial to combine multiple features. The performance of the proposed algorithm is remarkable when only a small amount of labeled training data are available. Yi Yang 0001, Jingkuan Song, Zi Huang, Zhigang Ma, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 6 |
| 2012 | Learning to predict health status of geriatric patients from observational dataabstractData for diagnosis and clinical studies are now typically gathered by hand. While more detailed, exhaustive behavioral assessments scales have been developed, they have the drawback of being too time consuming and manual assessment can be subjective. Besides, clinical knowledge is required for accurate manual assessment, for which extensive training is needed. Therefore our great research challenge is to leverage machine learning techniques to better understand patients health status automatically based on continuous computer observations. In this paper, we study the problem of health status prediction for geriatric patients using observational data. In the first part of this paper, we propose a distance metric learning algorithm to learn a Mahalanobis distance which is more precise for similarity measures. In the second part, we propose a robust classifier based on ℓ2,1-norm regression to predict the geriatric patients' health status. We test the algorithm on a dataset collected from a nursing home. Experiment shows that our algorithm achieves encouraging performance. Yi Yang 0001, Alex Hauptmann 0001, Ming-yu Chen 0001, Yang Cai 0002, Ashok Bharucha, Howard D. Wactlar |
CIBCB | 2 |
| 2012 | Action recognition by exploring data distribution and feature correlationabstractHuman action recognition in videos draws strong research interest in computer vision because of its promising applications for video surveillance, video annotation, interactive gaming, etc. However, the amount of video data containing human actions is increasing exponentially, which makes the management of these resources a challenging task. Given a database with huge volumes of unlabeled videos, it is prohibitive to manually assign specific action types to these videos. Considering that it is much easier to obtain a small number of labeled videos, a practical solution for organizing them is to build a mechanism which is able to conduct action annotation automatically by leveraging the limited labeled videos. Motivated by this intuition, we propose an automatic video annotation algorithm by integrating semi-supervised learning and shared structure analysis into a joint framework for human action recognition. We apply our algorithm on both synthetic and realistic video datasets, including KTH [20], CareMedia dataset [1], Youtube action [12] and its extended version, UCF50 [2]. Extensive experiments demonstrate that the proposed algorithm outperforms the compared algorithms for action recognition. Most notably, our method has a very distinct advantage over other compared algorithms when we have only a few labeled samples. Sen Wang 0001, Yi Yang 0001, Zhigang Ma, Xue Li 0001, Chaoyi Pang, Alex Hauptmann 0001 |
CVPR | 6 |
| 2012 | Activity Recognition from RGB-D Camera with 3D Local Spatio-temporal FeaturesabstractKinect, as a 3D digital capturing device, can collect the RGB and depth information of human activities rapidly. We study fusing the depth and RGB information for activity recognition. We introduce histogram color-based image thresholding to detect skin on human body, and use a GMM model to segment human hand areas. We design a new local descriptor, called a 3D Motion Scale-Invariant Feature Transform (3D MoSIFT), which can effectively detect interesting points based on both RGB and depth information, and consequently encode the visual and motion information from both to describe the interesting points. Experiments, based on a video dataset collected by a Kinect camera, show that adding depth information in the descriptor can distinctly improve the accuracy of human activity recognition. We introduce the F1-score measurement to evaluate and compare our performance with the other algorithms. Yue Ming 0001, Qiuqi Ruan, Alex Hauptmann 0001 |
ICME | 3 |
| 2012 | Constrained keypoint quantization: towards better bag-of-words model for large-scale multimedia retrievalabstractBag-of-words models are among the most widely used and successful representations in multimedia retrieval. However, the quantization error which is introduced when mapping keypoints to visual words is one of the main drawbacks of the bag-of-words model. Although some techniques, such as soft-assignment to bags [23] and query expansion [27], have been introduced to deal with the problem, the performance gain is always at the cost of longer query response time, which makes them difficult to apply to large-scale multimedia retrieval applications. In this paper, we propose a simple "constrained keypoint quantization" method which can effectively reduce the overall quantization error of the bag-of-words representation and greatly improve the retrieval efficiency at the same time. The central idea of the proposed quantization method is that if a keypoint is far away from all visual words, we simply remove it. At first glance, this simple strategy seems naive and dangerous. However, we show that the proposed method has a solid theoretical background. Our experimental results on three widely used datasets for near duplicate image and video retrieval confirm that by removing a large amount of keypoints which have high quantization error, we obtain comparable or even better retrieval performance while dramatically boosting retrieval efficiency. Yang Cai 0002, Linjun Yang, Alex Hauptmann 0001 |
ICMR | 4 |
| 2012 | Beyond audio and video retrieval: towards multimedia summarizationabstractGiven the deluge of multimedia content that is becoming available over the Internet, it is increasingly important to be able to effectively examine and organize these large stores of information in ways that go beyond browsing or collaborative filtering. In this paper we review previous work on audio and video processing, and define the task of Topic-Oriented Multimedia Summarization (TOMS) using natural language generation: given a set of automatically extracted features from a video (such as visual concepts and ASR transcripts) a TOMS system will automatically generate a paragraph of natural language ("a recounting"), which summarizes the important information in a video belonging to a certain topic area, and provides explanations for why a video was matched and retrieved. We see this as a first step towards systems that will be able to discriminate visually similar, but semantically different videos, compare two videos and provide textual output or summarize a large number of videos at once. In this paper, we introduce our approach of solving the TOMS problem. We extract visual concept features and ASR transcription features from a given video, and develop a template-based natural language generation system to produce a textual recounting based on the extracted features. We also propose possible experimental designs for continuously evaluating and improving TOMS systems, and present results of a pilot evaluation of our initial system. Duo Ding, Florian Metze, Shourabh Rawat, Peter Schulam 0001, Susanne Burger, Ehsan Younessian, Michael G. Christel, Alex Hauptmann 0001 |
ICMR | 9 |
| 2012 | Classifier-specific intermediate representation for multimedia tasksabstractVideo annotation and multimedia classification play important roles in many applications such as video indexing and retrieval. To improve video annotation and event detection, researchers have proposed using intermediate concept classifiers with concept lexica to help understand the videos. Yet it is difficult to judge how many and what concepts would be sufficient for the particular video analysis task. Additionally, obtaining robust semantic concept classifiers requires a large number of positive training examples, which in turn has high human annotation cost. In this paper, we propose an approach that is able to automatically learn an intermediate representation from video features together with a classifier. The joint optimization of the two components makes them mutually beneficial and reciprocal. Effectively, the intermediate representation and the classifier are tightly correlated. The classifier dependent intermediate representation not only accurately reflects the task semantics but is also more suitable for the specific classifier. Thus we have created a discriminative semantic analysis framework based on a tightly-coupled intermediate representation. Several experiments on video annotation and multimedia event detection using real-world videos demonstrate the effectiveness of the proposed approach. Zhigang Ma, Alex Hauptmann 0001, Yi Yang 0001, Nicu Sebe |
ICMR | 2 |
| 2012 | Multimodal knowledge-based analysis in multimedia event detectionabstractMultimedia Event Detection (MED) is a multimedia retrieval task with the goal of finding videos of a particular event in a large-scale Internet video archive, given example videos and text descriptions. We focus on the multimodal knowledge-based analysis in MED where we utilize meaningful and semantic features such as Automatic Speech Recognition (ASR) transcripts, acoustic concept indexing (i.e. 42 acoustic concepts) and visual semantic indexing (i.e. 346 visual concepts) to characterize videos in archive. We study two scenarios where we either do or do not use the provided example videos. In the former, we propose a novel Adaptive Semantic Similarity (ASS) to measure textual similarity between ASR transcripts of videos. We also incorporate acoustic concept indexing and classification to retrieve test videos, specially with too few spoken words. In the latter 'ad-hoc' scenario where we do not have any example video, we use only the event kit description to retrieve test videos ASR transcripts and visual semantics. We also propose an event-specific fusion scheme to combine textual and visual retrieval outputs. Our results show the effectiveness of the proposed ASS and acoustic concept indexing methods and their complimentary role. We also conduct a set of experiments to assess the proposed framework for the 'ad-hoc' scenario. Ehsan Younessian, Teruko Mitamura, Alex Hauptmann 0001 |
ICMR | 3 |
| 2012 | Leveraging high-level and low-level features for multimedia event detectionabstractThis paper addresses the challenge of Multimedia Event Detection by proposing a novel method for high-level and low-level features fusion based on collective classification. Generally, the method consists of three steps: training a classifier from low-level features; encoding high-level features into graphs; and diffusing the scores on the established graph to obtain the final prediction. The final prediction is derived from multiple graphs each of which corresponds to a high-level feature. The paper investigates two graph construction methods using logarithmic and exponential loss functions, respectively and two collective classification algorithms, i.e. Gibbs sampling and Markov random walk. The theoretical analysis demonstrates that the proposed method converges and is computationally scalable and the empirical analysis on TRECVID 2011 Multimedia Event Detection dataset validates its outstanding performance compared to state-of-the-art methods, with an added benefit of interpretability. Lu Jiang 0004, Alex Hauptmann 0001, Guang Xiang |
ACM Multimedia | 2 |
| 2012 | Knowledge adaptation for ad hoc multimedia event detection with few exemplarsabstractMultimedia event detection (MED) has a significant impact on many applications. Though video concept annotation has received much research effort, video event detection remains largely unaddressed. Current research mainly focuses on sports and news event detection or abnormality detection in surveillance videos. Our research on this topic is capable of detecting more complicated and generic events. Moreover, the curse of reality, i.e., precisely labeled multimedia content is scarce, necessitates the study on how to attain respectable detection performance using only limited positive examples. Research addressing these two aforementioned issues is still in its infancy. In light of this, we explore Ad Hoc MED, which aims to detect complicated and generic events by using few positive examples. To the best of our knowledge, our work makes the first attempt on this topic. As the information from these few positive examples is limited, we propose to infer knowledge from other multimedia resources to facilitate event detection. Experiments are performed on real-world multimedia archives consisting of several challenging events. The results show that our approach outperforms several other detection algorithms. Most notably, our algorithm outperforms SVM by 43% and 14% comparatively in Average Precision when using Gaussian and Χ2 kernel respectively. Zhigang Ma, Yi Yang 0001, Yang Cai 0002, Nicu Sebe, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2012 | Double Fusion for Multimedia Event Detection
Zhen-Zhong Lan, Shoou-I Yu, Wei Liu 0015, Alex Hauptmann 0001 |
MMM | 5 |
| 2012 | Symbiotic Black-Box Tracker
Yue Gao 0002, Alex Hauptmann 0001, Rongrong Ji, Boaz J. Super |
MMM | 3 |
| 2012 | Cross-domain video concept detection: A joint discriminative and generative active learning approach
Yang Liu 0007, Alex Hauptmann 0001, Zhang Xiong 0001 |
Expert Syst. Appl. | 4 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 6 |
| 2012 | Web-Scale Multimedia Processing and Applications [Scanning the Issue]abstractThe articles in this special issue focus on web-scale multimedia processing as well as applications for its use. Edward Y. Chang, Shih-Fu Chang, Alex Hauptmann 0001, Thomas S. Huang, Malcolm Slaney |
Proc. IEEE | 3 |
| 2012 | A Framework for Classifier Adaptation for Large-Scale Multimedia DataabstractMachine learning techniques have been used extensively to build models for the analysis and retrieval of multimedia data. The explosion of multimedia data on the Web poses a great challenge to such techniques not simply because of the sheer data volume, but also because of the heterogeneity of the data. With data from a wide variety of domains, models trained from one domain do not generalize well to other domains, while at the same time it is prohibitively expensive to build new models for each and every domain due to the high cost for labeling training examples. In this paper, we tackle the heterogeneity challenge in large-scale multimedia data using cross-domain model adaptation for better performance and reduced human cost. Specifically, we investigate the problem of adapting supervised classifiers trained from one or more source domains to a new classifier for a target domain that has only limited labeled examples. The foundation of our work is a general framework for function-level classifier adaptation based on the regularized loss minimization principle, which adapts a classifier by directly modifying its decision function. Under this framework, one can derive concrete adaptation algorithms by plugging in any loss and regularization functions, among which we elaborate on adaptive support vector machines (a-SVM). We further extend this framework for multiclassifier adaptation, namely adapting multiple existing classifiers into a classifier for the target domain, in a way that the contributions of these existing classifiers are automatically determined. We evaluate the proposed approaches in cross-domain semantic concept detection based on TRECVID corpora. The results show that our approaches outperform existing (adaptation and nonadaptation) methods in terms of accuracy and/or efficiency, and adaptation from multiple classifiers offers further benefits. Jun Yang 0003, Alex Hauptmann 0001 |
Proc. IEEE | 3 |
| 2012 | Spline Regression Hashing for Fast Image SearchabstractTechniques for fast image retrieval over large databases have attracted considerable attention due to the rapid growth of web images. One promising way to accelerate image search is to use hashing technologies, which represent images by compact binary codewords. In this way, the similarity between images can be efficiently measured in terms of the Hamming distance between their corresponding binary codes. Although plenty of methods on generating hash codes have been proposed in recent years, there are still two key points that needed to be improved: 1) how to precisely preserve the similarity structure of the original data and 2) how to obtain the hash codes of the previously unseen data. In this paper, we propose our spline regression hashing method, in which both the local and global data similarity structures are exploited. To better capture the local manifold structure, we introduce splines developed in Sobolev space to find the local data mapping function. Furthermore, our framework simultaneously learns the hash codes of the training data and the hash function for the unseen data, which solves the out-of-sample problem. Extensive experiments conducted on real image datasets consisting of over one million images show that our proposed method outperforms the state-of-the-art techniques. Yang Liu 0098, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 5 |
| 2012 | Web and Personal Image Annotation by Mining Label Correlation With Relaxed Visual Graph EmbeddingabstractThe number of digital images rapidly increases, and it becomes an important challenge to organize these resources effectively. As a way to facilitate image categorization and retrieval, automatic image annotation has received much research attention. Considering that there are a great number of unlabeled images available, it is beneficial to develop an effective mechanism to leverage unlabeled images for large-scale image annotation. Meanwhile, a single image is usually associated with multiple labels, which are inherently correlated to each other. A straightforward method of image annotation is to decompose the problem into multiple independent single-label problems, but this ignores the underlying correlations among different labels. In this paper, we propose a new inductive algorithm for image annotation by integrating label correlation mining and visual similarity mining into a joint framework. We first construct a graph model according to image visual features. A multilabel classifier is then trained by simultaneously uncovering the shared structure common to different labels and the visual graph embedded label prediction matrix for image annotation. We show that the globally optimal solution of the proposed framework can be obtained by performing generalized eigen-decomposition. We apply the proposed framework to both web image annotation and personal album labeling using the NUS-WIDE, MSRA MM 2.0, and Kodak image data sets, and the AUC evaluation metric. Extensive experiments on large-scale image databases collected from the web and personal album show that the proposed algorithm is capable of utilizing both labeled and unlabeled data for image annotation and outperforms other algorithms. Yi Yang 0001, Fei Wu 0001, Feiping Nie 0001, Heng Tao Shen, Yueting Zhuang, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 6 |
| 2012 | Discriminating Joint Feature Analysis for Multimedia Data UnderstandingabstractIn this paper, we propose a novel semi-supervised feature analyzing framework for multimedia data understanding and apply it to three different applications: image annotation, video concept detection and 3-D motion data analysis. Our method is built upon two advancements of the state of the art: (1)l2, 1-norm regularized feature selection which can jointly select the most relevant features from all the data points. This feature selection approach was shown to be robust and efficient in literature as it considers the correlation between different features jointly when conducting feature selection; (2) manifold learning which analyzes the feature space by exploiting both labeled and unlabeled data. It is a widely used technique to extend many algorithms to semi-supervised scenarios for its capability of leveraging the manifold structure of multimedia data. The proposed method is able to learn a classifier for different applications by selecting the discriminating features closely related to the semantic concepts. The objective function of our method is non-smooth and difficult to solve, so we design an efficient iterative algorithm with fast convergence, thus making it applicable to practical applications. Extensive experiments on image annotation, video concept detection and 3-D motion data analysis are performed on different real-world data sets to demonstrate the effectiveness of our algorithm. Zhigang Ma, Feiping Nie 0001, Yi Yang 0001, Jasper R. R. Uijlings, Nicu Sebe, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 6 |
| 2011 | People detection based on appearance and motion modelsabstractThe main contribution of this paper is a new people detection algorithm based on motion information. The algorithm builds a people motion model based on the Implicit Shape Model (ISM) Framework and the MoSIFT descriptor. We also propose a detection system that integrates appearance, motion and tracking information. Experimental results over sequences extracted from the TRECVID dataset show that our new people motion detector produces results comparable to the state of the art and that the proposed multimodal fusion system improves the obtained results combining the three information sources. Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez |
AVSS | 2 |
| 2010 | Explicit and implicit concept-based video retrieval with bipartite graph propagation modelabstractThe major scientific problem for content-based video retrieval is the semantic gap. Generally speaking, there are two appropriate ways to bridge the semantic gap: the first one is from human perspective (top-down) and the other one is from computer perspective (bottom-up). The top-down method defines a concept lexicon from human perspective, trains the detector for each concept based on supervised learning, and then indexes the corpus with concept detectors. Since each concept has an explicit semantic meaning, we call this concept as an explicit concept. The bottom-up approach directly discovers the underlying latent topics from video corpus by machine perspective using an unsupervised learning. The video corpus is indexed subsequently by these latent topics. As opposite to explicit concepts, we name latent topics as implicit concepts. Given the explicit concept set is pre-defined and independent of the corpus, it is impossible to completely describe corpus and users' queries. On the other hand, the implicit concepts are dynamic and dependent on the corpus, which is able to fully describe corpus and users' queries. Therefore, combining explicit and implicit concepts could be a promising way to bridge the semantic gap effectively. In this paper, a Bipartite Graph Propagation Model (BGPM) is applied to automatically balance influences from explicit and implicit concepts. Concept nodes with strong connections to queries are reinforced no matter explicit or implicit. Demonstrated by the experiments on TREVID 2008 video dataset, BGPM successfully fuses explicit and implicit concepts to achieve a significant improvement on 48 search tasks. Juan Cao 0001, Yongdong Zhang 0001, Jintao Li 0001, Ming-yu Chen 0001, Alex Hauptmann 0001 |
ACM Multimedia | 6 |
| 2010 | ACM international workshop on very-large-scale multimedia corpus, mining and retrieval (VLS-MCMR'10)abstractThe purpose of this workshop is to bring together researchers interested in the construction and analysis of Very Large Scale Multimedia Corpus, as well as the methodologies to Mine and Retrieve information from them. The Workshop will provide a forum to consolidate key issues related to research on very large scale multimedia dataset such as the construction of dataset, creation of ground truth, sharing and extension of such resources in terms of ground truth, features, algorithms and tools etc. The Workshop will discuss and formulate action plan towards these goals. Benoit Huet, Tat-Seng Chua, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2010 | Hybrid active learning for cross-domain video concept detectionabstractCross-domain video concept detection is a challenging task due to the distribution difference between the source domain and target domain. In order to avoid expensive labeling the target-domain data, Active Learning can be used to incrementally learn a target classifier by reusing the one in the source domain. It uses a discriminative query strategy and picks the most ambiguous samples to label, which could fail if the distribution difference is too large. In this paper, to deal with large difference in data distributions, we propose a generative query strategy which is then combined with the existing discriminative one to yield a hybrid method. This method adaptively fits the distribution differences and gives a mixture strategy that performs more robustly compared to both single strategies. Experimental results on TRECVID semantic concept detection task demonstrate superior performance of our hybrid method. Ming-yu Chen 0001, Alex Hauptmann 0001, Zhang Xiong 0001 |
ACM Multimedia | 4 |
| 2010 | Exploiting multi-level parallelism for low-latency activity recognition in streaming videoabstractVideo understanding is a computationally challenging task that is critical not only for traditionally throughput-oriented applications such as search but also latency-sensitive interactive applications such as surveillance, gaming, videoconferencing, and vision-based user interfaces. Enabling these types of video processing applications will require not only new algorithms and techniques, but new runtime systems that optimize latency as well as throughput. In this paper, we present a runtime system called Sprout that achieves low latency by exploiting the parallelism inherent in video understanding applications. We demonstrate the utility of our system on an activity recognition application that employs a robust new descriptor called MoSIFT, which explicitly augments appearance features with motion information. MoSIFT outperforms previous recognition techniques, but like other state-of-the-art techniques, it is computationally expensive -- a sequential implementation runs 100 times slower than real time. We describe the implementation of the activity recognition application on Sprout, and show that it can accurately recognize activities at full frame rate (25 fps) and low latency on a challenging airport surveillance video corpus. Ming-yu Chen 0001, Lily B. Mummert, Padmanabhan Pillai, Alex Hauptmann 0001, Rahul Sukthankar |
MMSys | 4 |
| 2010 | Representations of Keypoint-Based Semantic Concept Detection: A Comprehensive StudyabstractBased on the local keypoints extracted as salient image patches, an image can be described as a ¿bag-of-visual-words (BoW)¿ and this representation has appeared promising for object and scene classification. The performance of BoW features in semantic concept detection for large-scale multimedia databases is subject to various representation choices. In this paper, we conduct a comprehensive study on the representation choices of BoW, including vocabulary size, weighting scheme, stop word removal, feature selection, spatial information, and visual bi-gram. We offer practical insights in how to optimize the performance of BoW by choosing appropriate representation choices. For the weighting scheme, we elaborate a soft-weighting method to assess the significance of a visual word to an image. We experimentally show that the soft-weighting outperforms other popular weighting schemes such as TF-IDF with a large margin. Our extensive experiments on TRECVID data sets also indicate that BoW feature alone, with appropriate representation choices, already produces highly competitive concept detection performance. Based on our empirical findings, we further apply our method to detect a large set of 374 semantic concepts. The detectors, as well as the features and detection scores on several recent benchmark data sets, are released to the multimedia community. Yu-Gang Jiang 0001, Jun Yang 0003, Chong-Wah Ngo, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 4 |
| 2009 | ACM SIGMM the first workshop on web-scale multimedia corpus (WSMC09)abstract10.1145/1631272.1631551 Benoit Huet, Jinhui Tang 0001, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2009 | Identifying news videos' ideological perspectives using emphatic patterns of visual conceptsabstractTelevision news has become the predominant way of understanding the world around us, but individual news broadcasters can frame or mislead an audience's understanding of political and social issues. We are developing a computer system that can automatically identify highly biased television news and encourage audiences to seek news stories from contrasting viewpoints. But can computers identify the ideological perspective from which a news video was produced? We propose a method based on an empathic pattern of visual concepts: news broadcasters holding contrasting ideological beliefs appear to emphasize different subsets of visual concepts. We formalize the emphatic patterns and propose a statistical model. We evaluate the proposed model on a large broadcast news video archive with promising experimental results. Weihao Lin 0003, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2009 | Real-Time Near-Duplicate Elimination for Web Video Search With Content and ContextabstractWith the exponential growth of social media, there exist huge numbers of near-duplicate web videos, ranging from simple formatting to complex mixture of different editing effects. In addition to the abundant video content, the social web provides rich sets of context information associated with web videos, such as thumbnail image, time duration and so on. At the same time, the popularity of Web 2.0 demands for timely response to user queries. To balance the speed and accuracy aspects, in this paper, we combine the contextual information from time duration, number of views, and thumbnail images with the content analysis derived from color and local points to achieve real-time near-duplicate elimination. The results of 24 popular queries retrieved from YouTube show that the proposed approach integrating content and context can reach real-time novelty re-ranking of web videos with extremely high efficiency, where the majority of duplicates can be rapidly detected and removed from the top rankings. The speedup of the proposed approach can reach 164 times faster than the effective hierarchical method proposed in, with just a slight loss of performance. Xiao Wu 0001, Chong-Wah Ngo, Alex Hauptmann 0001, Hung-Khoon Tan |
IEEE Trans. Multim. | 3 |
| 2008 | Vox Populi Annotation: Measuring Intensity of Ideological Perspectives by Aggregating Group Judgments
Weihao Lin 0003, Alex Hauptmann 0001 |
LREC | 2 |
| 2008 | A Joint Topic and Perspective Model for Ideological Discourse
Weihao Lin 0003, Eric P. Xing, Alex Hauptmann 0001 |
ECML/PKDD (2) | 3 |
| 2008 | Measuring novelty and redundancy with multiple modalities in cross-lingual broadcast news
Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
Comput. Vis. Image Underst. | 2 |
| 2008 | Video Retrieval Based on Semantic ConceptsabstractAn approach using many intermediate semantic concepts is proposed with the potential to bridge the semantic gap between what a color, shape, and texture-based ldquolow-levelrdquo image analysis can extract from video and what users really want to find, most likely using text descriptions of their information needs. Semantic concepts such as cars, planes, roads, people, animals, and different types of scenes (outdoor, night time, etc.) can be automatically detected in the video with reasonable accuracy. This leads us to ask how can they be used automatically and how does a user (or a retrieval system) translate the user's information need into a selection of related concepts that would help find the relevant video clips, from the large list of available concepts. We illustrate how semantic concept retrieval can be automatically exploited by mapping queries into query classes and through pseudo-relevance feedback. We also provide evidence how a semantic concept can be utilized by users in interactive retrieval, through interfaces that provide affordances of explicit concept selection and search, concept filtering, and relevance feedback. How many concepts we actually need and how accurately they need to be detected and linked through various relationships is specified in the ontology structure. Alex Hauptmann 0001, Michael G. Christel |
Proc. IEEE | 1 |
| 2008 | Multimodal News Story Clustering With Pairwise Visual Near-Duplicate ConstraintabstractStory clustering is a critical step for news retrieval, topic mining, and summarization. Nonetheless, the task remains highly challenging owing to the fact that news topics exhibit clusters of varying densities, shapes, and sizes. Traditional algorithms are found to be ineffective in mining these types of clusters. This paper offers a new perspective by exploring the pairwise visual cues deriving from near-duplicate keyframes (NDK) for constraint-based clustering. We propose a constraint-driven co-clustering algorithm (CCC), which utilizes the near-duplicate constraints built on top of text, to mine topic-related stories and the outliers. With CCC, the duality between stories and their underlying multimodal features is exploited to transform features in low-dimensional space with normalized cut. The visual constraints are added directly to this new space, while the traditional DBSCAN is revisited to capitalize on the availability of constraints and the reduced dimensional space. We modify DBSCAN with two new characteristics for story clustering: 1) constraint-based centroid selection and 2) adaptive radius. Experiments on TRECVID-2004 corpus demonstrate that CCC with visual constraints is more capable of mining news topics of varying densities, shapes and sizes, compared with traditionalk-means, DBSCAN, and spectral co-clustering algorithms. Xiao Wu 0001, Chong-Wah Ngo, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 3 |
| 2007 | Query expansion using probabilistic local feedback with application to multimedia retrievalabstractAs one of the most effective query expansion approaches, local feedback is able to automatically discover new query terms and improve retrieval accuracy for different retrieval models. However, the performance of local feedback is heavily dependent on the assumption that most top-ranked documents are relevant to the query topic. Although this assumption might be sensible for ad-hoc text retrieval, it is usually violated in many other retrieval tasks such as multimedia retrieval. In this paper, we develop a robust local analysis approach called probabilistic local feedback (PLF) based on a discriminative probabilistic retrieval framework. The proposed model is effective for improving retrieval accuracy without assuming the most top-ranked documents are relevant. It also provides a sound probabilistic interpretation and a convergence guarantee on the iterative result updating process. Although derived from variational techniques, this approach only involves an iterative process of simple operations on ranking features and thus can be computed efficiently in practice. Our multimedia retrieval experiments on TRECVID'03-'05 collections have demonstrated the advantage of the proposed PLF approaches which can achieve noticeable gains in terms of mean average precision over various baseline methods and PRF-augmented results. Alex Hauptmann 0001 |
CIKM | 2 |
| 2007 | Undirected Graphical Models for Video Analysis and ClassificationabstractAccurate and efficient video classification and retrieval demands the fusion of multimodal information and the use of intermediate representations. This paper describes an undirected graphical model based on exponential-family harmonium, which derives intermediate semantic representations of video data by jointly modeling the textual and image information in the video. We propose an extension of the model to derive category-specific video representation and integrate video classification as a part of the modeling process. We report satisfactory classification performance on a set of 15 video categories from TRECVID collection as well as comparison on the effectiveness of different inference algorithms. Yan Liu 0002, Jun Yang 0003, Alex Hauptmann 0001 |
ICME | 3 |
| 2007 | Novelty detection for cross-lingual news stories with visual duplicates and speech transcriptsabstractAn overwhelming volume of news videos from different channels and languages is available today, which demands automatic management of this abundant information. To effectively search, retrieve, browse and track cross-lingual news stories, a news story similarity measure plays a critical role in assessing the novelty and redundancy among them. In this paper, we explore the novelty and redundancy detection with visual duplicates and speech transcripts for cross-lingual news stories. News stories are represented by a sequence of keyframes in the visual track and a set of words extracted from speech transcript in the audio track. A major difference to pure text documents is that the number of keyframes in one story is relatively small compared to the number of words and there exist a large number of non-near duplicate keyframes. These features make the behavior of similarity measures different compared to traditional textual collections. Furthermore, the textual features and visual features complement each other for news stories. They can be further combined to boost the performance. Experiments on the TRECVID-2005 cross-lingual news video corpus show that approaches on textual features and visual features demonstrate different performance, and measures on visual features are quite effective. Overall, the cosine distance on keyframes is still a robust measure. Language models built on visual features demonstrate promising performance. The fusion of textual and visual features improves overall performance. Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2007 | Practical elimination of near-duplicates from web video searchabstractCurrent web video search results rely exclusively on text keywords or user-supplied tags. A search on typical popular video often returns many duplicate and near-duplicate videos in the top results. This paper outlines ways to cluster and filter out the near-duplicate video using a hierarchical approach. Initial triage is performed using fast signatures derived from color histograms. Only when a video cannot be clearly classified as novel or near-duplicate using global signatures, we apply a more expensive local feature based near-duplicate detection which provides very accurate duplicate analysis through more costly computation. The results of 24 queries in a data set of 12,790 videos retrieved from Google, Yahoo! and YouTube show that this hierarchical approach can dramatically reduce redundant video displayed to the user in the top result set, at relatively small computational cost. Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2007 | Cross-domain video concept detection using adaptive svmsabstractMany multimedia applications can benefit from techniques for adapting existing classifiers to data with different distributions. One example is cross-domain video concept detection which aims to adapt concept classifiers across various video domains. In this paper, we explore two key problems for classifier adaptation: (1) how to transform existing classifier(s) into an effective classifier for a new dataset that only has a limited number of labeled examples, and (2) how to select the best existing classifier(s) for adaptation. For the first problem, we propose Adaptive Support Vector Machines (A-SVMs) as a general method to adapt one or more existing classifiers of any type to the new dataset. It aims to learn the "delta function" between the original and adapted classifier using an objective function similar to SVMs. For the second problem, we estimate the performance of each existing classifier on the sparsely-labeled new dataset by analyzing its score distribution and other meta features, and select the classifiers with the best estimated performance. The proposed method outperforms several baseline and competing methods in terms of classification accuracy and efficiency in cross-domain concept detection in the TRECVID corpus. Jun Yang 0003, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2007 | Harmonium Models for Semantic Video Representation and ClassificationabstractAccurate and efficient video classification demands the fusion of multimodal information and the use of intermediate representations. Combining the two ideas into the one framework, we propose a probabilistic approach for video classification using intermediate semantic representations derived from multi-modal features. Based on a class of bipartite undirected graphical models named harmonium, our approach represents the video data as latent semantic topics derived by jointly modeling the transcript keywords and color-histogram features, and performs classification using these latent topics under a unified framework. We show satisfactory classification performance of our approach on a benchmark dataset as well as interesting insights into the data. Jun Yang 0003, Yan Liu 0002, Eric P. Xing, Alex Hauptmann 0001 |
SDM | 4 |
| 2007 | A review of text and image retrieval approaches for broadcast news video
Alex Hauptmann 0001 |
Inf. Retr. | 2 |
| 2007 | Can High-Level Concepts Fill the Semantic Gap in Video Retrieval? A Case Study With Broadcast NewsabstractA number of researchers have been building high-level semantic concept detectors such as outdoors, face, building, to help with semantic video retrieval. Our goal is to examine how many concepts would be needed, and how they should be selected and used. Simulating performance of video retrieval under different assumptions of concept detection accuracy, we find that good retrieval can be achieved even when detection accuracy is low, if sufficiently many concepts are combined. We also derive suggestions regarding the types of concepts that would be most helpful for a large concept lexicon. Since our user study finds that people cannot predict which concepts will help their query, we also suggest ways to find the best concepts to use. Ultimately, this paper concludes that "concept-based" video retrieval with fewer than 5000 concepts, detected with a minimal accuracy of 10% mean average precision is likely to provide high accuracy results in broadcast news retrieval. Alex Hauptmann 0001, Weihao Lin 0003, Michael G. Christel, Howard D. Wactlar |
IEEE Trans. Multim. | 1 |
| 2006 | Are These Documents Written from Different Perspectives? A Test of Different Perspectives Based on Statistical Distribution DivergenceabstractIn this paper we investigate how to automatically determine if two document collections are written from different perspectives. By perspectives we mean a point of view, for example, from the perspective of Democrats or Republicans. We propose a test of different perspectives based on distribution divergence between the statistical models of two collections. Experimental results show that the test can successfully distinguish document collections of different perspectives from other types of collections. Weihao Lin 0003, Alex Hauptmann 0001 |
ACL | 2 |
| 2006 | Which Side are You on? Identifying Perspectives at the Document and Sentence Levels
Weihao Lin 0003, Theresa Wilson, Janyce Wiebe, Alex Hauptmann 0001 |
CoNLL | 4 |
| 2006 | Which Thousand Words are Worth a Picture? Experiments on Video Retrieval using a Thousand ConceptsabstractIn contrast to traditional video retrieval that represents visual content with low-level features (e.g. color and texture), emerging concept-based video retrieval allows users to search video archives by specifying a limited number of high-level concepts (e.g. outdoors and car). Recent studies have demonstrated the feasibility of concept-based retrieval, but a fundamental question remains: what kinds of concepts should we index? We analyze a large video archive annotated with more than a thousand high-level concepts, and develop guidelines for choosing concepts of high utility to video retrieval Weihao Lin 0003, Alex Hauptmann 0001 |
ICME | 2 |
| 2006 | Label Disambiguation and Sequence Modeling for Identifying Human Activities from Wearable Physiological SensorsabstractWearable physiological sensors can provide a faithful record of a patient's physiological states without constant attention of caregivers. A computer program that can infer human activities from physiological recordings will be an valuable tool for physicians. In this paper we investigate to what extent current machine learning algorithms can correctly identify human activities from physiological sensors. We further identify two challenges that developers need to address. The first problem is that the labels of training data are inevitably noisy due to difficulties of annotating thousands hours of data. The second problem lies in the continuous nature of human activities, which violates the independence assumption made by many learning algorithms. We approach the first problem of noisy labeling in the multiple-label framework, and develop a conditional Markov models to take temporal context into consideration. We evaluate the proposed methods on 12,000 hours of the physiological recordings. The results show that support vector machines are effective to identify human activities from physiological signals, and efforts of disambiguating noisy labels are worthwhile Weihao Lin 0003, Alex Hauptmann 0001 |
ICME | 2 |
| 2006 | Mining Relationship Between Video Concepts using Probabilistic Graphical ModelsabstractFor large scale automatic semantic video characterization, it is necessary to learn and model a large number of semantic concepts. These semantic concepts do not exist in isolation to each other and exploiting this relationship between multiple video concepts could be a useful source to improve the concept detection accuracy. In this paper, we describe various multi-concept relational learning approaches via a unified probabilistic graphical model representation and propose using numerous graphical models to mine the relationship between video concepts that have not been applied before. Their performances in video semantic concept detection are evaluated and compared on two TRECVID'05 video collections Ming-yu Chen 0001, Alex Hauptmann 0001 |
ICME | 3 |
| 2006 | Extreme video retrieval: joint maximization of human and computer performanceabstractWe present an efficient system for video search that maximizes the use of human bandwidth, while at the same time exploiting the machine's ability to learn in real-time from user selected relevant video clips. The system exploits the human capability for rapidly scanning imagery augmenting it with an active learning loop, which attempts to always present the most relevant material based on the current information. Two versions of the human interface were evaluated, one with variable page sizes and manual paging, the other with a fixed page size and automatic paging. Both require absolute attention and focus of the user for optimal performance. In either case, as users search and find relevant results, the system can invisibly re-rank its previous best guesses using a number of knowledge sources, such as image similarity, text similarity, and temporal proximity. Experimental evidence shows a significant improvement using the combined extremes of human and machine power over either approach alone. Alex Hauptmann 0001, Weihao Lin 0003, Jun Yang 0003, Ming-yu Chen 0001 |
ACM Multimedia | 1 |
| 2006 | 3WNews: who, where, and when in news videoabstractWe describe 3WNews as a novel system for browsing news video by the people (who) and locations (where) appearing in the footage as well as the time (when) of news events. The people names, locations, and time expressions are recognized from video transcript and their ambiguous references are resolved. As a key advantage, 3WNews distinguishes the people and locations that actually appear in the video from those merely mentioned in the transcript, and uses them as (better) indexes for browsing. It also supports browsing of news video by event time instead of broadcasting time. Jun Yang 0003, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2006 | Probabilistic latent query analysis for combining multiple retrieval sourcesabstractCombining the output from multiple retrieval sources over the same document collection is of great importance to a number of retrieval tasks such as multimedia retrieval, web retrieval and meta-search. To merge retrieval sources adaptively according to query topics, we propose a series of new approaches called probabilistic latent query analysis (pLQA), which can associate non-identical combination weights with latent classes underlying the query space. Compared with previous query independent and query-class based combination methods, the proposed approaches have the advantage of being able to discover latent query classes automatically without using prior human knowledge, to assign one query to a mixture of query classes, and to determine the number of query classes under a model selection principle. Experimental results on two retrieval tasks, i.e., multimedia retrieval and meta-search, demonstrate that the proposed methods can uncover sensible latent classes from training data, and can achieve considerable performance gains. Alex Hauptmann 0001 |
SIGIR | 2 |
| 2006 | A Discriminative Learning Framework with Pairwise Constraints for Video Object ClassificationabstractTo deal with the problem of insufficient labeled data in video object classification, one solution is to utilize additional pairwise constraints that indicate the relationship between two examples, i.e., whether these examples belong to the same class or not. In this paper, we propose a discriminative learning approach which can incorporate pairwise constraints into a conventional margin-based learning framework. Different from previous work that usually attempts to learn better distance metrics or estimate the underlying data distribution, the proposed approach can directly model the decision boundary and, thus, require fewer model assumptions. Moreover, the proposed approach can handle both labeled data and pairwise constraints in a unified framework. In this work, we investigate two families of pairwise loss functions, namely, convex and nonconvex pairwise loss functions, and then derive three pairwise learning algorithms by plugging in the hinge loss and the logistic loss functions. The proposed learning algorithms were evaluated using a people identification task on two surveillance video data sets. The experiments demonstrated that the proposed pairwise learning algorithms considerably outperform the baseline classifiers using only labeled data and two other pairwise learning algorithms with the same amount of pairwise constraints. Jian Zhang 0003, Jie Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | Learning rich semantics from news video archives by style analysisabstractWe propose a generic and robust framework for news video indexing which we founded on a broadcast news production model. We identify within this model four production phases, each providing useful metadata for annotation. In contrast to semiautomatic indexing approaches which exploit this information at production time, we adhere to an automatic data-driven approach. To that end, we analyze a digital news video using a separate set of multimodal detectors for each production phase. By combining the resulting production-derived features into a statistical classifier ensemble, the framework facilitates robust classification of several rich semantic concepts in news video; rich meaning that concepts share many similarities in their production process. Experiments on an archive of 120 hours of news video from the 2003 TRECVID benchmark show that a combined analysis of production phases yields the best results. In addition, we demonstrate that the accuracy of the proposed style analysis framework for classification of several rich semantic concepts is state-of-the-art. Cees Snoek, Marcel Worring, Alex Hauptmann 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2005 | Putting active learning into multimedia applications: dynamic definition and refinement of concept classifiersabstractThe authors developed an extensible system for video exploitation that puts the user in control to better accommodate novel situations and source material. Visually dense displays of thumbnail imagery in storyboard views are used for shot-based video exploration and retrieval. The user can identify a need for a class of audiovisual detection, adeptly and fluently supply training material for that class, and iteratively evaluate and improve the resulting automatic classification produced via multiple modality active learning and SVM. By iteratively reviewing the output of the classifier and updating the positive and negative training samples with less effort than typical for relevance feedback systems, the user can play an active role in directing the classification process while still needing to truth only a very small percentage of the multimedia data set. Examples are given illustrating the iterative creation of a classifier for a concept of interest to be included in subsequent investigations, and for a concept typically deemed irrelevant to be weeded out in follow-up queries. Filtering and browsing tools making use of existing and iteratively added concepts put the user further in control of the multimedia browsing and retrieval process. Ming-yu Chen 0001, Michael G. Christel, Alex Hauptmann 0001, Howard D. Wactlar |
ACM Multimedia | 3 |
| 2005 | Multiple instance learning for labeling faces in broadcasting news videoabstractLabeling faces in news video with their names is an interesting research problem which was previously solved using supervised methods that demand significant user efforts on labeling training data. In this paper, we investigate a more challenging setting of the problem where there is no complete information on data labels. Specifically, by exploiting the uniqueness of a face's name, we formulate the problem as a special multi-instance learning (MIL) problem, namely exclusive MIL or eMIL problem, so that it can be tackled by a model trained with partial labeling information as the anonymity judgment of faces, which requires less user effort to collect. We propose two discriminative probabilistic learning methods named Exclusive Density (ED) and Iterative ED for eMIL problems. Experiments on the face labeling problem shows that the performance of the proposed approaches are superior to the traditional MIL algorithms and close to the performance achieved by supervised methods trained with complete data labels. Jun Yang 0003, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2005 | Revisiting the effect of topic set size on retrieval errorabstractNo abstract available. Weihao Lin 0003, Alex Hauptmann 0001 |
SIGIR | 2 |
| 2005 | Mining Associated Text and Images with Dual-Wing Harmoniums
Eric P. Xing, Alex Hauptmann 0001 |
UAI | 3 |
| 2004 | A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification
Jian Zhang 0003, Jie Yang 0001, Alex Hauptmann 0001 |
CVPR (2) | 4 |
| 2004 | Searching for a specific person in broadcast news videoabstractPeople as news subjects play an important role in broadcast news and finding a specific person is a major challenge for multimedia retrieval. Beyond mere content-based general retrieval, this task requires exploitation of the structure, time sequence and meaning of news content. We introduce a comprehensive approach to discovering clues for finding a specific person in broadcast news video. Various information aspects are investigated, including text information, timing information, scene detection, and face recognition. Experimental results on the TREC 2003 video search task show that our approach can achieve surprisingly high performance by exploiting broadly diverse information to find specific named people. Ming-yu Chen 0001, Alex Hauptmann 0001 |
ICASSP (3) | 2 |
| 2004 | Towards robust face recognition from multiple viewsabstractThe work presents a novel approach to aid face recognition: Using multiple views of a face, we construct a 3D model instead of directly using the 2D images for recognition. Our framework is designed for videos, which contain many instances of a target face from a sequence of slightly differing views, as opposed to a single static picture of the face. Specifically, we reconstruct the 3D face shapes from two orthogonal views and select features based on pairwise distances between landmark points on the model using Fisher's linear discriminant. While 3D face shape reconstruction is sensitive to the quality of the feature point localization, our experiments show that 3D reconstruction together with the regularized Fisher's linear discriminant can provide highly accurate face recognition from multiple facial views. Experiments on the Carnegie Mellon PIE (pose, illumination and expressions) database containing the faces of 68 people, with at least 3 expressions under varying lighting conditions, demonstrate vastly improved performance. Ming-yu Chen 0001, Alex Hauptmann 0001 |
ICME | 2 |
| 2004 | Comparison and combination of two novel commercial detection methodsabstractDetection and removal of commercials plays an important role when searching for important broadcast news video material. Two novel approaches are proposed based on two distinctive characteristics of commercials, namely, repetitive use of commercials over time and distinctive color and audio features. Furthermore, proposed strategies for combining the results of the two methods yield even better performance. Experiments show over 90% recall and precision on a test set of 5 hours of ABC and CNN broadcast news data. Pinar Duygulu, Ming-yu Chen 0001, Alex Hauptmann 0001 |
ICME | 3 |
| 2004 | Merging rank lists from multiple sources in video classificationabstractMultimedia corpora increasingly consist of data from multiple sources, with different characteristics that can be exploited by specialized applications. This paper focuses on video classification over multiple-source collections, and addresses the question whether classifiers should train from individual sources or from a full data set across all sources. If training separately, how can rank lists from different sources be merged effectively? We formulate the problem of merging ranked lists as learning a function mapping from local scores to global scores, and propose a learning method based on logistic regression. In our experiments we find that source characteristics are very important for video classification. Moreover, our method of learning mapping functions performs significantly better than merging methods without explicitly learning the mapping junctions. Weihao Lin 0003, Alex Hauptmann 0001 |
ICME | 2 |
| 2004 | Modeling timing features in broadcast news video classificationabstractBroadcast news programs are well-structured video, and timing can be a strong predictor for specific types of news reports. However, learning a classifier using timing features may not be an easy task when training data are noisy. We approach the problem from the generative model perspective, and approximate the class density in a non-parametric fashion. The results show that timing is a simple but extremely effective feature, and our method can achieve significantly better performance than a discriminative classifier. Weihao Lin 0003, Alex Hauptmann 0001 |
ICME | 2 |
| 2004 | Detection of TV news monologues by style analysisabstractWe propose a method for detection of semantic concepts in produced video based on style analysis. Recognition of concepts is done by applying a classifier ensemble to the detected style elements. As a case study we present a method for detecting the concept of news subject monologues. Our approach had the best average precision performance amongst 26 submissions in the 2003 TRECVID benchmark Cees Snoek, Marcel Worring, Alex Hauptmann 0001 |
ICME | 3 |
| 2004 | Multi-class active learning for video semantic feature extractionabstractActive learning has been demonstrated to be a useful tool to reduce human labeling effort for many multimedia applications, especially for those handling large video collections. However, most of the previous work on active learning has focused on only binary classification, which greatly limits the applicability of active learning. We present a multi-class active learning approach which extends active learning from binary classification to multi-class classification using a unified representation with margin-based loss functions. The experimental results on the TREC03 semantic feature extraction task shows that the proposed active learning approach works effectively even with a significantly reduced amount of labeled data. Alex Hauptmann 0001 |
ICME | 2 |
| 2004 | Successful approaches in the TREC video retrieval evaluationsabstractThis paper reviews successful approaches in evaluations of video retrieval over the last three years. The task involves the search and retrieval of shots from MPEG digitized video recordings using a combination of automatic speech, image and video analysis and information retrieval technologies. The search evaluations are grouped into interactive (with a human in the loop) and noninteractive (where the human merely enters the query into the system) submissions. Most non-interactive search approaches have relied extensively on text retrieval, and only recently have image-based features contributed reliably to improved search performance. Interactive approaches have substantially outperformed all non-interactive approaches, with most systems relying heavily on the user’s ability to refine queries and reject spurious answers. We will examine both the successful automatic search approaches and the user interface techniques that have enabled high performance video retrieval. Alex Hauptmann 0001, Michael G. Christel |
ACM Multimedia | 1 |
| 2004 | Learning query-class dependent weights in automatic video retrievalabstractCombining retrieval results from multiple modalities plays a crucial role for video retrieval systems, especially for automatic video retrieval systems without any user feedback and query expansion. However, most of current systems only utilize query independent combination or rely on explicit user weighting. In this work, we propose using query-class dependent weights within a hierarchial mixture-of-expert framework to combine multiple retrieval results. We first classify each user query into one of the four predefined categories and then aggregate the retrieval results with query-class associated weights, which can be learned from the development data efficiently and generalized to the unseen queries easily. Our experimental results demonstrate that the performance with query-class dependent weights can considerably surpass that with the query independent weights. Jun Yang 0003, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2004 | Naming every individual in news video monologuesabstractNaming every individual person appearing in broadcast news videos with names detected from the video transcript leads to better access of the news video content. In this paper, we approach this challenging problem with a statistical learning method. Two categories of information extracted from multiple video modalities have been explored, namely features, which help distinguish the true name of every person, as well as constraints, which reveal the relationships among the names of different persons. The person-naming problem is formulated into a learning framework which predicts the most likely name for each person based on the features, and refines the predictions using the constraints. Experiments conducted on ABC World New Tonight and CNN Headline News videos demonstrate that this approach outperforms a non-learning alternative by a large amount. Jun Yang 0003, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2003 | On predicting rare classes with SVM ensembles in scene classificationabstractScene classification is an important technique to infer high-level semantic scene categories from low-level visual features. However, in the real world the positive data for many scenes may be rare, which degrades the performance of many classifiers. In this paper, we propose SVM ensembles to address the rare class problem. Various classifier combination strategies are investigated, including majority voting, sum rule, neural network gater and hierarchical SVMs. We also compare our method with two other common approaches for dealing with the rare class problem. Our experimental results show that hierarchical SVMs can achieve significantly better and more stable performance than other strategies, as well as high computational efficiency. Yan Liu 0002, Rong Jin 0001, Alex Hauptmann 0001 |
ICASSP (3) | 4 |
| 2003 | Automatically Labeling Video Data Using Multi-class Active LearningabstractLabeling video data is an essential prerequisite for many vision applications that depend on training data, such as visual information retrieval, object recognition, and human activity modelling. However, manually creating labels is not only time-consuming but also subject to human errors, and eventually, becomes impossible for a very large amount of data (e.g. 24/7 surveillance video). To minimize the human effort in labeling, we propose a unified multiclass active learning approach for automatically labeling video data. We include extending active learning from binary classes to multiple classes and evaluating several practical sample selection strategies. The experimental results show that the proposed approach works effectively even with a significantly reduced amount of labeled data. The best sample selection strategy can achieve more than a 50% error reduction over random sample selection. Jie Yang 0001, Alex Hauptmann 0001 |
ICCV | 3 |
| 2003 | Learning to identify video shots with people based on face detectionabstractWe examine how to identify video shots with at least two humans using only detected face information. While face detection is much more reliable than shape based people classification in broadcast video, one particular difficulty is that, when there are several humans in an image, the accuracy of face detection is usually significantly degraded, which leads to poor performance in identifying shots of 'people'. Furthermore, while our standard face detector works from individual still images, we propose using the statistics of face information of images within a whole shot as additional evidence in deciding whether or not a video shot belongs to the 'people' category. Empirically, we studied which statistics of face information are more informative than others and how to combine different statistics together in order to achieve better prediction. Rong Jin 0001, Alex Hauptmann 0001 |
ICME | 2 |
| 2003 | Supervised classification for video shot segmentationabstractIn this paper, we explore supervised classification methods for video shot segmentation. We transform the temporal segmentation problem into a multi-class categorization issue. This approach provides a uniform framework for using different kinds of features extracted from the video and for detecting various types of shot boundaries. The approach utilizes manual labeled training data and a simple classification structure, which eliminates arbitrary thresholds and achieves more reliable estimation than previous threshold-based methods. Contrastive experiments on 13 videos (/spl sim/4 hours) show excellent performance on the 2001 TREC video track shot classification task in terms of precision and recall. Yanjun Qi, Alex Hauptmann 0001, Ting Liu 0005 |
ICME | 2 |
| 2003 | A Faster Iterative Scaling Algorithm for Conditional Exponential Model
Rong Jin 0001, Jian Zhang 0003, Alex Hauptmann 0001 |
ICML | 4 |
| 2003 | Modified Logistic Regression: An Approximation to SVM and Its Applications in Large-Scale Text Categorization
Jian Zhang 0003, Rong Jin 0001, Yiming Yang 0002, Alex Hauptmann 0001 |
ICML | 4 |
| 2003 | The combination limit in multimedia retrievalabstractCombining search results from multimedia sources is crucial for dealing with heterogeneous multimedia data, particularly in multimedia retrieval where a final ranked list of items of interest is returned sorted by confidence or relevance. However, relatively little attention has been given to combination functions, especially their upper bound performance limits. This paper presents a theoretical framework for studying upper bounds for two types of combination functions. A general upper bound and two approximations are proposed for monotonic combination functions. We also studied the upper bounds for linear combination functions using a global optimization technique. Our experimental results show that the choice of combination functions has a considerable influence to retrieval performance. Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2003 | Negative pseudo-relevance feedback in content-based video retrievalabstractVideo information retrieval requires a system to find information relevant to a query which may be represented simultaneously in different ways through a text description, audio, still images and/or video sequences. We present a novel approach that uses pseudo-relevance feedbackfrom retrieved items that are NOT similar to the query items without further inquiring user feedback. We provide insight into this approach using a statistical model and suggest a score combination scheme via posterior probability estimation. An evaluation on the 2002 TREC Video Trackqueries shows that this technique can improve video retrieval performance on a real collection. We believe that negative pseudo-relevance feedbackshows great promise for very difficult multimedia retrieval tasks, especially when combined with other different retrieval algorithms. 1. Alex Hauptmann 0001, Rong Jin 0001 |
ACM Multimedia | 2 |
| 2003 | Web Image Retrieval Re-Ranking with Relevance ModelabstractWeb image retrieval is a challenging task that requires efforts from image processing, link structure analysis, and Web text retrieval. Since content-based image retrieval is still considered very difficult, most current large-scale Web image search engines exploit text and link structure to "understand" the content of the Web images. However, local text information, such as caption, filenames and adjacent text, is not always reliable and informative. Therefore, global information should be taken into account when a Web image retrieval system makes relevance judgment. We propose a re-ranking method to improve Web image retrieval by reordering the images retrieved from an image search engine. The re-ranking process is based on a relevance model, which is a probabilistic model that evaluates the relevance of the HTML document linking to the image, and assigns a probability of relevance. The experiment results showed that the re-ranked image retrieval achieved better performance than original Web image retrieval, suggesting the effectiveness of the re-ranking method. The relevance model is learned from the Internet without preparing any training data and independent of the underlying algorithm of the image search engines. The re-ranking process should be applicable to any image search engines with little effort. Weihao Lin 0003, Rong Jin 0001, Alex Hauptmann 0001 |
Web Intelligence | 3 |
| 2002 | A New Probabilistic Model for Title Generation
Rong Jin 0001, Alex Hauptmann 0001 |
COLING | 2 |
| 2002 | Using a probabilistic source model for comparing imagesabstractWe propose a probabilistic model for image retrieval. To obtain the similarity between the query image IQ and any image I' in the collection, the model computes the probability of generating the image I' given the observation of the query image I/sub Q/. We compare our probabilistic model for image retrieval with a color histogram based image retrieval method and the IBM QBIC image search engine. The evaluation used the 11-hour video retrieval collection (80,000 extracted images) and associated queries from the 2001 TREC-10 information retrieval evaluations. The experimental results show that the probabilistic model dramatically outperforms the color histogram based image retrieval method and the IBM QBIC image search engine by 40%. Rong Jin 0001, Alex Hauptmann 0001 |
ICIP (3) | 2 |
| 2002 | Collages as dynamic summaries for news videoabstractThis paper introduces the video collage, a novel effective interface for browsing and interpreting video collections. The paper discusses how collages are automatically produced, illustrates their use, and evaluates their effectiveness as summaries across news stories. Collages are presentations of text and images derived from multiple video sources, which provide an interactive visualization for a set of video documents, summarizing their contents and providing a navigation aid for further exploration. The dynamic creation of collages is based on user context, e.g., an originating query, coupled with automatic processing to refine the candidate imagery. Named entity identification and common phrase extraction provides descriptive text. The dynamic manipulation of collages allows user-directed browsing and reveals additional detail. The utility of collages as summaries is examined with respect to other published news summaries. Michael G. Christel, Alex Hauptmann 0001, Howard D. Wactlar, Tobun Dorbin Ng |
ACM Multimedia | 2 |
| 2002 | News video classification using SVM-based multimodal classifiers and combination strategiesabstractVideo classification is the first step toward multimedia content understanding. When video is classified into conceptual categories, it is usually desirable to combine evidence from multiple modalities. However, combination strategies in previous studies were usually ad hoc. We investigate a meta-classification combination strategy using Support Vector Machine, and compare it with probability-based strategies. Text features from closed-captions and visual features from images are combined to classify broadcast news video. The experimental results show that combining multimodal classifiers can significantly improve recall and precision, and our meta-classification strategy gives better precision than the approach of taking the product of the posterior probabilities. Weihao Lin 0003, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2002 | Title language model for information retrievalabstractIn this paper, we propose a new language model, namely, a title language model, for information retrieval. Different from the traditional language model used for retrieval, we define the conditional probability P(QID) as the probability of using query Q as the title for document D. We adopted the statistical translation model learned from the title and document pairs in the collection to compute the probability P(QID). To avoid the sparse data problem, we propose two new smoothing methods. In the experiments with four different TREC document collections, the title language model for information retrieval with the new smoothing method outperforms both the traditional language model and the vector space model for IR significantly. Rong Jin 0001, Alex Hauptmann 0001, ChengXiang Zhai |
SIGIR | 2 |
| 2002 | Language model for IR using collection informationabstractInformation retrieval using meta data can be traced back to the early age of IR where documents are represented by the controlled vocabulary. In this paper, we explore the usage of meta-data information under the framework of language model. We present a new language model that is able to take advantage of the category information for documents to improve the retrieval accuracy. We compare the new language model with the traditional language model over the TREC4 dataset where the collection information for documents is obtained using the k-means clustering method. The new language model outperforms the traditional language model, which verifies our statement. Rong Jin 0001, Luo Si, Alex Hauptmann 0001, Jamie Callan |
SIGIR | 3 |
| 2001 | Title Generation Using a Training Corpus
Rong Jin 0001, Alex Hauptmann 0001 |
CICLing | 2 |
| 2001 | Learning to Select Good Title Words: An New Approach based on Reverse Information Retrieval
Rong Jin 0001, Alex Hauptmann 0001 |
ICML | 2 |
| 2001 | Title Generation for Machine-Translated Documents
Rong Jin 0001, Alex Hauptmann 0001 |
IJCAI | 2 |
| 2001 | Meta-scoring: Automatically Evaluating Term Weighting Schemes in IR without Precision-RecallabstractIn this paper, we present a method that can automatically evaluate performance of different term weighting schemes in information retrieval without resorting to precision-recall based on human relevance judgments. Specifically, the problem is: given two document-term matrixes generated from two different term weighting schemes, can we tell which term weighting scheme will performance better than the other? We propose a meta-scoring function, which takes as input the document-term matrix generated by some term weighting scheme and computes a goodness score from the document-term matrix. In our experiments, we found out that this score is highly correlated with the precision-recall measurement for all the collections and term weighting schema we tried. Thus, we conclude that our meta-scoring function can be a substitute for the precision-recall measurement that needs relevance judgments of human subject. Furthermore, this meta-scoring function is not limited only to text information retrieval can be applied to fields such as image and DNA retrieval. Rong Jin 0001, Christos Falusos, Alex Hauptmann 0001 |
SIGIR | 3 |
| 2000 | Title generation for spoken broadcast news using a training corpusabstractThe problem of title generation involves finding the essence of a document and expressing it in only a few words. The results of a query to the Informedia Digital Video Library are summarized through an automatically generated title for each retrieved news story. When the document is errorful, as with speech-recognized broadcast news stories, the title creation challenge becomes even greater. We implemented a set of title word selection strategies and evaluated them on an independent test corpus of 579 broadcast news documents, comparing manual transcription results to automatically recognized speech using the CMU Sphinx speech recognition system with a 64000-word broadcast news language model. Using a training collection of 21190 transcribed broadcast news stories, we trained several systems to produce appropriate title words, i.e. Naïve Bayesian approach with full vocabulary, Naïve Bayesian approach with limited vocabulary, nearest neighbor approach and extractive approach. The F1 results shows that the nearest neighbor approach is a quick and easy way of generating good titles for speech recognized documents (F1 = 15.2%), while a Nave Bayesian approach with limited vocabulary also does well on our F1 measure (F1 = 21.6%), which ignores word order in the titles. Overall, the results show that title generation for speech recognized news documents is possible at a level approaching the accuracy of titles generated for perfect text transcriptions. One surprising phenomenon is that extractive approach performances slightly better for speech recognized documents than for manual transcripts. Rong Jin 0001, Alex Hauptmann 0001 |
INTERSPEECH | 2 |
| 1999 | Selection for acoustic coverage from unlimited speech extracted from closed-captioned TVabstractGiven unlimited amounts of speech training data, it is desirable to predict informative subsets that will still improve the resulting acoustic model. We present a triphone frequency threshold measure for predicting informative subsets from vast amounts of speech. Results with single pass decoding show that acoustic models built from our selection-based speech set perform better than when trained on similar amounts of non-selected speech, and perform similar to models built from the original, larger amount of speech. 1. INTRODUCTION Large amounts of hand-transcribed speech training data are required to build or improve speech recognition systems using current technologies. Speech data collection and its expensive manual transcription have been a bottleneck for improving speech systems. LDC (Linguistic Data Consortium) has made available 200 hours of manually transcribed speech in 1998 and about 600 hours will be available in 1999. Additionally, we have demonstrated a technique for aut... Photina Jaeyun Jang, Alex Hauptmann 0001 |
EUROSPEECH | 2 |
| 1999 | Laughter extracted from television closed captions as speech recognizer training dataabstractClosed captions in television broadcasts, intended to aid the hearing impaired, also have potential as training data for speech-recognition software. Use of closed captions for automatic extraction of virtually unlimited training data has already been demonstrated [1]. This paper reports some preliminary work on the use of non-speech sound tokens included in closed captions to extract training data to augment a speech recognizer's repertoire of non-speech phonemes. A small experiment was performed to pinpoint laughter sounds in television news broadcasts using the Informedia Digital Video Library's retrieval capabilities, which automatically exploit closed captions. The snippets found were used to retrain a speech recognizer. A small test showed a small but significant gain in performance. In the future we plan to develop this approach into a fully automatic procedure for extracting training data for non-speech sounds. INTRODUCTION Many television broadcasts are accompanied by closed... Paul E. Kennedy, Alex Hauptmann 0001 |
EUROSPEECH | 2 |
| 1998 | Hierarchical cluster language modeling with statistical rule extraction for rescoring n-best hypotheses during speech decoding
Photina Jaeyun Jang, Alex Hauptmann 0001 |
ICSLP | 2 |
| 1998 | Speech Recognition for a Digital Video LibraryabstractThe standard method for making the full content of audio and video material searchable is to annotate it with human-generated meta-data that describes the content in a way that the search can understand, as is done in the creation of multimedia CD-ROMs. However, for the huge amounts of data that could usefully be included in digital video and audio libraries, the cost of producing this meta-data is prohibitive. In the Informedia Digital Video Library, the production of the meta-data supporting the library interface is automated using techniques derived from artificial intelligence (AI) research. By applying speech recognition together with natural language processing, information retrieval, and image analysis, an interface has been produced that helps users locate the information they want, and navigate or browse the digital video library more effectively. Specific interface components include automatic titles, filmstrips, video skims, word location marking, and representative frames for shots. Both the user interface and the information retrieval engine within Informedia are designed for use with automatically derived meta-data, much of which depends on speech recognition for its production. Some experimental information retrieval results will be given, supporting a basic premise of the Informedia project: That speech recognition generated transcripts can make multimedia material searchable. The Informedia project emphasizes the integration of speech recognition, image processing, natural language processing, and information retrieval to compensate for deficiencies in these individual technologies. © 1998 John Wiley & Sons, Inc. Michael Witbrock, Alex Hauptmann 0001 |
J. Am. Soc. Inf. Sci. | 2 |
| 1997 | Indexing and search of multimodal informationabstractThe Informedia Digital Library Project allows full content indexing and retrieval of text, audio and video material. The integration of speech recognition, image processing, natural language processing and information retrieval overcomes limits in each technology to create a useful system. In order to answer the question how good speech recognition has to be in order to be useful and usable for indexing and retrieving speech recognizer generated transcripts, some empirical evidence is presented that illustrates the degradation of information retrieval at different levels of speech accuracy. In our experiments, word error rates up to 25% did not significantly impact information retrieval and error rates of 50% still provided 85 to 95% of the recall and precision relative to fully accurate transcripts in the same retrieval system. Alex Hauptmann 0001, Howard D. Wactlar |
ICASSP | 1 |
| 1995 | Speech recognition in the Informedia Digital Video Library: uses and limitationsabstractIn principle, speech recognition technology can make any spoken data useful for library indexing and retrieval. The paper describes the Informedia Digital Video Library project and discusses how speech recognition is used for transcript creation from video, alignment with hand-generated transcripts, query interface and audio paragraph segmentation. The results show that speech recognition accuracy varies dramatically depending on the quality and type of data used. Our information retrieval experiments also show that reasonable recall and precision can be obtained with moderate speech recognition accuracy. Finally we discuss some active areas of speech research relevant to the digital video library problem. Alex Hauptmann 0001 |
ICTAI | 1 |
| 1995 | Speech for Multimedia Information RetrievalabstractNo abstract available. Alex Hauptmann 0001, Michael Witbrock, Alexander I. Rudnicky |
ACM Symposium on User Interface Software and Technology | 1 |
| 1995 | Demonstration of a Reading Coach that ListensabstractNo abstract available. Jack Mostow, Alex Hauptmann 0001, Steven F. Roth |
ACM Symposium on User Interface Software and Technology | 2 |
| 1994 | A Reading Coach that Listens: (Edited) Video Transcript
Jack Mostow, Alex Hauptmann 0001, Steven F. Roth, Matthew Kane, Adam Swift, Lin Lawrence Chase, Bob Weide |
AAAI | 2 |
| 1994 | A Prototype Reading Coach that Listens
Jack Mostow, Steven F. Roth, Alex Hauptmann 0001, Matthew Kane |
AAAI | 3 |
| 1993 | Towards a Reading Coach that Listens: Automated Detection of Oral Reading Errors
Jack Mostow, Alex Hauptmann 0001, Lin Lawrence Chase, Steven F. Roth |
AAAI | 2 |
| 1993 | SPEAKEZ: a first experiment in concatenation synthesis from a large corpus
Alex Hauptmann 0001 |
EUROSPEECH | 1 |
| 1993 | Speech recognition applied to reading assistance for children: a baseline language model
Alex Hauptmann 0001, Lin Lawrence Chase, Jack Mostow |
EUROSPEECH | 1 |
| 1993 | Gestures with Speech for Graphic Manipulation
Alex Hauptmann 0001, Paul McAvinney |
Int. J. Man Mach. Stud. | 1 |
| 1991 | From Syntax to Meaning in Natural Language Processing
Alex Hauptmann 0001 |
AAAI | 1 |
| 1991 | Models for evaluating interaction protocols in speech recognitionabstractRecognitionerrors complicate the assessment of speech systems.This paper presents a new approach to modeling spoken language interaction protocols, based on finite Markov chains.An interaction protocol, prescribed by the interface design, defines a set of primitive transaction steps and the order of their execut ion.The efficiency of an interface depends on the interaction protocol as well as the cost of each different transaction step.Markov chains provide a simple and computationally eflicient method for modeling errorful systems.They allow for detailed comparisons between different interaction protocols and between different modalities.The method is illustrated by application to example protocols. Alexander I. Rudnicky, Alex Hauptmann 0001 |
CHI | 2 |
| 1991 | JANUS: a speech-to-speech translation system using connectionist and symbolic processing strategiesabstractThe authors present JANUS, a speech-to-speech translation system that utilizes diverse processing strategies including dynamic programming, stochastic techniques, connectionist learning, and traditional AI knowledge representation approaches. JANUS translates continuously spoken English utterances into Japanese and German speech utterances. The overall system performance on a corpus of conference registration conversations is 87%. Two versions of JANUS are compared: one using a LR parser (JANUS 1) and one using a connectionist parser (JANUS 2). Performance results were mixed, with JANUS 1 deriving benefit from a tighter language model and JANUS 2 benefitting from greater flexibility.> Alex Waibel, Ajay N. Jain, Arthur E. McNair, Hiroaki Saito 0001, Alex Hauptmann 0001, Joe Tebelskis |
ICASSP | 5 |
| 1989 | Speech and gestures for graphic image manipulationabstractAn experiment was conducted with people using gestures and speech to manipulate graphic images on a computer screen. A human was substituted for the recognition devices. The analysis showed that people strongly prefer to use both gestures and speech for the graphics manipulation and that they intuitively use multiple hands and multiple fingers in all three dimensions. There was surprising uniformity and simplicity in the gestures and speech. The analysis of these results provides strong encouragement for future development of integrated multi-modal interaction systems. Alex Hauptmann 0001 |
CHI | 1 |
| 1989 | Layering Predictions: Flexible Use of Dialog Expectation in Speech Recognition
Sheryl R. Young, Wayne H. Ward, Alex Hauptmann 0001 |
IJCAI | 3 |
| 1988 | Using Dialog-Level Knowledge Sources to Improve Speech Recognition
Alex Hauptmann 0001, Sheryl R. Young, Wayne H. Ward |
AAAI | 1 |
| 1988 | Parsing spoken phrases despite missing wordsabstractThe authors compare the recognition accuracy obtained in forming sentence hypotheses using island-driven sentence parsers with parsers that hypothesize sentences in left-to-right fashion. Island-driven parsing algorithms are especially valuable in speech recognition systems because they can function more gracefully when not all of the correct words of an utterance were produced by the word hypothesizer. The inputs to both types of parsers consist of a lattice of candidate words, which are identified by their begin and end times, and the quality of the acoustic-phonetic match. Grammatical constraints are expressed by trigram models of sequences of lexical and semantic labels. The authors found that the island-driven parser produces parses with a higher percentage of correct words than the left-to-right parser is all cases considered. When the quality of the input lattices is extremely high, differences in parsing accuracy can be directly attributed to the superior ability of the island-driven parser to handle lattices with missing words. With lower-quality input, the accuracy of both types of parsers degrades, which is due to the creation of garden-path hypotheses and a lack of good words to serve as seeds for island formation.> Wayne H. Ward, Alex Hauptmann 0001, Richard M. Stern, Thomas Chanak |
ICASSP | 2 |
| 1988 | Talking to Computers: An Empirical Investigation
Alex Hauptmann 0001, Alexander I. Rudnicky |
Int. J. Man Mach. Stud. | 1 |
| 1987 | Sentence parsing with weak grammatical constraintsabstractThis paper compares the recognition accuracy obtained in forming sentence hypotheses using several parsers based on different types of weak statistical models of syntax and semantics. The inputs to the parsers were word hypotheses generated from simulated acoustic-phonetic labels. Grammatical constraints are expressed by trigram models of sequences of lexical or semantic labels, or by a finite-state network of the semantic labels. When the input to the parser is of high quality, the more restrictive trigram models were found to perform as well as or better than the finite-state language model. The more restrictive trigram and network models of language produce better recognition accuracy when all correct words are actually hypothesized, but strong constraints can degrade performance when many correct words are missing from the parser input. Richard M. Stern, Wayne H. Ward, Alex Hauptmann 0001, Juan Leon |
ICASSP | 3 |
| 1986 | Parsing Spoken Language: A Semantic Caseframe Approach
Philip J. Hayes, Alex Hauptmann 0001, Jaime G. Carbonell, Masaru Tomita |
COLING | 2 |
| 1986 | On quick word spotting techniquesabstractThis paper describes 2 efficiency enhancements for speaker-independent connected spoken word spotting. The first enhancement, LESS COST, uses phoneme concatenation to reduce the cost of computing the local distance between each reference and input pattern point. The second enhancement, Coarse DP refinement, reduces the cost of dynamic time warping for only a small error rate penalty. An experiment confirmed these techniques. Seiichi Nakagawa, Alex Hauptmann 0001, Masaru Tomita |
ICASSP | 2 |