Mohan Kankanhalli

dblp:09/3613 · also Mohan S. Kankanhalli · DBLP profile ↗
← Back
382ranked-venue papers
14as first author
118since 2021 · last 2026
0000-0002-4846-2015ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 268 · 8 first-author · 65 since 2021Artificial intelligence and machine learning · 99 · 4 first-author · 62 since 2021Computer networks · 33 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 27 · 7 since 2021Security and privacy · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 6Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Object-Centric Framework for Video Moment Retrieval
abstract
Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.
Yongkang Wong, Satoshi Yamazaki, Jianquan Liu, Mohan Kankanhalli
AAAI5
2026 Aggregating Diverse Cue Experts for AI-Generated Image Detection
abstract
The rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN), a novel framework that integrates different yet complementary cues as input. MCAN employs a mixture-of-encoders adapter to dynamically process these cues, enabling more adaptive and robust feature representation. Our cues include the input image itself, which represents the overall content, and high-frequency components that emphasize edge details. Additionally, we introduce a Chromatic Inconsistency (CI) cue, which normalizes intensity values and captures noise information introduced during the image acquisition process in real images, making these noise patterns more distinguishable from those in AI-generated content. Unlike prior methods, MCAN employs a multi-cue aggregation strategy, leveraging spatial, frequency, and chromaticity-based cues. These cues are intrinsically more indicative of real images, enhancing cross-model generalization. Extensive experiments on the GenImage, Chameleon, and UniversalFakeDetect benchmark validate the state-of-the-art performance of MCAN. In the GenImage dataset, MCAN outperforms the best state-of-the-art method by up to 7.4\% in average ACC across eight different image generators.
Shuwei Li, Mohan Kankanhalli, Robby T. Tan
AAAI3
2026 Editorial to special issue on selected extended works from 9th international conference on computer vision & image processing (CVIP) 2024
Mohan Kankanhalli, Balasubramanian Raman, M. Subrahmanyam 0001, Jagadeesh Kakarla, Sambit Bakshi
Image Vis. Comput.1
2026 TailorEdit: An Adaptive Framework for Instruction-Guided Fashion Image Editing
abstract
Fashion image editing has garnered significant attention due to its growing demand in e-commerce, social media, and virtual try-on applications. However, existing methods are typically designed for specific editing tasks in isolation, lacking a unified framework capable of handling diverse editing requirements. This work addresses this limitation from two critical perspectives. First, we constructInstructFashion, a large-scale, high-quality dataset specifically curated for instruction-guided fashion image editing. It is generated through carefully designed pipelines that cover four distinct editing tasks. Second, we proposeTailorEdit, an adaptive framework for instruction-guided fashion image editing. It integrates human segmentation map-based denoising guidance, modular LoRA-based editing experts, and a dynamic expert routing mechanism to enable precise and semantically coherent modifications. Extensive quantitative and qualitative evaluations demonstrate that TailorEdit consistently outperforms state-of-the-art methods in terms of realism, coherence, and instruction adherence. Our code is available at https://github.com/EndaJude/TailorEdit.
Xiaoling Gu, Lingda Zhu, Yongkang Wong, Zhou Yu 0001, Huan Li 0003, Zizhao Wu, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.7
2026 Distill to Delete: Unlearning in Graph Networks With Knowledge Distillation
abstract
Graph unlearning has emerged as a pivotal method to delete information from an already trained graph neural network (GNN). One may delete nodes, a class of nodes, edges, or a class of edges. An unlearning method enables the GNN model to comply with data protection regulations (i.e., the right to be forgotten), adapt to evolving data distributions, and reduce the GPU-hours carbon footprint by avoiding repetitive retraining. Removing specific graph elements from graph data is challenging due to the inherent intricate relationships and neighborhood dependencies. Existing partitioning and aggregation-based methods have limitations due to their poor handling of local graph dependencies and additional overhead costs. Our work takes a novel approach to address these challenges in graph unlearning through knowledge distillation, as it distills to delete in GNN (D2DGN). It is an efficient model-agnostic distillation framework where the complete graph knowledge is divided and marked for retention and deletion. It performs distillation with response-based soft targets and feature-based node embedding while minimizing KL-divergence. The unlearned model effectively removes the influence of the deleted graph elements while preserving knowledge about the retained graph elements. D2DGN surpasses the performance of existing methods when evaluated on various real-world graph datasets by up to $\mathbf {43.1\%}$ (AUC) in edge and node unlearning tasks. Other notable advantages include better efficiency, better performance in removing target elements, preservation of performance for the retained elements, and zero overhead costs. Source code: https://github.com/MachineUnlearn/D2DGN.
Yash Sinha, Murari Mandal, Mohan Kankanhalli
IEEE Trans. Neural Networks Learn. Syst.3
2026 Towards Generalizable Deepfake Detection by Primary Region Regularization
abstract
The existing deepfake detection methods have reached a bottleneck in generalizing to unseen forgeries and manipulation approaches. Based on the observation that the deepfake detectors exhibit a preference for overfitting specific primary regions in input, this article enhances the generalization capability from a novel regularization perspective. This can be simply achieved by augmenting the images through primary region removal, thereby preventing the detector from over-relying on data bias. Our method consists of two stages, namely the static localization for primary region maps, as well as the dynamic exploitation of primary region masks. The proposed method can be seamlessly integrated into different backbones without affecting their inference efficiency. We conduct extensive experiments over five widely used deepfake datasets—DFDC, DF-1.0, Celeb-DF, WildDF, and FFIW with seven backbones. Our method demonstrates an average performance improvement of 6% across different backbones and performs competitively with several state-of-the-art baselines.
Harry Cheng 0002, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.5
2026 3D Scenes Motion Planning and Generation with Motion Diffusion Probabilistic Model
abstract
Generating natural and realistic human motion sequences under the constraints of 3D scenes is a highly challenging task, requiring not only the precise modeling of dynamic variations in human joints but also the rigorous consideration of intricate interactions between the human body and the surrounding environment. While recent advances in deep generative models show great potential in tackling these challenges, existing methods often result in unnatural human motions and human–environment penetration during generation. In order to cope with these issues, we propose a novel approach that divides human motion generation into two stages. The first stage employs a bidirectional long short-term memory network incorporated with full-connected layers to generate motion trajectory under the input conditions including the starting and ending positions and orientations of the human model and scene feature point clouds extracted from the surrounding environment. In the second stage, we design a conditional diffusion model, guided by the trajectory generated in the first stage and the embedding of 3D scene information, to generate human motion sequences within 3D scenes. We evaluate our framework through extensive experiments on the PROX datasets, which validates its effectiveness. The results show that our method significantly outperforms existing ones in enhancing human motion naturalness and reasonableness, and reducing human penetration.
Yubao Sun, Guiyu Xia, Qingshan Liu 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.5
2026 Introduction to the Special Issue on Multimodal Video Understanding and Analysis with Foundation Models
Fan Liu 0008, Hanjia Lyu, Yinwei Wei, Hehe Fan, Djamila Aouada, Jiebo Luo 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Multi-Modal Recommendation Unlearning for Legal, Licensing, and Modality Constraints
abstract
User data spread across multiple modalities has popularized multi-modal recommender systems (MMRS). They recommend diverse content such as products, social media posts, TikTok reels, etc., based on a user-item interaction graph. With rising data privacy demands, recent methods propose unlearning private user data from uni-modal recommender systems (RS). However, methods for unlearning item data related to outdated user preferences, revoked licenses, and legally requested removals are still largely unexplored. Previous RS unlearning methods are unsuitable for MMRS due to the incompatibility of their matrix-based representation with the multi-modal user-item interaction graph. Moreover, their data partitioning step degrades performance on each shard due to poor data heterogeneity and requires costly performance aggregation across shards. This paper introduces MMRecUn, the first approach known to us for unlearning in MMRS and unlearning item data. Given a trained RS model, MMRecUn employs a novel Reverse Bayesian Personalized Ranking (BPR) objective to enable the model to forget marked data. The reverse BPR attenuates the impact of user-item interactions within the forget set, while the forward BPR reinforces the significance of user-item interactions within the retain set. Our experiments demonstrate that MMRecUn outperforms baseline methods across various unlearning requests when evaluated on benchmark MMRS datasets. MMRecUn achieves recall performance improvements of up to 49.85% compared to baseline methods and is up to 1.3× faster than the Gold model, which is trained on retain set from scratch. MMRecUn offers significant advantages, including superiority in removing target interactions, preserving retained interactions, and zero overhead costs compared to previous methods.
Yash Sinha, Murari Mandal, Mohan Kankanhalli
AAAI3
2025 Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models
abstract
Data synthesis has become a crucial research area in large language models (LLMs), especially for generating high-quality instruction fine-tuning data to enhance downstream performance.In code generation, a key application of LLMs, manual annotation of code instruction data is costly.Recent methods, such as Code Evol-Instruct and OSS-Instruct, leverage LLMs to synthesize large-scale code instruction data, significantly improving LLM coding capabilities.However, these approaches face limitations due to unidirectional synthesis and randomness-driven generation, which restrict data quality and diversity.To overcome these challenges, we introduce Tree-of-Evolution (ToE), a novel framework that models code instruction synthesis process with a tree structure, exploring multiple evolutionary paths to alleviate the constraints of unidirectional generation.Additionally, we propose optimizationdriven evolution, which refines each generation step based on the quality of the previous iteration.Experimental results across five widely-used coding benchmarks-HumanEval, MBPP, EvalPlus, LiveCodeBench, and Big-CodeBench-demonstrate that base models fine-tuned on just 75k data synthesized by our method achieve comparable or superior performance to the state-of-the-art open-weight Code LLM, Qwen2.5-Coder-Instruct, which was finetuned on millions of samples.
Hongzhan Lin 0001, Mohan Kankanhalli, Jing Ma 0004
ACL (1)5
2025 VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
abstract
Large multimodal models (LMMs) with advanced video analysis capabilities have recently garnered significant attention. However, most evaluations rely on traditional methods like multiple-choice question answering in benchmarks such as VideoMME and LongVideoBench, which are prone to lack the depth needed to capture the complex demands of real-world users. To address this limitation—and due to the prohibitive cost and slow pace of human annotation for video tasks—we introduce VideoAutoArena, an arena-style benchmark inspired by LMSYS Chatbot Arena’s framework, designed to automatically assess LMMs’ video analysis abilities. VideoAutoArena utilizes user simulation to generate open-ended, adaptive questions that rigorously assess model performance in video understanding. The benchmark features an automated, scalable evaluation framework, incorporating a modified ELO Rating System for fair and continuous comparisons across multiple LMMs. To validate our automated judging system, we construct a "gold standard" using a carefully curated subset of human annotations, demonstrating that our arena strongly aligns with human judgment while maintaining scalability. Additionally, we introduce a fault-driven evolution strategy, progressively increasing question complexity to push models toward handling more challenging video analysis scenarios. Experimental results demonstrate that VideoAutoArena effectively differentiates among state-of-the-art LMMs, providing insights into model strengths and areas for improvement. To further streamline our evaluation, we introduce VideoAutoBench as an auxiliary benchmark, where human annotators label winners in a subset of VideoAutoArena battles. We use GPT-4o as a judge to compare responses against these human-validated answers. Together, VideoAu-Toarena and VideoAutoBench offer a cost-effective, and scalable framework for evaluating LMMs in user-centric video analysis.
Dongxu Li 0003, Jing Ma 0004, Mohan Kankanhalli, Junnan Li 0001
CVPR5
2025 Joint Vision-Language Social Bias Removal for CLIP
abstract
Vision-Language (V-L) pre-trained models such as CLIP show promising capabilities in various downstream tasks. Despite this promise, V-L models are notoriously limited by their inherent social biases. A typical demonstration is that V-L models often produce biased predictions against specific groups of people, significantly undermining their real-world applicability. Existing approaches endeavor to mitigate the social bias problem in V-L models by removing biased attribute information from model embeddings. However, after our revisiting of these methods, we find that their bias removal is frequently accompanied by greatly compromised V-L alignment capabilities. We then reveal that this performance degradation stems from the unbalanced debiasing in image and text embeddings. To address this issue, we propose a novel V-L debiasing framework to align image and text biases followed by removing them from both modalities. By doing so, our method achieves multi-modal bias mitigation while maintaining the V-L alignment in the debiased embeddings. Additionally, we advocate a new evaluation protocol that can 1) holistically quantify the model debiasing and V-L alignment ability, and 2) evaluate the generalization of social bias removal models. We believe this work will offer new insights and guidance for future studies addressing the social bias problem in CLIP. Our code can be found at https://github.com/haoyusimon/VL_Debiasing.
Mohan Kankanhalli
CVPR3
2025 SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency
abstract
While contrastive pre-training is widely employed, its data efficiency problem has remained relatively under-explored thus far. Existing methods often rely on static coreset selection algorithms to pre-identify important data for training. However, this static nature renders them unable to dynamically track the data usefulness throughout pre-training, leading to subpar pre-trained models. To address this challenge, our paper introduces a novel dynamic bootstrapping dataset pruning method. It involves pruning data preparation followed by dataset mutation operations, both of which undergo iterative and dynamic updates. We apply this method to two prevalent contrastive pre-training frameworks: \textbf{CLIP} and \textbf{MoCo}, representing vision-language and vision-centric domains, respectively. In particular, we individually pre-train seven CLIP models on two large-scale image-text pair datasets, and two MoCo models on the ImageNet dataset, resulting in a total of 16 pre-trained models. With a data pruning rate of 30-35\% across all 16 models, our method exhibits only marginal performance degradation (less than \textbf{1\%} on average) compared to corresponding models trained on the full dataset counterparts across various downstream datasets, and also surpasses several baselines with a large performance margin. Additionally, the byproduct from our method, \ie coresets derived from the original datasets after pre-training, also demonstrates significant superiority in terms of downstream performance over other static coreset selection approaches.
Mohan Kankanhalli
ICCV2
2025 Strong Preferences Affect the Robustness of Preference Models and Value Alignment
abstract
Value alignment, which aims to ensure that large language models (LLMs) and other AI agents behave in accordance with human values, is critical for ensuring safety and trustworthiness of these systems. A key component of value alignment is the modeling of human preferences as a representation of human values. In this paper, we investigate the robustness of value alignment by examining the sensitivity of preference models. Specifically, we ask: how do changes in the probabilities of some preferences affect the predictions of these models for other preferences? To answer this question, we theoretically analyze the robustness of widely used preference models by examining their sensitivities to minor changes in preferences they model. Our findings reveal that, in the Bradley-Terry and the Placket-Luce model, the probability of a preference can change significantly as other preferences change, especially when these preferences are dominant (i.e., with probabilities near zero or one). We identify specific conditions where this sensitivity becomes significant for these models and discuss the practical implications for the robustness and safety of value alignment in AI systems.
Ziwei Xu 0001, Mohan Kankanhalli
ICLR2
2025 GroMo25: ACM Multimedia 2025 Grand Challenge for Plant Growth Modeling with Multiview Images
Shreya Bansal, Ruchi Bhatt, Amanpreet Chander, Malya Singh, Mohan Kankanhalli, Abdulmotaleb El Saddik, Mukesh Saini
ACM Multimedia6
2025 FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
abstract
Proactive Deepfake detection via robust watermarks has seen interest ever since passive Deepfake detectors encountered challenges in identifying high-quality synthetic images. However, while demonstrating reasonable detection performance, they lack localization functionality and explainability in detection results. Additionally, the unstable robustness of watermarks can significantly affect the detection performance. In this study, we propose novel fractal watermarks for proactive Deepfake detection and localization, namely FractalForensics. Benefiting from the characteristics of fractals, we devise a parameter-driven watermark generation pipeline that derives fractal-based watermarks and performs one-way encryption of the selected parameters. Subsequently, we propose a semi-fragile watermarking framework for watermark embedding and recovery, trained to be robust against benign image processing operations and fragile when facing Deepfake manipulations in a black-box setting. Moreover, we introduce an entry-to-patch strategy that implicitly embeds the watermark matrix entries into image patches at corresponding positions, achieving localization of Deepfake manipulations. Extensive experiments demonstrate satisfactory robustness and fragility of our approach against common image processing operations and Deepfake manipulations, outperforming state-of-the-art semi-fragile watermarking algorithms and passive detectors for Deepfake detection. Furthermore, by highlighting the areas manipulated, our method provides explainability for the proactive Deepfake detection results.
Tianyi Wang 0006, Harry Cheng 0002, Minghui Liu 0001, Mohan Kankanhalli
ACM Multimedia4
2025 A New Dataset and Benchmark for Grounding Multimodal Misinformation
abstract
The proliferation of online misinformation videos poses serious societal risks. Current datasets and detection methods primarily target binary classification or single-modality localization based on post-processed data, lacking the interpretability needed to counter persuasive misinformation. In this paper, we introduce the task of Grounding Multimodal Misinformation (GroundMM), which verifies multimodal content and localizes misleading segments across modalities. We present the first real-world dataset for this task, GroundLie360, featuring a taxonomy of misinformation types, fine-grained annotations across text, speech, and visuals, and validation with Snopes evidence and annotator reasoning. We also propose a VLM-based, QA-driven baseline, FakeMark, using single and cross-modal cues for effective detection and grounding. Our experiments highlight the challenges of this task and lay a foundation for explainable multimodal misinformation detection. Dataset will be released at https://github.com/yangbingjian/GroundLie360.
Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu 0008, Zheng Wang 0007, Mohan Kankanhalli
ACM Multimedia6
2025 Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
Yu-Wei Zhan, Fan Liu 0008, Xin Luo 0006, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia6
2025 Fair Deepfake Detectors Can Generalize
abstract
Deepfake detection models face two critical challenges: generalization to unseen manipulations and demographic fairness among population groups. However, existing approaches often demonstrate that these two objectives are inherently conflicting, revealing a trade-off between them. In this paper, we, for the first time, uncover and formally define a causal relationship between fairness and generalization. Building on the back-door adjustment, we show that controlling for confounders (data distribution and model capacity) enables improved generalization via fairness interventions. Motivated by this insight, we propose Demographic Attribute-insensitive Intervention Detection (DAID), a plug-and-play framework composed of: i) Demographic-aware data rebalancing, which employs inverse-propensity weighting and subgroup-wise feature normalization to neutralize distributional biases; and ii) Demographic-agnostic feature aggregation, which uses a novel alignment loss to suppress sensitive-attribute signals. Across three cross-domain benchmarks, DAID consistently achieves superior performance in both fairness and generalization compared to several state-of-the-art detectors, validating both its theoretical foundation and practical effectiveness.
Harry Cheng 0002, Minghui Liu 0001, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
NeurIPS6
2025 The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
abstract
The vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise. However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evaluations, often with minimal effort. This dual high performance in both attack and defense gives rise to a fundamental and perplexing paradox. To gain a deep understanding of this issue and thus further help strengthen the trustworthiness of VLLMs, this paper makes three key contributions: i) One tentative explanation for VLLMs being prone to jailbreak attacks--inclusion of vision inputs, as well as its in-depth analysis. ii) The recognition of a largely ignored problem in existing VLLM defense mechanisms--over-prudence. The problem causes these defense methods to exhibit unintended abstention, even in the presence of benign inputs, thereby undermining their reliability in faithfully defending against attacks. iii) A simple safety-aware method--LLM-Pipeline. Our method repurposes the more advanced guardrails of LLMs on the fly, serving as an effective alternative detector prior to VLLM response. Last but not least, we find that the two representative evaluation methods for jailbreak often exhibit chance agreement. This limitation makes it potentially misleading when evaluating attack strategies or defense mechanisms. We believe the findings from this paper offer useful insights to rethink the foundational development of VLLM safety with respect to benchmark datasets, defense strategies, and evaluation methods.
Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli
NeurIPS4
2025 Image-Based Virtual Try-On: A Survey
Dan Song 0006, Xuanpu Zhang, Weizhi Nie, Ruofeng Tong 0001, Mohan Kankanhalli, Anan Liu
Int. J. Comput. Vis.6
2025 Learning to Predict Gradients for Semi-Supervised Continual Learning
abstract
A key challenge for machine intelligence is to learn new visual concepts without forgetting the previously acquired knowledge. Continual learning (CL) is aimed toward addressing this challenge. However, there still exists a gap between CL and human learning. In particular, humans are able to continually learn from the samples associated with known or unknown labels in their daily lives, whereas existing CL and semi-supervised CL (SSCL) methods assume that the training samples are associated with known labels. Specifically, we are interested in two questions: 1) how to utilize unrelated unlabeled data for the SSCL task and 2) how unlabeled data affect learning and catastrophic forgetting in the CL task. To explore these issues, we formulate a new SSCL method, which can be generically applied to existing CL models. Furthermore, we propose a novel gradient learner to learn from labeled data to predict gradients on unlabeled data. In this way, the unlabeled data can fit into the supervised CL framework. We extensively evaluate the proposed method on mainstream CL methods, adversarial CL (ACL), and semi-supervised learning (SSL) tasks. The proposed method achieves state-of-the-art performance on classification accuracy and backward transfer (BWT) in the CL setting while achieving the desired performance on classification accuracy in the SSL setting. This implies that the unlabeled images can enhance the generalizability of CL models on the predictive ability of unseen data and significantly alleviate catastrophic forgetting. The code is available at https://github.com/luoyan407/grad_prediction.git.
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Understanding Before Recommendation: Semantic Aspect-Aware Review Exploitation via Large Language Models
abstract
Recommendation systems harness user–item interactions like clicks and reviews to learn their representations. Previous studies improve recommendation accuracy and interpretability by modeling user preferences across various aspects and intents. However, the aspects and intents are inferred directly from user reviews or behavior patterns, suffering from the data noise and the data sparsity problem. Furthermore, it is difficult to understand the reasons behind recommendations due to the challenges of interpreting implicit aspects and intents. To address these constraints, we harness the sentiment analysis capabilities of Large Language Models (LLMs) to enhance the accuracy and interpretability of the conventional recommendation methods. Specifically, inspired by the deep semantic understanding offered by LLMs, we introduce a chain-based prompting strategy to uncover semantic aspect-aware interactions, which provide clearer insights into user behaviors at a fine-grained semantic level. To incorporate the rich interactions of various aspects, we propose the simple yet effective Semantic Aspect-Based Graph Convolution Network (SAGCN). By performing graph convolutions on multiple semantic aspect graphs, SAGCN efficiently combines embeddings across multiple semantic aspects for final user and item representations. The effectiveness of the SAGCN was evaluated on four publicly available datasets through extensive experiments, which revealed that it outperforms all other competitors. Furthermore, interpretability analysis experiments were conducted to demonstrate the interpretability of incorporating semantic aspects into the model.
Fan Liu 0008, Huilin Chen 0002, Zhiyong Cheng 0001, Liqiang Nie, Mohan Kankanhalli
ACM Trans. Inf. Syst.6
2025 Implications of Privacy Regulations on Video Surveillance Systems
abstract
Advanced video surveillance systems (VSS), which collect information of every individual who passes through a surveilled area, have become ubiquitous due to its utility for security. However, such proactive monitoring threatens the individual's privacy due to the public's lack of control over personal data. Additionally, individuals or organizations may unethically misuse VSS for other purposes (e.g., individual profiling and unwarranted monitoring). To safeguard individual privacy, various governments have introduced mandatory information privacy regulations (e.g., GDPR, PDPA, and CCPA) to provide extensive guidelines for the purpose of achieving identity confidentiality. Currently, there is a gap between the information privacy regulations and VSS. This article aims to bridge this gap through four contributions. First, this article conceptualizes VSS as comprising various data stages based on the idea of data lifecycle and studies the implications of existing regulations on VSS. Second, we conducted a survey in ASEAN and European regions to understand the public perception of data risks at each data stage. Third, we review existing privacy-enhancing technologies and its relation to each data stage. Finally, we discuss open research problems in order to realize privacy-aware VSS.
Kajal Kansal, Yongkang Wong, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2024 PELA: Learning Parameter-Efficient Models with Low-Rank Approximation
abstract
Applying a pre-trained large model to downstream tasks is prohibitive under resource-constrained conditions. Recent dominant approaches for addressing efficiency issues involve adding a few learnable parameters to the fIxed backbone model. This strategy, however, leads to more challenges in loading large models for downstream finetuning with limited resources. In this paper, we propose a novel method for increasing the parameter efficiency of pretrained models by introducing an intermediate pre-training stage. To this end, we first employ low-rank approximation to compress the original large model and then devise a feature distillation module and a weight perturbation regularization module. These modules are specifically designed to enhance the low-rank model. In particular, we update only the low-rank model while freezing the backbone parameters during pre-training. This allows for direct and efficient utilization of the low-rank model for downstream finetuning tasks. The proposed method achieves both efficiencies in terms of required parameters and computation time while maintaining comparable results with minimal modifications to the backbone architecture. Specifically, when applied to three vision-only and one vision-language Transformer models, our approach often demonstrates a merely rvO.6 point decrease in performance while reducing the original parameter size by 1/3 to 2/3. We release our code at link.
Guangzhi Wang, Mohan Kankanhalli
CVPR3
2024 Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-Robustness
abstract
Human-Object Interaction (HOI) Detection constitutes an important aspect of human-centric scene understanding, which requires precise object detection and interaction recognition. Despite increasing advancement in detection, recognizing subtle and intricate interactions remains challenging. Recent methods have endeavored to leverage the rich semantic representation from pretrained CLIP, yet fail to efficiently capture finer-grained spatial features that are highly informative for interaction discrimination. In this work, instead of solely using representations from CLIP, we fill the gap by proposing a spatial adapter that efficiently utilizes the multi-scale spatial information in the pretrained detector. This leads to a bilateral adaptation that mutually produces complementary features. To further improve interaction recognition under occlusion, which is common in crowded scenarios, we propose an Occluded Part Extrapolation module that guides the model to recover the spatial details from manually occluded feature maps. Moreover, we design a Conditional Contextual Mining module that further mines informative contextual clues from the spatial features via a tailored cross-attention mechanism. Extensive experiments on V-COCO and HICO-DET benchmarks demonstrate that our method significantly outperforms prior art on both standard and zero-shot settings, resulting in new state-of-the-art performance. Additional ablation studies further validate the effectiveness of each component in our method.
Guangzhi Wang, Ziwei Xu 0001, Mohan Kankanhalli
CVPR4
2024 Finetuning Text-to-Image Diffusion Models for Fairness
abstract
The rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for minority groups. In this work, we frame fairness as a distributional alignment problem. Our solution consists of two main technical contributions: (1) a distributional alignment loss that steers specific characteristics of the generated images towards a user-defined target distribution, and (2) adjusted direct finetuning of diffusion model's sampling process (adjusted DFT), which leverages an adjusted gradient to directly optimize losses defined on the generated images. Empirically, our method markedly reduces gender, racial, and their intersectional biases for occupational prompts. Gender bias is significantly reduced even when finetuning just five soft tokens. Crucially, our method supports diverse perspectives of fairness beyond absolute equality, which is demonstrated by controlling age to a 75% young and 25% old distribution while simultaneously debiasing gender and race. Finally, our method is scalable: it can debias multiple concepts at once by simply including these prompts in the finetuning data. We share code and various fair diffusion model adaptors at https://sail-sg.github.io/finetune-fair-diffusion/.
Tianyu Pang, Yongkang Wong, Mohan Kankanhalli
ICLR6
2024 An LLM can Fool Itself: A Prompt-Based Adversarial Attack
abstract
The wide-ranging applications of large language models (LLMs), especially in safety-critical domains, necessitate the proper evaluation of the LLM’s adversarial robustness. This paper proposes an efficient tool to audit the LLM’s adversarial robustness via a prompt-based adversarial attack (PromptAttack). PromptAttack converts adversarial textual attacks into an attack prompt that can cause the victim LLM to output the adversarial sample to fool itself. The attack prompt is composed of three important components: (1) original input (OI) including the original sample and its ground-truth label, (2) attack objective (AO) illustrating a task description of generating a new sample that can fool itself without changing the semantic meaning, and (3) attack guidance (AG) containing the perturbation instructions to guide the LLM on how to complete the task by perturbing the original sample at character, word, and sentence levels, respectively. Besides, we use a fidelity filter to ensure that PromptAttack maintains the original semantic meanings of the adversarial examples. Further, we enhance the attack power of PromptAttack by ensembling adversarial examples at different perturbation levels. Comprehensive empirical results using Llama2 and GPT-3.5 validate that PromptAttack consistently yields a much higher attack success rate compared to AdvGLUE and AdvGLUE++. Interesting findings include that a simple emoji can easily mislead GPT-3.5 to make wrong predictions. Our source code is available at https://github.com/GodXuxilie/PromptAttack.
Xilie Xu, Keyi Kong, Ning Liu 0014, Li-Zhen Cui 0001, Di Wang 0015, Jingfeng Zhang, Mohan Kankanhalli
ICLR7
2024 AutoLoRa: An Automated Robust Fine-Tuning Framework
abstract
Robust Fine-Tuning (RFT) is a low-cost strategy to obtain adversarial robustness in downstream applications, without requiring a lot of computational resources and collecting significant amounts of data. This paper uncovers an issue with the existing RFT, where optimizing both adversarial and natural objectives through the feature extractor (FE) yields significantly divergent gradient directions. This divergence introduces instability in the optimization process, thereby hindering the attainment of adversarial robustness and rendering RFT highly sensitive to hyperparameters. To mitigate this issue, we propose a low-rank (LoRa) branch that disentangles RFT into two distinct components: optimizing natural objectives via the LoRa branch and adversarial objectives via the FE. Besides, we introduce heuristic strategies for automating the scheduling of the learning rate and the scalars of loss terms. Extensive empirical evaluations demonstrate that our proposed automated RFT disentangled via the LoRa branch (AutoLoRa) achieves new state-of-the-art results across a range of downstream tasks. AutoLoRa holds significant practical utility, as it automatically converts a pre-trained FE into an adversarially robust model for downstream tasks without the need for searching hyperparameters. Our source code is available at [the GitHub](https://github.com/GodXuxilie/RobustSSL_Benchmark/tree/main/Finetuning_Methods/AutoLoRa).
Xilie Xu, Jingfeng Zhang, Mohan Kankanhalli
ICLR3
2024 Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning
abstract
Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending the context of vision and language, we introduce Multimodal Composition Learning (MCL) for the purpose of mapping or aligning the vision and language input. In particular, we introduce two tasks: Multimodal-Context Captioning (MC-Cap) and Multimodal-Context Retrieval (MC-Ret) to guide a frozen LLM in comprehending the vision and language context. These specialized tasks are crafted to improve the LLM’s capacity for efficient processing and utilization of multimodal inputs, thereby enhancing its proficiency in generating more accurate text or visual representations. Extensive experiments on both retrieval tasks (i.e., zero-shot composed image retrieval, visual storytelling image retrieval and visual dialog image retrieval) and text generation tasks (i.e., visual question answering) demonstrate the effectiveness of the proposed method. The code is available at: https://github.com/dhg-wei/MCL.
Hehe Fan, Yongkang Wong, Yi Yang 0001, Mohan Kankanhalli
ICML5
2024 MCM: Multi-condition Motion Synthesis Framework
Zeyu Ling, Bo Han 0003, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
IJCAI5
2024 EcoVal: An Efficient Data Valuation Framework for Machine Learning
abstract
Quantifying the value of data within a machine learning workflow can play a pivotal role in making more strategic decisions in machine learning initiatives. The existing Shapley value based frameworks for data valuation in machine learning are computationally expensive as they require considerable amount of repeated training of the model to obtain the Shapley value. In this paper, we introduce an efficient data valuation framework EcoVal, to estimate the value of data for machine learning models in a fast and practical manner. Instead of directly working with individual data sample, we determine the value of a cluster of similar data points. This value is further propagated amongst all the member cluster points. We show that the overall value of the data can be determined by estimating the intrinsic and extrinsic value of each data. This is enabled by formulating the performance of a model as aproduction function, a concept which is popularly used to estimate the amount of output based on factors like labor and capital in a traditional free economic market. We provide a formal proof of our valuation technique and elucidate the principles and mechanisms that enable its accelerated performance. We demonstrate the real-world applicability of our method by showcasing its effectiveness for both in-distribution and out-of-sample data. This work addresses one of the core challenges of efficient data valuation at scale in machine learning models. The code is available at https://github.com/respai-lab/ecoval.
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan, Bowei Chen 0001, Mohan Kankanhalli
KDD6
2024 Diffusion Facial Forgery Detection
abstract
Detecting diffusion-generated images has recently developed as an emerging research area. Existing diffusion-based datasets predominantly focus on general image generation. However, facial forgeries, which pose severe social risks, have remained less explored thus far. To address this gap, this paper introduces DiFF, a comprehensive dataset dedicated to face-focused diffusion-generated images. DiFF comprises over 500,000 images that are synthesized using thirteen distinct generation methods under four conditions. In particular, this dataset utilizes 30,000 carefully collected textual and visual prompts, ensuring the synthesis of images with both high fidelity and semantic consistency. We conduct extensive experiments on the DiFF dataset via human subject tests and several representative forgery detection methods. The results demonstrate that the binary detection accuracies of both human observers and automated detectors often fall below 30%, revealing insights on the challenges in detecting diffusion-generated facial forgeries. Moreover, our experiments demonstrate that DiFF, compared to previous facial forgery datasets, contains a more diverse and realistic range of forgeries, showcasing its potential to aid in the development of more generalized detectors. Finally, we propose an edge graph regularization approach to effectively enhance the generalization capability of existing detectors.
Harry Cheng 0002, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia5
2024 Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
abstract
Recommendation algorithms predict user preferences by correlating user and item representations derived from historical interaction patterns. In pursuit of enhanced performance, many methods focus on learning robust and independent representations by disentangling the intricate factors within interaction data across various modalities in an unsupervised manner. However, such an approach obfuscates the discernment of how specific factors (e.g., category or brand) influence the outcomes, making it challenging to regulate their effects. In response to this challenge, we introduce a novel method called Attribute-Driven Disentangled Representation Learning (short for AD-DRL), which explicitly incorporates attributes from different modalities into the disentangled representation learning process. By assigning a specific attribute to each factor in multimodal features, AD-DRL can disentangle the factors at both attribute and attribute-value levels. To obtain robust and independent representations for each factor associated with a specific attribute, we first disentangle the representations of features both within and across different modalities. Moreover, we further enhance the robustness of the representations by fusing the multimodal features of the same factor. Empirical evaluations conducted on three public real-world datasets substantiate the effectiveness of AD-DRL, as well as its interpretability and controllability.
Fan Liu 0008, Yinwei Wei, Zhiyong Cheng 0001, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia6
2024 Perplexity-aware Correction for Robust Alignment with Noisy Preferences
abstract
Alignment techniques are critical in ensuring that large language models (LLMs) output helpful and harmless content by enforcing the LLM-generated content to align with human preferences. However, the existence of noisy preferences (NPs), where the responses are mistakenly labelled as chosen or rejected, could spoil the alignment, thus making the LLMs generate useless and even malicious content. Existing methods mitigate the issue of NPs from the loss perspective by adjusting the alignment loss based on a clean validation dataset. Orthogonal to these loss-oriented methods, we propose perplexity-aware correction (PerpCorrect) from the data perspective for robust alignment which detects and corrects NPs based on the differences between the perplexity of the chosen and rejected responses (dubbed as PPLDiff). Intuitively, a higher PPLDiff indicates a higher probability of the NP because a rejected/chosen response which is mistakenly labelled as chosen/rejected is less preferable to be generated by an aligned LLM, thus having a higher/lower perplexity. PerpCorrect works in three steps: (1) PerpCorrect aligns a surrogate LLM using the clean validation data to make the PPLDiff able to distinguish clean preferences (CPs) and NPs. (2) PerpCorrect further aligns the surrogate LLM by incorporating the reliable clean training data whose PPLDiff is extremely small and reliable noisy training data whose PPLDiff is extremely large after correction to boost the discriminatory power. (3) Detecting and correcting NPs according to the PPLDiff obtained by the aligned surrogate LLM to obtain a denoised training dataset for robust alignment. Comprehensive experiments validate that our proposed PerpCorrect can achieve state-of-the-art alignment performance under NPs. Notably, PerpCorrect demonstrates practical utility by requiring only a modest amount of validation data and being compatible with various alignment techniques. Our code is available at [PerpCorrect](https://github.com/luxinyayaya/PerpCorrect).
Keyi Kong, Xilie Xu, Di Wang 0015, Jingfeng Zhang, Mohan Kankanhalli
NeurIPS5
2024 TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
abstract
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inefficient language supervision in recent web-collected video-text datasets. In this paper, we introduce Text-Only Pre-Alignment (TOPA), a novel approach to extend large language models (LLMs) for video understanding, without the need for pre-training on real video data. Specifically, we first employ an advanced LLM to automatically generate Textual Videos comprising continuous textual frames, along with corresponding annotations to simulate real video-text data. Then, these annotated textual videos are used to pre-align a language-only LLM with the video modality. To bridge the gap between textual and real videos, we employ the CLIP model as the feature extractor to align image and text modalities. During text-only pre-alignment, the continuous textual frames, encoded as a sequence of CLIP text features, are analogous to continuous CLIP image features, thus aligning the LLM with real video representation. Extensive experiments, including zero-shot evaluation and finetuning on various video understanding tasks, demonstrate that TOPA is an effective and efficient framework for aligning video content with LLMs. In particular, without training on any video data, the TOPA-Llama2-13B model achieves a Top-1 accuracy of 51.0% on the challenging long-form video understanding benchmark, Egoschema. This performance surpasses previous video-text pre-training approaches and proves competitive with recent GPT-3.5 based video agents.
Hehe Fan, Yongkang Wong, Mohan Kankanhalli, Yi Yang 0001
NeurIPS4
2024 Privacy-Enhancing Person Re-identification Framework - A Dual-Stage Approach
abstract
In this work, we show that deep learning-based re-identification (Re-ID) models, albeit trained only with a Re-ID objective (i.e. if two samples belong to the same identity), encode personally identifiable information (PII) in the learned features that may lead to serious privacy concerns. In cognizance of the modern privacy regulations on protecting PII, we propose a novel dual-stage person Re-ID framework that (1) suppresses the PII from the discriminative features, and (2) introduces a controllable privacy mechanism through differential privacy. The former is achieved with a self-supervised de-identification (De-ID) decoder and an adversarial-identity (Adv-ID) module, whereas the latter mechanism leverages a controllable privacy budget to generate a privacy-protected gallery with a Gaussian noise generator. Furthermore, we introduce the notion of a privacy metric to quantify the privacy leakage in Re-ID features which is not explicitly examined in prior work. We demonstrate the feasibility of our approach in achieving a better trade-off between utility and privacy through rigorous experiments on person Re-ID benchmarks.
Kajal Kansal, Yongkang Wong, Mohan Kankanhalli
WACV3
2024 A Comprehensive Picture of Factors Affecting User Willingness to Use Mobile Health Applications
abstract
Mobile health (mHealth) applications have become increasingly valuable in preventive healthcare and in reducing the burden on healthcare organizations. The aim of this article is to investigate the factors that influence user acceptance of mHealth apps and identify the underlying structure that shapes users’ behavioral intention. An online study that employed factorial survey design with vignettes was conducted, and a total of 1,669 participants from eight countries across four continents were included in the study. Structural equation modeling was employed to quantitatively assess how various factors collectively contribute to users’ willingness to use mHealth apps. The results indicate that users’ digital literacy has the strongest impact on their willingness to use them, followed by their online habit of sharing personal information. Users’ concerns about personal privacy only had a weak impact. Furthermore, users’ demographic background, such as their country of residence, age, ethnicity, and education, has a significant moderating effect. Our findings have implications for app designers, healthcare practitioners, and policymakers. Efforts are needed to regulate data collection and sharing and promote digital literacy among the general population to facilitate the widespread adoption of mHealth apps.
Shaojing Fan, Ramesh Jain 0001, Mohan Kankanhalli
ACM Trans. Comput. Heal.3
2024 Multi-Modal Meta-Transfer Fusion Network for Few-Shot 3D Model Classification
Heyu Zhou, Anan Liu, Chenyu Zhang 0003, Qianyi Zhang, Mohan Kankanhalli
Int. J. Comput. Vis.6
2024 Multi2Human: Controllable human image generation with multimodal controls
Xiaoling Gu, Shengwenzhuo Xu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli
Neurocomputing7
2024 UNK-VQA: A Dataset and a Probe Into the Abstention Ability of Multi-Modal Large Models
abstract
Teaching Visual Question Answering (VQA) models to refrain from answering unanswerable questions is necessary for building a trustworthy AI system. Existing studies, though have explored various aspects of VQA but somewhat ignored this particular attribute. This paper aims to bridge the research gap by contributing a comprehensive dataset, called UNK-VQA. The dataset is specifically designed to address the challenge of questions that models do not know. To this end, we first augment the existing data via deliberate perturbations on either the image or question. In specific, we carefully ensure that the question-image semantics remain close to the original unperturbed distribution. By this means, the identification of unanswerable questions becomes challenging, setting our dataset apart from others that involve mere image replacement. We then extensively evaluate the zero- and few-shot performance of several emerging multi-modal large models and discover their significant limitations when applied to our dataset. Additionally, we also propose a straightforward method to tackle these unanswerable questions. This dataset, we believe, will serve as a valuable benchmark for enhancing the abstention capability of VQA models, thereby leading to increased trustworthiness of AI systems. We have made the dataset available to facilitate further exploration in this area.
Fangkai Jiao, Zhiqi Shen 0002, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Gaussian Distribution-Aware Commonsense Knowledge Learning for Scene Graph Generation
abstract
Knowledge-based Scene Graph Generation (SGG) requires external commonsense knowledge beyond the visual scene to infer the relation between objects. Such knowledge can be obtained in a variety of forms, such as vision, text, and graph. However, there are two drawbacks as follows: 1) commonsense knowledge essentially has uncertainty, but current works usually represent knowledge in a deterministic manner, which is not well matched to its nature, 2) using commonsense knowledge without denoising will introduce irrelevant information. This can increase the burden on the relation classifier and only obtain marginal gains over a large amount of data. In this paper, we propose a novel Gaussian distribution-aware commonsense knowledge learning method for SGG. First, we associate each object pair with a Gaussian distribution, which parametrizes visual context and commonsense as mean and variance, respectively. We prove that Gaussian modeling can provide a probabilistic soft space to measure the uncertainty of external knowledge, which allows diverse predictions. Second, to reduce semantic noise in commonsense, we sample multiple variables from the Gaussian distribution and train multi-expert classifiers, which can be dynamically examined for the ensemble softmax classification. Extensive comparative experiments on two benchmarks confirm that our method can achieve competitive performance against the state-of-the-art. Ablation studies verify the essential roles of individual components. Moreover, the visualization of multi-expert classifiers confirms our ability to integrate commonsense for relation inference.
Hongshuo Tian, Ning Xu 0003, Mohan Kankanhalli, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Balanced Class-Incremental 3D Object Classification and Retrieval
abstract
Most existing 3D object classification and retrieval algorithms rely on one-off supervised learning on closed 3D object sets and tend to provide rigid convolutional neural networks with little scalability. Such limitations substantially restrict their potential to learn newly emerged 3D object classes continually in the real world. Aiming to go beyond these limitations, we innovatively propose two new and challenging tasks: class-incremental 3D object classification (CI-3DOC) and class-incremental 3D object retrieval (CI-3DOR), the key to which is class-incremental 3D representation learning. It expects the network to update continually to learn new 3D class representations without forgetting the previously learned ones. To this end, we design a novel balanced distillation network(BDNet)that uses a dual supervision mechanism to balance between consolidating old knowledge (stability) and adapting to new 3D object classes (plasticity) carefully. On the one hand, we employ stability-based supervision to retain the stable and discriminative information of old classes that greatly benefit both classification and retrieval tasks. On the other hand, we use plasticity-based supervision to improve the network's generalization for learning new class 3D representations by transferring knowledge from a temporary teacher network to the current model. By properly handling the relationship between the two modules, we achieve a surprising performance improvement. Furthermore, considering there is no available dataset for evaluation, we build two 3D datasets, INOR-1 and INOR-2, to evaluate these two new tasks. Extensive experimental results demonstrate that our method can significantly outperform other state-of-the-art class-incremental learning methods. Even if we store 500-1000 fewer 3D objects than SOTA methods,BDNetstill achieves comparable performance.
Anan Liu, Haochun Lu, Heyu Zhou, Tianbao Li 0001, Mohan Kankanhalli
IEEE Trans. Knowl. Data Eng.5
2024 Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering
abstract
The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in questions and employ a simple graph to aggregate features without considering relative relations between objects, which may lead to inferior performance. In this paper, we propose a Keyword-aware Relative Spatio-Temporal (KRST) graph network for VideoQA. First, to make question features aware of keywords, we employ an attention mechanism to assign high weights to keywords during question encoding. The keyword-aware question features are then used to guide video graph construction. Second, because relations are relative, we integrate the relative relation modeling to better capture the spatio-temporal dynamics among object nodes. Moreover, we disentangle the spatio-temporal reasoning into an object-level spatial graph and a frame-level temporal graph, which reduces the impact of spatial and temporal relation reasoning on each other. Extensive experiments on the TGIF-QA, MSVD-QA and MSRVTT-QA datasets demonstrate the superiority of our KRST over multiple state-of-the-art methods.
Hehe Fan, Dongyun Lin, Ying Sun 0001, Mohan Kankanhalli, Joo-Hwee Lim
IEEE Trans. Multim.5
2024 Learning to Agree on Vision Attention for Visual Commonsense Reasoning
abstract
Visual Commonsense Reasoning (VCR) remains a significant yet challenging research problem in the realm of visual reasoning. A VCR model generally aims at answering a textual question regarding an image, followed by the rationale prediction for the preceding answering process. Though these two processes are sequential and intertwined, existing methods always consider them as two independent matching-based instances. They, therefore, ignore the pivotal relationship between the two processes, leading to sub-optimal model performance. This paper presents a novel visual attention alignment method to efficaciously handle these two processes in a unified framework. To achieve this, we first design a re-attention module for aggregating the vision attention map produced in each process. Thereafter, the resultant two sets of attention maps are carefully aligned to guide the two processes to make decisions based on the same image regions. We apply this method to both conventional attention and the recent Transformer models and carry out extensive experiments on the VCR benchmark dataset. The results demonstrate that with the attention alignment module, our method achieves a considerable improvement over the baseline methods, evidently revealing the feasibility of the coupling of the two processes as well as the effectiveness of the proposed method.
Kejie Wang, Fan Liu 0008, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Multim.6
2024 Fast Yet Effective Machine Unlearning
abstract
Unlearning the data observed during the training of a machine learning (ML) model is an important task that can play a pivotal role in fortifying the privacy and security of ML-based applications. This article raises the following questions: 1) can we unlearn a single or multiple class(es) of data from an ML model without looking at the full training data even once? and 2) can we make the process of unlearning fast and scalable to large datasets, and generalize it to different deep networks? We introduce a novel machine unlearning framework with error-maximizing noise generation and impair-repair based weight manipulation that offers an efficient solution to the above questions. An error-maximizing noise matrix is learned for the class to be unlearned using the original model. The noise matrix is used to manipulate the model weights to unlearn the targeted class of data. We introduce impair and repair steps for a controlled manipulation of the network weights. In the impair step, the noise matrix along with a very high learning rate is used to induce sharp unlearning in the model. Thereafter, the repair step is used to regain the overall performance. With very few update steps, we show excellent unlearning while substantially retaining the overall model accuracy. Unlearning multiple classes requires a similar number of update steps as for a single class, making our approach scalable to large problems. Our method is quite efficient in comparison to the existing methods, works for multiclass unlearning, does not put any constraints on the original optimization mechanism or network design, and works well in both small and large-scale vision tasks. This work is an important step toward fast and easy implementation of unlearning in deep networks. Source code: https://github.com/vikram2000b/Fast-Machine-Unlearning.
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, Mohan Kankanhalli
IEEE Trans. Neural Networks Learn. Syst.4
2024 Cluster-Based Graph Collaborative Filtering
abstract
Graph Convolution Networks (GCNs) have significantly succeeded in learning user and item representations for recommendation systems. The core of their efficacy is the ability to explicitly exploit the collaborative signals from both the first- and high-order neighboring nodes. However, most existing GCN-based methods overlook the multiple interests of users while performing high-order graph convolution. Thus, the noisy information from unreliable neighbor nodes (e.g., users with dissimilar interests) negatively impacts the representation learning of the target node. Additionally, conducting graph convolution operations without differentiating high-order neighbors suffers the over-smoothing issue when stacking more layers, resulting in performance degradation. In this article, we aim to capture more valuable information from high-order neighboring nodes while avoiding noise for better representation learning of the target node. To achieve this goal, we propose a novel GCN-based recommendation model, termed Cluster-based Graph Collaborative Filtering (ClusterGCF). This model performs high-order graph convolution on cluster-specific graphs, which are constructed by capturing the multiple interests of users and identifying the common interests among them. Specifically, we design an unsupervised and optimizable soft node clustering approach to classify user and item nodes into multiple clusters. Based on the soft node clustering results and the topology of the user–item interaction graph, we assign the nodes with probabilities for different clusters to construct the cluster-specific graphs. To evaluate the effectiveness of ClusterGCF, we conducted extensive experiments on four publicly available datasets. Experimental results demonstrate that our model can significantly improve recommendation performance.
Fan Liu 0008, Zhiyong Cheng 0001, Liqiang Nie, Mohan Kankanhalli
ACM Trans. Inf. Syst.5
2024 Unsupervised Domain Adaptation by Causal Learning for Biometric Signal-based HCI
abstract
Biometric signal based human-computer interface (HCI) has attracted increasing attention due to its wide application in healthcare, entertainment, neurocomputing, and so on. In recent years, deep learning-based approaches have made great progress on biometric signal processing. However, the state-of-the-art (SOTA) approaches still suffer from model degradation across subjects or sessions. In this work, we propose a novel unsupervised domain adaptation approach for biometric signal-based HCI via causal representation learning. Specifically, three kinds of interventions on biometric signals (i.e., subjects, sessions, and trials) can be selected to generalize deep models across the selected intervention. In the proposed approach, a generative model is trained for producing intervened features that are subsequently used for learning transferable and causal relations with three modes. Experiments on the EEG-based emotion recognition task and sEMG-based gesture recognition task are conducted to confirm the superiority of our approach. An improvement of +0.21% on the task of inter-subject EEG-based emotion recognition is achieved using our approach. Besides, on the task of inter-session sEMG-based gesture recognition, our approach achieves improvements of +1.47%, +3.36%, +1.71%, and +1.01% on sEMG datasets including CSL-HDEMG, CapgMyo DB-b, 3DC, and Ninapro DB6, respectively. The proposed approach also works on the task of inter-trial sEMG-based gesture recognition and an average improvement of +0.66% on Ninapro databases is achieved. These experimental results show the superiority of the proposed approach compared with the SOTA unsupervised domain adaptation methods on HCIs based on biometric signal.
Qingfeng Dai, Yongkang Wong, Guofei Sun, Zhou Zhou 0012, Mohan Kankanhalli, Weidong Geng
ACM Trans. Multim. Comput. Commun. Appl.6
2024 PAINT: Photo-realistic Fashion Design Synthesis
abstract
In this article, we investigate a new problem of generating a variety of multi-view fashion designs conditioned on a human pose and texture examples of arbitrary sizes, which can replace the repetitive and low-level design work for fashion designers. To solve this challenging multi-modal image translation problem, we propose a novel Photo-reAlistic fashIon desigN synThesis (PAINT) framework, which decomposes the framework into three manageable stages. In the first stage, we employ a Layout Generative Network (LGN) to transform an input human pose into a series of person semantic layouts. In the second stage, we propose a Texture Synthesis Network (TSN) to synthesize textures on all transformed semantic layouts. Specifically, we design a novel attentive texture transfer mechanism for precisely expanding texture patches to the irregular clothing regions of the target fashion designs. In the third stage, we leverage an Appearance Flow Network (AFN) to generate the fashion design images of other viewpoints from a single-view observation by learning 2D multi-scale appearance flow fields. Experimental results demonstrate that our method is capable of generating diverse photo-realistic multi-view fashion design images with fine-grained appearance details conditioned on the provided multiple inputs. The source code and trained models are available at https://github.com/gxl-groups/PAINT .
Xiaoling Gu, Jie Huang 0033, Yongkang Wong, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Recurrent Appearance Flow for Occlusion-Free Virtual Try-On
abstract
Image-based virtual try-on aims at transferring a target in-shop garment onto a reference person, and has garnered significant attention from the research communities recently. However, previous methods have faced severe challenges in handling occlusion problems. To address this limitation, we classify occlusion problems into three types based on the reference person’s arm postures: single-arm occlusion , two-arm non-crossed occlusion , and two-arm crossed occlusion . Specifically, we propose a novel Occlusion-Free Virtual Try-On Network (OF-VTON) that effectively overcomes these occlusion challenges. The OF-VTON framework consists of two core components: (i) a new Recurrent Appearance Flow based Deformation (RAFD) model that robustly aligns the in-shop garment to the reference person by adopting a multi-task learning strategy . This model jointly produces the dense appearance flow to warp the garment and predicts a human segmentation map to provide semantic guidance for the subsequent image synthesis model. (ii) a powerful Multi-mask Image SynthesiS (MISS) model that generates photo-realistic try-on results by introducing a new mask generation and selection mechanism . Experimental results demonstrate that our proposed OF-VTON significantly outperforms existing state-of-the-art methods by mitigating the impact of occlusion problems. Our code is available at https://github.com/gxl-groups/OF-VTON .
Xiaoling Gu, Junkai Zhu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.7
2023 Can Bad Teaching Induce Forgetting? Unlearning in Deep Networks Using an Incompetent Teacher
abstract
Machine unlearning has become an important area of research due to an increasing need for machine learning (ML) applications to comply with the emerging data privacy regulations. It facilitates the provision for removal of certain set or class of data from an already trained ML model without requiring retraining from scratch. Recently, several efforts have been put in to make unlearning to be effective and efficient. We propose a novel machine unlearning method by exploring the utility of competent and incompetent teachers in a student-teacher framework to induce forgetfulness. The knowledge from the competent and incompetent teachers is selectively transferred to the student to obtain a model that doesn't contain any information about the forget data. We experimentally show that this method generalizes well, is fast and effective. Furthermore, we introduce the zero retrain forgetting (ZRF) metric to evaluate any unlearning method. Unlike the existing unlearning metrics, the ZRF score does not depend on the availability of the expensive retrained model. This makes it useful for analysis of the unlearned model after deployment as well. We present results of experiments conducted for random subset forgetting and class forgetting on various deep networks and across different application domains. Code is at: https://github.com/vikram2000b/bad-teaching- unlearning
Vikram S. Chundawat, Ayush K. Tarun, Murari Mandal, Mohan Kankanhalli
AAAI4
2023 Text to Point Cloud Localization with Relation-Enhanced Transformer
abstract
Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on a text-to-point-cloud cross-modal localization problem. Given a textual query, it aims to identify the described location from city-scale point clouds. The task involves two challenges. 1) In city-scale point clouds, similar ambient instances may exist in several locations. Searching each location in a huge point cloud with only instances as guidance may lead to less discriminative signals and incorrect results. 2) In textual descriptions, the hints are provided separately. In this case, the relations among those hints are not explicitly described, leaving the difficulties of learning relations to the agent itself. To alleviate the two challenges, we propose a unified Relation-Enhanced Transformer (RET) to improve representation discriminability for both point cloud and nature language queries. The core of the proposed RET is a novel Relation-enhanced Self-Attention (RSA) mechanism, which explicitly encodes instance (hint)-wise relations for the two modalities. Moreover, we propose a fine-grained cross-modal matching method to further refine the location predictions in a subsequent instance-hint matching stage. Experimental results on the KITTI360Pose dataset demonstrate that our approach surpasses the previous state-of-the-art method by large margins.
Guangzhi Wang, Hehe Fan, Mohan Kankanhalli
AAAI3
2023 PointListNet: Deep Learning on 3D Point Lists
abstract
Deep neural networks on regular 1D lists (e.g., natural languages) and irregular 3D sets (e.g., point clouds) have made tremendous achievements. The key to natural language processing is to model words and their regular order dependency in texts. For point cloud understanding, the challenge is to understand the geometry via irregular point coordinates, in which point-feeding orders do not matter. However, there are a few kinds of data that exhibit both regular 1 D list and irregular 3D set structures, such as proteins and non-coding RNAs. In this paper, we refer to them as 3D point lists and propose a Transformer-style PointListNet to model them. First, PointListNet employs non-parametric distance-based attention because we find sometimes it is the distance, instead of the feature or type, that mainly determines how much two points, e.g., amino acids, are correlated in the micro world. Second, different from the vanilla Transformer that directly performs a simple linear transformation on inputs to generate values and does not explicitly model relative relations, our PointListNet integrates the 1D order and 3D Euclidean displacements into values. We conduct experiments on protein fold classification and enzyme reaction classification. Experimental results show the effectiveness of the proposed PointListNet.
Hehe Fan, Linchao Zhu, Yi Yang 0001, Mohan Kankanhalli
CVPR4
2023 DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face Alignment
abstract
Sensitivity to severe occlusion and large view angles limits the usage scenarios of the existing monocular 3D dense face alignment methods. The state-of-the-art 3DMM-based method, directly regresses the model's coefficients, underutilizing the low-level 2D spatial and semantic information, which can actually offer cues for face shape and orientation. In this work, we demonstrate how modeling 3D facial geometry in image and model space jointly can solve the occlusion and view angle problems. Instead of predicting the whole face directly, we regress image space features in the visible facial region by dense prediction first. Subsequently, we predict our model's coefficients based on the regressed feature of the visible regions, leveraging the prior knowledge of whole face geometry from the morphable models to complete the invisible regions. We further propose a fusion network that combines the advantages of both the image and model space predictions to achieve high robustness and accuracy in unconstrained scenarios. Thanks to the proposed fusion module, our method is robust not only to occlusion and large pitch and roll view angles, which is the bene- fit of our image space approach, but also to noise and large yaw angles, which is the benefit of our model space method. Comprehensive evaluations demonstrate the superior performance of our method compared with the state-of-the-art methods. On the 3D dense face alignment task, we achieve 3.80% NME on the AFLW2000-3D dataset, which outperforms the state-of-the-art method by 5.5%. Code is available at https://github.com/1hyfst/DSFNet.
Heyuan Li, Bo Wang 0019, Yu Cheng 0009, Mohan Kankanhalli, Robby T. Tan
CVPR4
2023 Continuous-Discrete Convolution for Geometry-Sequence Modeling in Proteins
Hehe Fan, Zhangyang Wang, Yi Yang 0001, Mohan Kankanhalli
ICLR4
2023 Deep Regression Unlearning
abstract
With the introduction of data protection and privacy regulations, it has become crucial to remove the lineage of data on demand from a machine learning (ML) model. In the last few years, there have been notable developments in machine unlearning to remove the information of certain training data efficiently and effectively from ML models. In this work, we explore unlearning for the regression problem, particularly in deep learning models. Unlearning in classification and simple linear regression has been considerably investigated. However, unlearning in deep regression models largely remains an untouched problem till now. In this work, we introduce deep regression unlearning methods that generalize well and are robust to privacy attacks. We propose the Blindspot unlearning method which uses a novel weight optimization process. A randomly initialized model, partially exposed to the retain samples and a copy of the original model are used together to selectively imprint knowledge about the data that we wish to keep and scrub off the information of the data we wish to forget. We also propose a Gaussian fine tuning method for regression unlearning. The existing unlearning metrics for classification are not directly applicable to regression unlearning. Therefore, we adapt these metrics for the regression setting. We conduct regression unlearning experiments for computer vision, natural language processing and forecasting applications. Our methods show excellent performance for all these datasets across all the metrics. Source code: https://github.com/ayu987/deep-regression-unlearning
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, Mohan Kankanhalli
ICML4
2023 Sample Less, Learn More: Efficient Action Recognition via Frame Feature Restoration
abstract
Training an effective video action recognition model poses significant computational challenges, particularly under limited resource budgets. Current methods primarily aim to either reduce model size or utilize pre-trained models, limiting their adaptability to various backbone architectures. This paper investigates the issue of over-sampled frames, a prevalent problem in many approaches yet it has received relatively little attention. Despite the use of fewer frames being a potential solution, this approach often results in a substantial decline in performance. To address this issue, we propose a novel method to restore the intermediate features for two sparsely sampled and adjacent video frames. This feature restoration technique brings a negligible increase in computational requirements compared to resource-intensive image encoders, such as ViT. To evaluate the effectiveness of our method, we conduct extensive experiments on four public datasets, including Kinetics-400, ActivityNet, UCF-101, and HMDB-51. With the integration of our method, the efficiency of three commonly used baselines has been improved by over 50%, with a mere 0.5% reduction in recognition accuracy. In addition, our method also surprisingly helps improve the generalization ability of the models under zero-shot settings.
Harry Cheng 0002, Liqiang Nie, Zhiyong Cheng 0001, Mohan Kankanhalli
ACM Multimedia5
2023 NarSUM '23: The 2nd Workshop on User-Centric Narrative Summarization of Long Videos
abstract
With video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization, etc.
Mohan Kankanhalli, Ioannis Patras, Jianquan Liu, Yongkang Wong, Takahiro Komamizu, Satoshi Yamazaki, Karen Stephen, Kajal Kansal
ACM Multimedia1
2023 Panel: Multimodal Large Foundation Models
abstract
The surprisingly fluent predictive performance of LLM (Large Language Models) as well as the high-quality photo-realistic rendering of Diffusion Models has heralded a new beginning in the area of Generative AI. Such kinds of deep learning based models with billions of parameters and pre-trained on massive-scale data-sets are also called Large Foundation Models (LFM). These models not only have caught the public imagination but also have led to an unprecedented surge in interest towards the applications of these models. Instead of the previous approach of developing AI models for specific tasks, more and more researchers are developing large task-agnostic models pre-trained on massive data, which can then be adapted to a variety of downstream tasks via fine-tuning, fewshot learning, or zero-shot learning. Some examples are ChatGPT, LLaMA, GPT-4, Flamingo, MidJourney, Stable-Diffusion and DALLE. Some of them can handle text (e.g., ChatGPT, LLaMA) while some others (e.g., GPT-4 and Flamingo) can utilize multimodal data and can hence be considered Multimodal Large Foundation Models (MLFM).
Mohan Kankanhalli, Marcel Worring
ACM Multimedia1
2023 Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
abstract
Visual Commonsense Reasoning (VCR) calls for explanatory reasoning behind question answering over visual scenes. To achieve this goal, a model is required to provide an acceptable rationale as the reason for the predicted answers. Progress on the benchmark dataset stems largely from the recent advancement of Vision-Language Transformers (VL Transformers). These models are first pre-trained on some generic large-scale vision-text datasets, and then the learned representations are transferred to the downstream VCR task. Despite their attractive performance, this paper posits that the VL Transformers do not exhibit visual commonsense, which is the key to VCR. In particular, our empirical results pinpoint several shortcomings of existing VL Transformers: small gains from pre-training, unexpected language bias, limited model architecture for the two inseparable sub-tasks, and neglect of the important object-tag correlation. With these findings, we tentatively suggest some future directions from the aspect of dataset, evaluation metric, and training tricks. We believe this work could make researchers revisit the intuition and goals of VCR, and thus help tackle the remaining challenges in visual reasoning.
Kejie Wang, Xiaolin Chen 0001, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia6
2023 Semantic-Guided Feature Distillation for Multimodal Recommendation
abstract
Multimodal recommendation exploits the rich multimodal information associated with users or items to enhance the representation learning for better performance. In these methods, end-to-end feature extractors (e.g., shallow/deep neural networks) are often adopted to tailor the generic multimodal features that are extracted from raw data by pre-trained models for recommendation. However, compact extractors, such as shallow neural networks, may find it challenging to extract effective information from complex and high-dimensional generic modality features. Conversely, DNN-based extractors may encounter the data sparsity problem in recommendation. To address this problem, we propose a novel model-agnostic approach called Semantic-guided Feature Distillation (SGFD), which employs a teacher-student framework to extract feature for multimodal recommendation. The teacher model first extracts rich modality features from the generic modality feature by considering both the semantic information of items and the complementary information of multiple modalities. SGFD then utilizes response-based and feature-based distillation loss to effectively transfer the knowledge encoded in the teacher model to the student model. To evaluate the effectiveness of our SGFD, we integrate SGFD into three backbone multimodal recommendation models. Extensive experiments on three public real-world datasets demonstrate that SGFD-enhanced models can achieve substantial improvement over their counterparts.
Fan Liu 0008, Huilin Chen 0002, Zhiyong Cheng 0001, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia5
2023 Combating Misinformation in the Era of Generative AI Models
abstract
Misinformation has been a persistent and harmful phenomenon affecting our society in various ways, including individuals' physical health and economic stability. With the rise of short video platforms and related applications, the spread of multi-modal misinformation, encompassing images, texts, audios, and videos have exacerbated these concerns. The introduction of generative AI models like ChatGPT and Stable Diffusion has further complicated matters, giving rise to Artificial Intelligence Generated Content (AIGC) and presenting new challenges in detecting and mitigating misinformation. Consequently, traditional approaches to misinformation detection and intervention have become inadequate in this evolving landscape. This paper explores the challenges posed by AIGC in the context of misinformation. It examines the issue from psychological and societal perspectives, and explores the subtle manipulation traces found in AIGC at signal, perceptual, semantic, and human levels. By scrutinizing manipulation traces such as signal manipulation, semantic inconsistencies, logical incoherence, and psychological strategies, our objective is to tackle AI-generated misinformation and provide a conceptual design of systematic explainable solution. Ultimately, we aim for this paper to contribute valuable insights into combating misinformation, particularly in the era of AIGC.
Danni Xu, Shaojing Fan, Mohan Kankanhalli
ACM Multimedia3
2023 Enhancing Adversarial Contrastive Learning via Adversarial Invariant Regularization
abstract
Adversarial contrastive learning (ACL) is a technique that enhances standard contrastive learning (SCL) by incorporating adversarial data to learn a robust representation that can withstand adversarial attacks and common corruptions without requiring costly annotations. To improve transferability, the existing work introduced the standard invariant regularization (SIR) to impose style-independence property to SCL, which can exempt the impact of nuisance style factors in the standard representation. However, it is unclear how the style-independence property benefits ACL-learned robust representations. In this paper, we leverage the technique of causal reasoning to interpret the ACL and propose adversarial invariant regularization (AIR) to enforce independence from style factors. We regulate the ACL using both SIR and AIR to output the robust representation. Theoretically, we show that AIR implicitly encourages the representational distance between different views of natural data and their adversarial variants to be independent of style factors. Empirically, our experimental results show that invariant regularization significantly improves the performance of state-of-the-art ACL methods in terms of both standard generalization and robustness on downstream tasks. To the best of our knowledge, we are the first to apply causal reasoning to interpret ACL and develop AIR for enhancing ACL-learned robust representations. Our source code is at https://github.com/GodXuxilie/Enhancing_ACL_via_AIR.
Xilie Xu, Jingfeng Zhang, Feng Liu 0003, Masashi Sugiyama, Mohan Kankanhalli
NeurIPS5
2023 Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset Selection
abstract
Adversarial contrastive learning (ACL) does not require expensive data annotations but outputs a robust representation that withstands adversarial attacks and also generalizes to a wide range of downstream tasks. However, ACL needs tremendous running time to generate the adversarial variants of all training data, which limits its scalability to large datasets. To speed up ACL, this paper proposes a robustness-aware coreset selection (RCS) method. RCS does not require label information and searches for an informative subset that minimizes a representational divergence, which is the distance of the representation between natural data and their virtual adversarial variants. The vanilla solution of RCS via traversing all possible subsets is computationally prohibitive. Therefore, we theoretically transform RCS into a surrogate problem of submodular maximization, of which the greedy search is an efficient solution with an optimality guarantee for the original problem. Empirically, our comprehensive results corroborate that RCS can speed up ACL by a large margin without significantly hurting the robustness transferability. Notably, to the best of our knowledge, we are the first to conduct ACL efficiently on the large-scale ImageNet-1K dataset to obtain an effective robust representation via RCS. Our source code is at https://github.com/GodXuxilie/Efficient_ACL_via_RCS.
Xilie Xu, Jingfeng Zhang, Feng Liu 0003, Masashi Sugiyama, Mohan Kankanhalli
NeurIPS5
2023 Emotional Attention: From Eye Tracking to Computational Modeling
abstract
Attending selectively to emotion-eliciting stimuli is intrinsic to human vision. In this research, we investigate how emotion-elicitation features of images relate to human selective attention. We create the EMOtional attention dataset (EMOd). It is a set of diverse emotion-eliciting images, each with (1) eye-tracking data from 16 subjects, (2) image context labels at both object- and scene-level. Based on analyses of human perceptions of EMOd, we report an emotion prioritization effect: emotion-eliciting content draws stronger and earlier human attention than neutral content, but this advantage diminishes dramatically after initial fixation. We find that human attention is more focused on awe eliciting and aesthetic vehicle and animal scenes in EMOd. Aiming to model the above human attention behavior computationally, we design a deep neural network (CASNet II), which includes a channel weighting subnetwork that prioritizes emotion-eliciting objects, and an Atrous Spatial Pyramid Pooling (ASPP) structure that learns the relative importance of image regions at multiple scales. Visualizations and quantitative analyses demonstrate the model's ability to simulate human attention behavior, especially on emotion-eliciting content.
Shaojing Fan, Zhiqi Shen 0002, Ming Jiang 0019, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Point Spatio-Temporal Transformer Networks for Point Cloud Video Modeling
abstract
Due to the inherent unorderliness and irregularity of point cloud, points emerge inconsistently across different frames in a point cloud video. To capture the dynamics in point cloud videos, tracking points and limiting temporal modeling range are usually employed to preserve spatio-temporal structure. However, as points may flow in and out across frames, computing accurate point trajectories is extremely difficult, especially for long videos. Moreover, when points move fast, even in a small temporal window, points may still escape from a region. Besides, using the same temporal range for different motions may not accurately capture the temporal structure. In this paper, we propose a Point Spatio-Temporal Transformer (PST-Transformer). To preserve the spatio-temporal structure, PST-Transformer adaptively searches related or similar points across the entire video by performing self-attention on point features. Moreover, our PST-Transformer is equipped with an ability to encode spatio-temporal structure. Because point coordinates are irregular and unordered but point timestamps exhibit regularities and order, the spatio-temporal encoding is decoupled to reduce the impact of the spatial irregularity on the temporal modeling. By properly preserving and encoding spatio-temporal structure, our PST-Transformer effectively models point cloud videos and shows superior performance on 3D action recognition and 4D semantic segmentation.
Hehe Fan, Yi Yang 0001, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Fair Representation: Guaranteeing Approximate Multiple Group Fairness for Unknown Tasks
abstract
Motivated by scenarios where data is used for diverse prediction tasks, we study whether fair representation can be used to guarantee fairness for unknown tasks and for multiple fairness notions. We consider seven group fairness notions that cover the concepts of independence, separation, and calibration. Against the backdrop of the fairness impossibility results, we explore approximate fairness. We prove that, although fair representation might not guarantee fairness for all prediction tasks, it does guarantee fairness for an important subset of tasks-the tasks for which the representation is discriminative. Specifically, all seven group fairness notions are linearly controlled by fairness and discriminativeness of the representation. When an incompatibility exists between different fairness notions, fair and discriminative representation hits the sweet spot that approximately satisfies all notions. Motivated by our theoretical findings, we propose to learn both fair and discriminative representations using pretext loss which self-supervises learning, and Maximum Mean Discrepancy as a fair regularizer. Experiments on tabular, image, and face datasets show that using the learned representation, downstream predictions that we are unaware of when learning the representation indeed become fairer. The fairness guarantees computed from our theoretical results are all valid.
Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 When and Why Static Images Are More Effective Than Videos
abstract
People often prefer videos over images in research and applications, believing that videos are more effective for eliciting human emotions and building machine intelligence. However, our research shows that this assumption is not always correct when it comes to evoking emotions in human observers. In this article, we compare thirteen emotions and two perceptions elicited by short videos (2-6 second, silent video clips) versus static frames extracted from the videos. We show that static frames and videos elicit most emotions similarly, but static frames elicit negative emotions more strongly than videos. We test two complementary explanations: differential activation of suspense and the peak-end rule. These findings help us to computationally model human reactions more faithfully with fewer video frames. Our interdisciplinary results have important implications for methods, theory, and applications in diverse fields, including social psychology, computer vision, mass media, and marketing.
Shaojing Fan, Zhiqi Shen 0002, Bryan L. Koenig, Tian-Tsong Ng, Mohan Kankanhalli
IEEE Trans. Affect. Comput.5
2023 Zero-Shot Machine Unlearning
abstract
Modern privacy regulations grant citizens the right to be forgotten by products, services and companies. In case of machine learning (ML) applications, this necessitates deletion of data not only from storage archives but also from ML models. Due to an increasing need for regulatory compliance required for ML applications,machine unlearningis becoming an emerging research problem. The right to be forgotten requests come in the form of removal of a certain set or class of data from the already trained ML model. Practical considerations preclude retraining of the model from scratch after discarding the deleted data. The few existing studies use either the whole training data, or a subset of training data, or some metadata stored during training to update the model weights for unlearning. However, strict regulatory compliance requires time-bound deletion of data. Thus, in many cases, no data related to the training process or training samples may be accessible even for the unlearning purpose. We therefore ask the question:is it possible to achieve unlearning with zero training samples?In this paper, we introduce the novel problem ofzero-shot machine unlearningthat caters for the extreme but practical scenario where zero original data samples are available for use. We then propose two novel solutions forzero-shot machine unlearningbased on (a) error minimizing-maximizing noise and (b) gated knowledge transfer. These methods remove the information of the forget data from the model while maintaining the model efficacy on the retain data. The zero-shot approach offers good protection against the model inversion attacks and membership inference attacks. We introduce a new evaluation metric,Anamnesis Index(AIN) to effectively measure the quality of the unlearning method. The experiments show promising results for unlearning in deep learning models on benchmark vision data-sets. The source code is available here: https://github.com/ayu987/zero-shot-unlearning.
Vikram S. Chundawat, Ayush K. Tarun, Murari Mandal, Mohan Kankanhalli
IEEE Trans. Inf. Forensics Secur.4
2023 Joint Answering and Explanation for Visual Commonsense Reasoning
abstract
Visual Commonsense Reasoning (VCR), deemed as one challenging extension of Visual Question Answering (VQA), endeavors to pursue a higher-level visual comprehension. VCR includes two complementary processes: question answering over a given image and rationale inference for answering explanation. Over the years, a variety of VCR methods have pushed more advancements on the benchmark dataset. Despite significance of these methods, they often treat the two processes in a separate manner and hence decompose VCR into two irrelevant VQA instances. As a result, the pivotal connection between question answering and rationale inference is broken, rendering existing efforts less faithful to visual reasoning. To empirically study this issue, we perform some in-depth empirical explorations in terms of both language shortcuts and generalization capability. Based on our findings, we then propose a plug-and-play knowledge distillation enhanced framework to couple the question answering and rationale inference processes. The key contribution lies in the introduction of a new branch, which serves as a relay to bridge the two processes. Given that our framework is model-agnostic, we apply it to the existing popular baselines and validate its effectiveness on the benchmark dataset. As demonstrated in the experimental results, when equipped with our method, these baselines all achieve consistent and significant performance improvements, evidently verifying the viability of processes coupling.
Kejie Wang, Yinwei Wei, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Image Process.6
2023 Disentangled Multimodal Representation Learning for Recommendation
abstract
Many multimodal recommender systems have been proposed to exploit the rich side information associated with users or items (e.g., user reviews and item images) for learning better user and item representations to improve the recommendation performance. Studies from psychology show that users have individual differences in the utilization of various modalities for organizing information. Therefore, for a certain factor of an item (such asappearanceorquality), the features of different modalities are of varying importance to a user. However, existing methods ignore the fact that different modalities contribute differently towards a user's preference on various factors of an item. In light of this, in this paper, we propose a novelDisentangled Multimodal Representation Learning(DMRL) recommendation model, which can capture users' attention to different modalities on each factor in user preference modeling. In particular, we employ a disentangled representation technique to ensure the features of different factors in each modality are independent of each other. A multimodal attention mechanism is then designed to capture users' modality preference for each factor. Based on the estimated weights obtained by the attention mechanism, we make recommendations by combining the preference scores of a user's preferences to each factor of the target item over different modalities. Extensive evaluation on five real-world datasets demonstrate the superiority of our method compared with existing methods.
Fan Liu 0008, Huilin Chen 0002, Zhiyong Cheng 0001, Anan Liu, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Multim.6
2023 Learning to Minimize the Remainder in Supervised Learning
abstract
The learning process of deep learning methods usually updates the model’s parameters in multiple iterations. Each iteration can be viewed as the first-order approximation of Taylor’s series expansion. The remainder, which consists of higher-order terms, is usually ignored in the learning process for simplicity. This learning scheme empowers various multimedia-based applications, such as image retrieval, recommendation system, and video search. Generally, multimedia data (e.g.images) are semantics-rich and high-dimensional, hence the remainders of approximations are possibly non-zero. In this work, we consider that the remainder is informative and study how it affects the learning process. To this end, we propose a new learning approach, namely gradient adjustment learning (GAL), to leverage the knowledge learned from the past training iterations to adjust vanilla gradients, such that the remainders are minimized and the approximations are improved. The proposed GAL is model- and optimizer-agnostic, and is easy to adapt to the standard learning framework. It is evaluated on three tasks,i.e.image classification, object detection, and regression, with state-of-the-art models and optimizers. The experiments show that the proposed GAL consistently enhances the evaluated models, whereas the ablation studies validate various aspects of the proposed GAL. The code is available athttps://github.com/luoyan407/gradient_adjustment.git.
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
IEEE Trans. Multim.3
2023 Semantic-Aware Triplet Loss for Image Classification
abstract
Successful image classification requires a discriminative representation learning model for images. To approach this idea, deep metric learning (DML), serving as building a basic feature space with a pre-defined metric, has demonstrated compelling performance over the years. DML is often implemented with a carefully crafted loss function, such as the representative triplet loss, which encourages a positive sample to be by a fixed margin closer to the anchor than the negative. Despite its efficacy, the negative samples are treated uniformly, rendering the feature space less informative since different negative samples can be largely different from the anchor. In this work, we, for the first time, propose to exploit the semantic information inherent in discrete class labels as an aid for the triplet loss. Specifically, we build a bi-level negative sampling strategy,i.e., strong negative and weak negative sampling, with the guidance of an external knowledge source, from which rich class semantics can be extracted. With several fine-grained and complementary triplet losses based on this strategy, our method is enhanced with semantic awareness for image classification. In addition, to coordinate with the complicated training dynamics, we devise an ad-hoc Semantic Relation Weighting module, which consistently inspects model states and dynamically adjusts the importance of each triplet loss. It is worth noting that our method is plug-and-play, and we thus test its validity over various backbones and knowledge sources. Both qualitative and quantitative experimental results on benchmark datasets demonstrate the effectiveness of employing semantics for image classification.
Guangzhi Wang, Ziwei Xu 0001, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Multim.5
2023 On Modality Bias Recognition and Reduction
abstract
Making each modality in multi-modal data contribute is of vital importance to learning a versatile multi-modal model. Existing methods, however, are often dominated by one or few of modalities during model training, resulting in sub-optimal performance. In this article, we refer to this problem as modality bias and attempt to study it in the context of multi-modal classification systematically and comprehensively. After stepping into several empirical analyses, we recognize that one modality affects the model prediction more just because this modality has a spurious correlation with instance labels. To primarily facilitate the evaluation on the modality bias problem, we construct two datasets, respectively, for the colored digit recognition and video action recognition tasks in line with the Out-of-Distribution (OoD) protocol. Collaborating with the benchmarks in the visual question answering task, we empirically justify the performance degradation of the existing methods on these OoD datasets, which serves as evidence to justify the modality bias learning. In addition, to overcome this problem, we propose a plug-and-play loss function method, whereby the feature space for each label is adaptively learned according to the training set statistics. Thereafter, we apply this method on 10 baselines in total to test its effectiveness. From the results on four datasets regarding the above three tasks, our method yields remarkable performance improvements compared with the baselines, demonstrating its superiority on reducing the modality bias problem.
Liqiang Nie, Harry Cheng 0002, Zhiyong Cheng 0001, Mohan Kankanhalli, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation with Reliable Voted Pseudo Labels
abstract
In this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unla-beled data by scaling up or down point clouds and then predicting the scales. Second, to capture the local structure in a self-supervised manner, we propose to project a 3D local area onto a 2D plane and then learn to reconstruct the squeezed region. Moreover, to effectively transfer the knowledge from source domain, we propose to vote pseudo labels for target samples based on the labels of their nearest source neighbors in the shared feature space. To avoid the noise caused by incorrect pseudo labels, we only select re-liable target samples, whose voting consistencies are high enough, for enhancing adaptation. The voting method is able to adaptively select more and more target samples during training, which in return facilitates adaptation because the amount of labeled target data increases. Experiments on PointDA (ModelNet-10, ShapeNet-10 and ScanNet-10) and Sim-to-Real (ModelNet-11, ScanObjectNN-11, ShapeNet-9 and ScanObjectNN-9) demonstrate the effectiveness of our method.
Hehe Fan, Xiaojun Chang, Wanyue Zhang, Ying Sun 0001, Mohan Kankanhalli
CVPR6
2022 DIOT: Detecting Implicit Obstacles from Trajectories
Yifan Lei, Mohan Kankanhalli, Anthony K. H. Tung
DASFAA (1)3
2022 Chairs Can Be Stood On: Overcoming Object Bias in Human-Object Interaction Detection
Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
ECCV (24)4
2022 Adversarial Attack and Defense for Non-Parametric Two-Sample Tests
abstract
Non-parametric two-sample tests (TSTs) that judge whether two sets of samples are drawn from the same distribution, have been widely used in the analysis of critical data. People tend to employ TSTs as trusted basic tools and rarely have any doubt about their reliability. This paper systematically uncovers the failure mode of non-parametric TSTs through adversarial attacks and then proposes corresponding defense strategies. First, we theoretically show that an adversary can upper-bound the distributional shift which guarantees the attack’s invisibility. Furthermore, we theoretically find that the adversary can also degrade the lower bound of a TST’s test power, which enables us to iteratively minimize the test criterion in order to search for adversarial pairs. To enable TST-agnostic attacks, we propose an ensemble attack (EA) framework that jointly minimizes the different types of test criteria. Second, to robustify TSTs, we propose a max-min optimization that iteratively generates adversarial pairs to train the deep kernels. Extensive experiments on both simulated and real-world datasets validate the adversarial vulnerabilities of non-parametric TSTs and the effectiveness of our proposed defense. Source code is available at https://github.com/GodXuxilie/Robust-TST.git.
Xilie Xu, Jingfeng Zhang, Feng Liu 0003, Masashi Sugiyama, Mohan Kankanhalli
ICML5
2022 Learning Realistic Patterns from Visually Unrealistic Stimuli: Generalization and Data Anonymization (Extended Abstract)
abstract
Good training data is a prerequisite to develop useful Machine Learning applications. However, in many domains existing data sets cannot be shared due to privacy regulations (e.g., from medical studies). This work investigates a simple yet unconventional approach for anonymized data synthesis to enable third parties to benefit from such anonymized data. We explore the feasibility of learning implicitly from visually unrealistic, task-relevant stimuli, which are synthesized by exciting the neurons of a trained deep neural network. As such, neuronal excitation can be used to generate synthetic stimuli. The stimuli data is used to train new classification models. Furthermore, we extend this framework to inhibit representations that are associated with specific individuals. Extensive comparative empirical investigation shows that different algorithms trained on the stimuli are able to generalize successfully on the same task as the original model.
Konstantinos Nikolaidis, Stein Kristiansen, Thomas Plagemann, Vera Goebel, Knut Liestøl, Mohan Kankanhalli, Gunn Marit Traaen, Britt Øverland, Harriet Akre, Lars Aakerøy, Sigurd Steinshamn
IJCAI6
2022 A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQA
abstract
Knowledge-based Visual Question Answering (VQA) expects models to rely on external knowledge for robust answer prediction. Though significant it is, this paper discovers several leading factors impeding the advancement of current state-of-the-art methods. On the one hand, methods which exploit the explicit knowledge take the knowledge as a complement for the coarsely trained VQA model. Despite their effectiveness, these approaches often suffer from noise incorporation and error propagation. On the other hand, pertaining to the implicit knowledge, the multi-modal implicit knowledge for knowledge-based VQA still remains largely unexplored. This work presents a unified end-to-end retriever-reader framework towards knowledge-based VQA. In particular, we shed light on the multi-modal implicit knowledge from vision-language pre-training models to mine its potential in knowledge reasoning. As for the noise problem encountered by the retrieval operation on explicit knowledge, we design a novel scheme to create pseudo labels for effective knowledge supervision. This scheme is able to not only provide guidance for knowledge retrieval, but also drop these instances potentially error-prone towards question answering. To validate the effectiveness of the proposed method, we conduct extensive experiments on the benchmark dataset. The experimental results reveal that our method outperforms existing baselines by a noticeable margin. Beyond the reported numbers, this paper further spawns several insights on knowledge utilization for future research with some empirical findings.
Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng 0001, Mohan Kankanhalli
ACM Multimedia6
2022 NarSUM '22: 1st Workshop on User-centric Narrative Summarization of Long Videos
abstract
With video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization etc.
Mohan Kankanhalli, Jianquan Liu, Yongkang Wong, Karen Stephen, Rishabh Sheoran, Anusha Bhamidipati
ACM Multimedia1
2022 Distance Matters in Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection has received considerable attention in the context of scene understanding. Despite the growing progress, we realize existing methods often perform unsatisfactorily on distant interactions, where the leading causes are two-fold: 1) Distant interactions are by nature more difficult to recognize than close ones. A natural scene often involves multiple humans and objects with intricate spatial relations, making the interaction recognition for distant human-object largely affected by complex visual context. 2) Insufficient number of distant interactions in datasets results in under-fitting on these instances. To address these problems, we propose a novel two-stage method for better handling distant interactions in HOI detection. One essential component in our method is a novel Far Near Distance Attention module. It enables information propagation between humans and objects, whereby the spatial distance is skillfully taken into consideration. Besides, we devise a novel Distance-Aware loss function which leads the model to focus more on distant yet rare interactions. We conduct extensive experiments on HICO-DET and V-COCO datasets. The results show that the proposed method surpass existing methods significantly, leading to new state-of-the-art results.
Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
ACM Multimedia4
2022 Compute to Tell the Tale: Goal-Driven Narrative Generation
abstract
Man is by nature a social animal. One important facet of human evolution is through narrative imagination, be it fictional or factual, and to tell the tale to other individuals. The factual narrative, such as news, journalism, field report, etc., is based on real-world events and often requires extensive human efforts to create. In the era of big data where video capture devices are commonly available everywhere, a massive amount of raw videos (including life-logging, dashcam or surveillance footage) are generated daily. As a result, it is rather impossible for humans to digest and analyze these video data. This paper reviews the problem of computational narrative generation where a goal-driven narrative (in the form of text with or without video) is generated from a single or multiple long videos. Importantly, the narrative generation problem makes itself distinguished from the existing literature by its focus on a comprehensive understanding of user goal, narrative structure and open-domain input. We tentatively outline a general narrative generation framework and discuss the potential research problems and challenges in this direction. Informed by the real-world impact of narrative generation, we then illustrate several practical use cases in Video Logging as a Service platform which enables users to get more out of the data through a goal-driven intelligent storytelling AI agent.
Yongkang Wong, Shaojing Fan, Ziwei Xu 0001, Karen Stephen, Rishabh Sheoran, Anusha Bhamidipati, Vivek Barsopia, Jianquan Liu, Mohan Kankanhalli
ACM Multimedia10
2022 Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action Segmentation
abstract
We propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output and the constraints. We propose a comprehensive set of constraints, which are implicit in data annotations, and incorporate them with deep networks via DTL. We evaluate the effectiveness of DTL on the temporal action segmentation task and observe improved performance and reduced logical errors in the output of different task models. Furthermore, we provide an extensive analysis to visualize the desirable effects of DTL.
Ziwei Xu 0001, Yogesh S. Rawat, Yongkang Wong, Mohan Kankanhalli, Mubarak Shah
NeurIPS4
2022 Privacy-Preserving Synthetic Data Generation for Recommendation Systems
abstract
Recommendation systems make predictions chiefly based on users' historical interaction data (e.g., items previously clicked or purchased). There is a risk of privacy leakage when collecting the users' behavior data for building the recommendation model. However, existing privacy-preserving solutions are designed for tackling the privacy issue only during the model training [32] and results collection [40] phases. The problem of privacy leakage still exists when directly sharing the private user interaction data with organizations or releasing them to the public. To address this problem, in this paper, we present a User Privacy Controllable Synthetic Data Generation model (short for UPC-SDG), which generates synthetic interaction data for users based on their privacy preferences. The generation model aims to provide certain privacy guarantees while maximizing the utility of the generated synthetic data at both data level and item level. Specifically, at the data level, we design a selection module that selects those items that contribute less to a user's preferences from the user's interaction data. At the item level, a synthetic data generation module is proposed to generate a synthetic item corresponding to the selected item based on the user's preferences. Furthermore, we also present a privacy-utility trade-off strategy to balance the privacy and utility of the synthetic data. Extensive experiments and ablation studies have been conducted on three publicly accessible datasets to justify our method, demonstrating its effectiveness in generating synthetic data under users' privacy preferences.
Fan Liu 0008, Zhiyong Cheng 0001, Huilin Chen 0002, Yinwei Wei, Liqiang Nie, Mohan Kankanhalli
SIGIR6
2022 Superclass-aware network for few-shot learning
Shuang Wu 0002, Mohan Kankanhalli, Anthony K. H. Tung
Comput. Vis. Image Underst.2
2022 My Health Sensor, My Classifier - Adapting a Trained Classifier to Unlabeled End-User Data
abstract
Sleep apnea is a common yet severely under-diagnosed sleep related disorder. Unattended sleep monitoring at home with low-cost sensors can be leveraged for condition detection, and Machine Learning offers a generalized solution for this task. However, patient characteristics, lack of sufficient training data, and other factors can imply a domain shift between training and end-user data and reduced task performance. In this work, we address this issue with the aim to achieve personalization based on the patient’s needs. We present an unsupervised domain adaptation (UDA) solution with the constraint that labeled source data are not directly available. Instead, a classifier trained on the source data is provided. Our solution iteratively labels target data sub-regions based on classifier beliefs, and trains new classifiers from the expanding dataset. Experiments with sleep monitoring datasets and various sensors show that our solution outperforms the classifier trained on the source domain, with a kappa coefficient improvement from 0.012 to 0.242. Additionally, we apply our solution to digit classification DA between three well-established datasets, to investigate its generalizability, and allow for related work comparisons. Even without direct access to the source data, it outperforms several well-established UDA methods in these datasets.
Konstantinos Nikolaidis, Stein Kristiansen, Thomas Plagemann, Vera Goebel, Knut Liestøl, Mohan Kankanhalli, Gunn Marit Traaen, Britt Øverland, Harriet Akre, Lars Aakerøy, Sigurd Steinshamn
ACM Trans. Comput. Heal.6
2022 One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli
Neurocomputing7
2022 Editorial
Meng Liu 0006, Yan Yan 0014, Tian Gan 0002, Mohan Kankanhalli
Multim. Syst.5
2022 Deep Hierarchical Representation of Point Cloud Videos via Spatio-Temporal Decomposition
abstract
In point cloud videos, point coordinates are irregular and unordered but point timestamps exhibit regularities and order. Grid-based networks for conventional video processing cannot be directly used to model raw point cloud videos. Therefore, in this work, we propose a point-based network that directly handles raw point cloud videos. First, to preserve the spatio-temporal local structure of point cloud videos, we design a point tube covering a local range along spatial and temporal dimensions. By progressively subsampling frames and points and enlarging the spatial radius as the point features are fed into higher-level layers, the point tube can capture video structure in a spatio-temporally hierarchical manner. Second, to reduce the impact of the spatial irregularity on temporal modeling, we decompose space and time when extracting point tube representations. Specifically, a spatial operation is employed to encode the local structure of each spatial region in a tube and a temporal operation is used to encode the dynamics of the spatial regions along the tube. Empirically, the proposed network shows strong performance on 3D action recognition, 4D semantic segmentation and scene flow estimation. Theoretically, we analyse the necessity to decompose space and time in point cloud video modeling and why the network outperforms existing methods.
Hehe Fan, Xin Yu 0002, Yi Yang 0001, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Entropy guided attention network for weakly-supervised action localization
Ying Sun 0001, Hehe Fan, Tao Zhuo, Joo-Hwee Lim, Mohan Kankanhalli
Pattern Recognit.6
2022 Recognition of Advertisement Emotions With Application to Computational Advertising
abstract
Advertisements (ads) often contain strong emotions to capture audience attention and convey an effective message. Still, little work has focused on affect recognition (AR) from ads employing audiovisual or user cues. This work (1) compiles an affective video ad dataset which evokes coherent emotions across users; (2) explores the efficacy of content-centric convolutional neural network (CNN) features for ad AR vis-ã-vis handcrafted audio-visual descriptors; (3) examines user-centric ad AR from Electroencephalogram (EEG) signals, and (4) demonstrates how better affect predictions facilitate effective computational advertising via a study involving 18 users. Experiments reveal that (a) CNN features outperform handcrafted audiovisual descriptors for content-centric AR; (b) EEG features encode ad-induced emotions better than content-based features; (c) Multi-task learning achieves optimal ad AR among a slew of classifiers and (d) Pursuant to (b), EEG features enable optimized ad insertion onto streamed video compared to content-based or manual insertion, maximizing ad recall and viewing experience.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan
IEEE Trans. Affect. Comput.4
2022 Understanding Atomic Hand-Object Interaction With Human Intention
abstract
Hand-object interaction plays a very important role when humans manipulate objects. While existing methods focus on improving hand-object recognition with fully automatic methods, human intention has been largely neglected in the recognition process, thus leading to undesirable interaction descriptions. To better interpret human-object interaction that is aligned to human intention, we argue that a reference specifying human intention should be taken into account. Thus, we propose a new approach to represent interactions while reflecting human purpose with three key factors,i.e., hand, object and reference. Specifically, we design a pattern ofhand-object, object-reference, hand, object, reference> (HOR) to recognize intention based atomic hand-object interactions. This pattern aims to model interactions with the states of hand, object, reference and their relationships. Furthermore, we design a simple yet effective Spatially Part-based (3+1)D convolutional neural network, namely SP(3+1)D, which leverages 3D and 1D convolutions to model visual dynamics and object position changes based on our HOR, respectively. With the help of our SP(3+1)D network, the recognition results are able to indicate human purposes accurately. To evaluate the proposed method, we annotate a Something-1.3k dataset, which contains 10 atomic hand-object interactions and about 130 videos for each interaction. Experimental results on Something-1.3k demonstrate the effectiveness of our SP(3+1)D network.
Hehe Fan, Tao Zhuo, Xin Yu 0002, Yi Yang 0001, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.5
2022 Video Snapshot Compressive Imaging Using Residual Ensemble Network
abstract
Video snapshot compressive imaging (SCI) system enables high-frame-rate imaging by projecting multiple frames into a 2D snapshot measurement during a single exposure, and the original video frames can be reconstructed by solving an optimization problem. However, existing methods usually cannot achieve a good balance between reconstruction time and reconstruction quality, which has become a major obstacle for practical application of video SCI. In order to cope with this issue, we propose a residual ensemble network to learn the explicit inverse mapping from the 2D snapshot measurement to the original video. Specifically, the proposed network aims to exploit the spatiotemporal correlations between video frames for improving reconstruction quality. The spatiotemporal correlations of video frames demonstrate multiple types, including intra-frame spatial correlation, inter-frame forward and backward temporal correlation. With the purpose of fully capturing these differentiated correlations, we design four sub-networks, namely, a pseudo-3D U-shape sub-network, two residual sub-networks, and a serial forward and backward recurrent sub-network, and further assemble these four sub-networks into an ensemble network through alternate residual links. This ensemble network can effectively fuse the predictions of each sub-network and maintain spatiotemporal consistency between video frames. We further design a compound loss function to guide the network learning, and the new video can be fast reconstructed by simply feeding its 2D snapshot measurement into the learned network. The experimental results demonstrate that our network can significantly improve the reconstruction quality while maintaining low computational cost.
Yubao Sun, Xunhao Chen, Mohan Kankanhalli, Qingshan Liu 0001, Junxia Li
IEEE Trans. Circuits Syst. Video Technol.3
2022 Monocular Image-Based 3-D Model Retrieval: A Benchmark
abstract
Monocular image-based 3-D model retrieval aims to search for relevant 3-D models from a dataset given one RGB image captured in the real world, which can significantly benefit several applications, such as self-service checkout, online shopping, etc. To help advance this promising yet challenging research topic, we built a novel dataset and organized the first international contest for monocular image-based 3-D model retrieval. Moreover, we conduct a thorough analysis of the state-of-the-art methods. Existing methods can be classified into supervised and unsupervised methods. The supervised methods can be analyzed based on several important aspects, such as the strategies of domain adaptation, view fusion, loss function, and similarity measure. The unsupervised methods focus on solving this problem with unlabeled data and domain adaptation. Seven popular metrics are employed to evaluate the performance, and accordingly, we provide a thorough analysis and guidance for future work. To the best of our knowledge, this is the first benchmark for monocular image-based 3-D model retrieval, which aims to help related research in multiview feature learning, domain adaptation, and information retrieval.
Dan Song 0006, Weizhi Nie, Wenhui Li 0001, Mohan Kankanhalli, Anan Liu
IEEE Trans. Cybern.4
2022 Unsupervised Spatial-Spectral Network Learning for Hyperspectral Compressive Snapshot Reconstruction
abstract
Hyperspectral compressive imaging takes advantage of compressive sensing theory to achieve coded aperture snapshot measurement without temporal scanning, and the entire 3-D spatial–spectral data is captured by a 2-D projection during a single integration period. Its core issue is how to reconstruct the underlying hyperspectral image (HSI) using compressive sensing reconstruction algorithms. Due to the diversity in the spectral response characteristics and wavelength range of different spectral imaging devices, previous works are often inadequate to capture complex spectral variations or lack the adaptive capacity to new hyperspectral imagers. In order to address these issues, we propose an unsupervised spatial–spectral network to reconstruct HSIs only from the compressive snapshot measurement. The proposed network acts as a conditional generative model conditioned on the snapshot measurement, and it exploits the spatial–spectral attention module to capture the joint spatial–spectral correlation of HSIs. The network parameters are optimized to make sure that the network output can closely match the given snapshot measurement according to the imaging model, thus the proposed network can adapt to different imaging settings, which can inherently enhance the applicability of the network. Extensive experiments upon multiple datasets demonstrate that our network can achieve better reconstruction results than the state-of-the-art methods.
Yubao Sun, Qingshan Liu 0001, Mohan Kankanhalli
IEEE Trans. Geosci. Remote. Sens.4
2022 Relation-Aware Compositional Zero-Shot Learning for Attribute-Object Pair Recognition
abstract
This paper proposes a novel model for recognizing images with composite attribute-object concepts, notably for composite concepts that are unseen during model training. We aim to explore the three key properties required by the task — relation-aware, consistent, and decoupled—to learn rich and robust features for primitive concepts that compose attribute-object pairs. To this end, we propose the Blocked Message Passing Network (BMP-Net). The model consists of two modules. The concept module generates semantically meaningful features for primitive concepts, whereas the visual module extracts visual features for attributes and objects from input images. A message passing mechanism is used in the concept module to capture the relations between primitive concepts. Furthermore, to prevent the model from being biased towards seen composite concepts and reduce the entanglement between attributes and objects, we propose a blocking mechanism that equalizes the information available to the model for both seen and unseen concepts. Extensive experiments and ablation studies on two benchmarks show the efficacy of the proposed model.
Ziwei Xu 0001, Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Multim.4
2022 Toward Region-Aware Attention Learning for Scene Graph Generation
abstract
Scene graph generation (SGGen) is a challenging task due to a complex visual context of an image. Intuitively, the human visual system can volitionally focus on attended regions by salient stimuli associated with visual cues. For example, to infer the relationship between man and horse, the interaction between human leg and horseback can provide strong visual evidence to predict the predicate ride. Besides, the attended region face can also help to determine the object man. Till now, most of the existing works studied the SGGen by extracting coarse-grained bounding box features while understanding fine-grained visual regions received limited attention. To mitigate the drawback, this article proposes a region-aware attention learning method. The key idea is to explicitly construct the attention space to explore salient regions with the object and predicate inferences. First, we extract a set of regions in an image with the standard detection pipeline. Each region regresses to an object. Second, we propose the object-wise attention graph neural network (GNN), which incorporates attention modules into the graph structure to discover attended regions for object inference. Third, we build the predicate-wise co-attention GNN to jointly highlight subject's and object's attended regions for predicate inference. Particularly, each subject-object pair is connected with one of the latent predicates to construct one triplet. The proposed intra-triplet and inter-triplet learning mechanism can help discover the pair-wise attended regions to infer predicates. Extensive experiments on two popular benchmarks demonstrate the superiority of the proposed method. Additional ablation studies and visualization further validate its effectiveness.
Anan Liu, Hongshuo Tian, Ning Xu 0003, Weizhi Nie, Yongdong Zhang 0001, Mohan Kankanhalli
IEEE Trans. Neural Networks Learn. Syst.6
2022 Enhanced 3D Shape Reconstruction With Knowledge Graph of Category Concept
abstract
Reconstructing three-dimensional (3D) objects from images has attracted increasing attention due to its wide applications in computer vision and robotic tasks. Despite the promising progress of recent deep learning–based approaches, which directly reconstruct the full 3D shape without considering the conceptual knowledge of the object categories, existing models have limited usage and usually create unrealistic shapes. 3D objects have multiple forms of representation, such as 3D volume, conceptual knowledge, and so on. In this work, we show that the conceptual knowledge for a category of objects, which represents objects as prototype volumes and is structured by graph, can enhance the 3D reconstruction pipeline. We propose a novel multimodal framework that explicitly combines graph-based conceptual knowledge with deep neural networks for 3D shape reconstruction from a single RGB image. Our approach represents conceptual knowledge of a specific category as a structure-based knowledge graph. Specifically, conceptual knowledge acts as visual priors and spatial relationships to assist the 3D reconstruction framework to create realistic 3D shapes with enhanced details. Our 3D reconstruction framework takes an image as input. It first predicts the conceptual knowledge of the object in the image, then generates a 3D object based on the input image and the predicted conceptual knowledge. The generated 3D object satisfies the following requirements: (1) it is consistent with the predicted graph in concept, and (2) consistent with the input image in geometry. Extensive experiments on public datasets (i.e., ShapeNet, Pix3D, and Pascal3D+) with 13 object categories show that (1) our method outperforms the state-of-the-art methods, (2) our prototype volume-based conceptual knowledge representation is more effective, and (3) our pipeline-agnostic approach can enhance the reconstruction quality of various 3D shape reconstruction pipelines.
Guofei Sun, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud Videos
abstract
Point cloud videos exhibit irregularities and lack of order along the spatial dimension where points emerge inconsistently across different frames. To capture the dynamics in point cloud videos, point tracking is usually employed. However, as points may flow in and out across frames, computing accurate point trajectories is extremely difficult. Moreover, tracking usually relies on point colors and thus may fail to handle colorless point clouds. In this paper, to avoid point tracking, we propose a novel Point 4D Transformer (P4Transformer) network to model raw point cloud videos. Specifically, P4Transformer consists of (i) a point 4D convolution to embed the spatio-temporal local structures presented in a point cloud video and (ii) a transformer to capture the appearance and motion information across the entire video by performing self-attention on the embedded local features. In this fashion, related or similar local areas are merged with attention weight rather than by explicit tracking. Extensive experiments, including 3D action recognition and 4D semantic segmentation, on four benchmarks demonstrate the effectiveness of our P4Transformer for point cloud video modeling.
Hehe Fan, Yi Yang 0001, Mohan Kankanhalli
CVPR3
2021 Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventions
abstract
3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a step towards training robust models for cross-domain pose estimation task, which brings together ideas from causal representation learning and generative adversarial networks. Specifically, this paper introduces a novel framework for causal representation learning which explicitly exploits the causal structure of the task. We consider changing domain as interventions on images under the data-generation process and steer the generative model to produce counterfactual features. This help the model learn transferable and causal relations across different domains. Our framework is able to learn with various types of unlabeled datasets. We demonstrate the efficacy of our proposed method on both human and hand pose estimation task. The experiment results show the proposed approach achieves state-of-the-art performance on most datasets for both domain adaptation and domain generalization settings.
Xiheng Zhang, Yongkang Wong, Juwei Lu, Mohan Kankanhalli, Weidong Geng
ICCV5
2021 PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
Hehe Fan, Xin Yu 0002, Yuhang Ding, Yi Yang 0001, Mohan Kankanhalli
ICLR5
2021 Geometry-aware Instance-reweighted Adversarial Training
Jingfeng Zhang, Jianing Zhu, Gang Niu 0001, Bo Han 0003, Masashi Sugiyama, Mohan Kankanhalli
ICLR6
2021 Effective Abstract Reasoning with Dual-Contrast Network
Tao Zhuo, Mohan Kankanhalli
ICLR2
2021 Human Attributes Prediction under Privacy-preserving Conditions
abstract
Human attributes prediction in visual media is a well-researched topic with a major focus on human faces. However, face images are often of high privacy concern as they can reveal an individual's identity. How to balance this trade-off between privacy and utility is a key problem among researchers and practitioners. In this study, we make one of the first attempts to investigate the human attributes (emotion, age, and gender) prediction under the different de-identification (eyes, lower-face, face, and head obfuscation) privacy scenarios. We first constructed the Diversity in People and Context Dataset (DPaC). We then performed a human study with eye-tracking on how humans recognize facial attributes without the presence of face and context. Results show that in an image, situational context is informative of a target's attributes. Motivated by our human study, we proposed a multi-tasking deep learning model - Context-Guided Human Attributes Prediction (CHAPNet), for human attributes prediction under privacy-preserving conditions. Extensive experiments on DPaC and three commonly used benchmark datasets demonstrate the superiority of CHAPNet in leveraging the situational context for a better interpretation of a target's attributes without the full presence of the target's face. Our research demonstrates the feasibility of visual analytics under de-identification for privacy.
Anshu Singh, Shaojing Fan, Mohan Kankanhalli
ACM Multimedia3
2021 Motion = Video - Content: Towards Unsupervised Learning of Motion Representation from Videos
abstract
Motion, according to its definition in physics, is the change in position with respect to time, regardless of the specific moving object and background. In this paper, we aim to learn appearance-independent motion representation in an unsupervised manner. The main idea is to separate motion from videos while leaving objects and background as content. Specifically, we design an encoder-decoder model which consists of a content encoder, a motion encoder and a video generator. To train the model, we leverage a one-step cycle-consistency in reconstruction within the same video and a two-step cycle-consistency in generation across different videos as self-supervised signals, and use adversarial training to remove the content representation from the motion representation. We demonstrate that the proposed framework can be used for conditional video generation and fine-grained action recognition.
Hehe Fan, Mohan Kankanhalli
MMAsia2
2021 Learning to Predict Trustworthiness with Steep Slope Loss
abstract
Understanding the trustworthiness of a prediction yielded by a classifier is critical for the safe and effective use of AI models. Prior efforts have been proven to be reliable on small-scale datasets. In this work, we study the problem of predicting trustworthiness on real-world large-scale datasets, where the task is more challenging due to high-dimensional features, diverse visual concepts, and a large number of samples. In such a setting, we observe that the trustworthiness predictors trained with prior-art loss functions, i.e., the cross entropy loss, focal loss, and true class probability confidence loss, are prone to view both correct predictions and incorrect predictions to be trustworthy. The reasons are two-fold. Firstly, correct predictions are generally dominant over incorrect predictions. Secondly, due to the data complexity, it is challenging to differentiate the incorrect predictions from the correct ones on real-world large-scale datasets. To improve the generalizability of trustworthiness predictors, we propose a novel steep slope loss to separate the features w.r.t. correct predictions from the ones w.r.t. incorrect predictions by two slide-like curves that oppose each other. The proposed loss is evaluated with two representative deep learning models, i.e., Vision Transformer and ResNet, as trustworthiness predictors. We conduct comprehensive experiments and analyses on ImageNet, which show that the proposed loss effectively improves the generalizability of trustworthiness predictors. The code and pre-trained trustworthiness predictors for reproducibility are available at \url{https://github.com/luoyan407/predict_trustworthiness}.
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
NeurIPS3
2021 Unsupervised Motion Representation Learning with Capsule Autoencoders
abstract
We propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and semantic-agnostic snippets. In the higher level, the snippets are aggregated to form full-length semantic-aware segments. For both levels, we represent motion with a set of learned transformation invariant templates and the corresponding geometric transformations by using capsule autoencoders of a novel design. This leads to a robust and efficient encoding of viewpoint changes. MCAE is evaluated on a novel Trajectory20 motion dataset and various real-world skeleton-based human action datasets. Notably, it achieves better results than baselines on Trajectory20 with considerably fewer parameters and state-of-the-art performance on the unsupervised skeleton-based action recognition task.
Ziwei Xu 0001, Yongkang Wong, Mohan Kankanhalli
NeurIPS4
2021 Learning Realistic Patterns from Visually Unrealistic Stimuli: Generalization and Data Anonymization
abstract
Good training data is a prerequisite to develop useful Machine Learning applications. However, in many domains existing data sets cannot be shared due to privacy regulations (e.g., from medical studies). This work investigates a simple yet unconventional approach for anonymized data synthesis to enable third parties to benefit from such anonymized data. We explore the feasibility of learning implicitly from visually unrealistic, task-relevant stimuli, which are synthesized by exciting the neurons of a trained deep neural network. As such, neuronal excitation can be used to generate synthetic stimuli. The stimuli data is used to train new classification models. Furthermore, we extend this framework to inhibit representations that are associated with specific individuals. We use sleep monitoring data from both an open and a large closed clinical study, and Electroencephalogram sleep stage classification data, to evaluate whether (1) end-users can create and successfully use customized classification models, and (2) the identity of participants in the study is protected. Extensive comparative empirical investigation shows that different algorithms trained on the stimuli are able to generalize successfully on the same task as the original model. Architectural and algorithmic similarity between new and original models play an important role in performance. For similar architectures, the performance is close to that of using the original data (e.g., Accuracy difference of 0.56%-3.82%, Kappa coefficient difference of 0.02-0.08). Further experiments show that the stimuli can provide state-ofthe-art resilience against adversarial association and membership inference attacks.
Konstantinos Nikolaidis, Stein Kristiansen, Thomas Plagemann, Vera Goebel, Knut Liestøl, Mohan Kankanhalli, Gunn Marit Traaen, Britt Øverland, Harriet Akre, Lars Aakerøy, Sigurd Steinshamn
J. Artif. Intell. Res.6
2021 Direction Concentration Learning: Enhancing Congruency in Machine Learning
abstract
One of the well-known challenges in computer vision tasks is the visual diversity of images, which could result in an agreement or disagreement between the learned knowledge and the visual content exhibited by the current observation. In this work, we first define such an agreement in a concepts learning process as congruency. Formally, given a particular task and sufficiently large dataset, the congruency issue occurs in the learning process whereby the task-specific semantics in the training data are highly varying. We propose a Direction Concentration Learning (DCL) method to improve congruency in the learning process, where enhancing congruency influences the convergence path to be less circuitous. The experimental results show that the proposed DCL method generalizes to state-of-the-art models and optimizers, as well as improves the performances of saliency prediction task, continual learning task, and classification task. Moreover, it helps mitigate the catastrophic forgetting problem in the continual learning task. The code is publicly available at https://github.com/luoyan407/congruency.
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 A new DCT-PCM method for license plate number detection in drone images
Hamam Mokayed, Palaiahnakote Shivakumara, Hock Woon Hon, Mohan Kankanhalli, Tong Lu 0002, Umapada Pal 0001
Pattern Recognit. Lett.4
2021 Scene Graph Inference via Multi-Scale Context Modeling
abstract
The scene graph generated for an image structurally represents its object interactions and it substantially aids image scene understanding. To the best of our knowledge, most current works on scene graph generation chiefly focus on pairwise object regions for object and relation inference while ignoring the global visual context outside of these regions. Guided by the intuition that object/relation inference can benefit from the visual context within an image, this paper proposes a multi-scale context modeling method, which can jointly discover and integrate the complementary object-centric and region-centric context for scene graph inference. While both the object-centric and region-centric contexts are separately modeled by their individual modules, a bi-directional message propagation strategy is designed to mutually reinforce the context modeling. A context-fused inference is then proposed to integrate the multi-scale context to guide scene graph inference. Extensive experiments establish that this method can achieve competitive performance compared to the state-of-the-art methods on three benchmarks. Additional ablation studies further validate its effectiveness. Code has been made available at: https://github.com/ningxu1990/MSCM.
Ning Xu 0003, Anan Liu, Yongkang Wong, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.6
2021 Unsupervised Abstract Reasoning for Raven's Problem Matrices
abstract
Raven's Progressive Matrices (RPM) is highly correlated with human intelligence, and it has been widely used to measure the abstract reasoning ability of humans. In this paper, to study the abstract reasoning capability of deep neural networks, we propose the first unsupervised learning method for solving RPM problems. Since the ground truth labels are not allowed, we design a pseudo target based on the prior constraints of the RPM formulation to approximate the ground-truth label, which effectively converts the unsupervised learning strategy into a supervised one. However, the correct answer is wrongly labelled by the pseudo target, and thus the noisy contrast will lead to inaccurate model training. To alleviate this issue, we propose to improve the model performance with negative answers. Moreover, we develop a decentralization method to adapt the feature representation to different RPM problems. Extensive experiments on three datasets demonstrate that our method even outperforms some of the supervised approaches. Our code is available at https://github.com/visiontao/ncd.
Tao Zhuo, Mohan Kankanhalli
IEEE Trans. Image Process.3
2021 Toward Multi-Modal Conditioned Fashion Image Translation
abstract
Having the capability to synthesize photo-realistic fashion product images conditioned on multiple attributes or modalities would bring many new exciting applications. In this work, we propose an end-to-end network architecture that built upon a new generative adversarial network for automatically synthesizing photo-realistic images of fashion products under multiple conditions. Given an input pose image that consists of a 2D skeleton pose and a sentence description of products, our model synthesizes a fashion image preserving the same pose and wearing the fashion products described as the text. Specifically, the generator$G$tries to generate realistic-looking fashion images based on a$\langle \mathsf {pose}, \mathsf {text} \rangle$pair condition to fool the discriminator. An attention network is added for enhancing the generator, which predicts a probability map indicating which part of the image needs to be attended for translation. In contrast, the discriminator$D$distinguishes real images from the translated ones based on the input pose image and text description. The discriminator is divided into two multi-scale sub-discriminators for improving image distinguishing task. Quantitative and qualitative analysis demonstrates that our method is capable of synthesizing realistic images that retain the poses of given images while matching the semantics of provided sentence descriptions.
Xiaoling Gu, Jun Yu 0002, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Multim.4
2021 Adversarial Learning for Personalized Tag Recommendation
abstract
We have recently seen great progress in image classification due to the success of deep convolutional neural networks and the availability of large-scale datasets. Most of the existing work focuses on single-label image classification. However, there are usually multiple tags associated with an image. The existing works on multi-label classification are mainly based on lab curated labels. Humans assign tags to their images differently, which is mainly based on their interests and personal tagging behavior. In this paper, we address the problem of personalized tag recommendation and propose an end-to-end deep network which can be trained on large-scale datasets. The user-preference is learned within the network in an unsupervised way where the network performs joint optimization for user-preference and visual encoding. A joint training of user-preference and visual encoding allows the network to efficiently integrate the visual preference with tagging behavior for a better user recommendation. In addition, we propose the use of adversarial learning, which enforces the network to predict tags resembling user-generated tags. We demonstrate the effectiveness of the proposed model on two different large-scale and publicly available datasets, YFCC100 M and NUS-WIDE. The proposed method achieves significantly better performance on both the datasets when compared to the baselines and other state-of-the-art methods. The code is publicly available at https://github.com/vyzuer/ALTReco.
Erik Quintanilla, Yogesh S. Rawat, Andrey Sakryukin, Mubarak Shah, Mohan Kankanhalli
IEEE Trans. Multim.5
2021 DeepDance: Music-to-Dance Motion Choreography With Adversarial Learning
abstract
The creation of improvised dancing choreographies is an important research field of cross-modal analysis. A key point of this task is how to effectively create and correlate music and dance with a probabilistic one-to-many mapping, which is essential to create realistic dances of various genres. To address this issue, we propose a GAN-based cross-modal association framework, DeepDance, which correlates two different modalities (dance motion and music) together, aiming at creating the desired dance sequence in terms of the input music. Its generator is to predictively produce the dance movements best-fit to current music piece by learning from examples. In another hand, its discriminator acts as an external evaluation from the audience and judges the whole performance. The generated dance movements and the corresponding input music are considered to be well-matched if the discriminator cannot distinguish the generated movements from the training samples according to the estimated probability. By adding motion consistency constraints in our loss function, the proposed framework is able to create long realistic dance sequences. To alleviate the problem of expensive and inefficient data collection, we propose an effective approach to create a large-scale dataset, YouTube-Dance3D, from open data source. Extensive experiments on currently available music-dance datasets and our YouTube-Dance3D dataset demonstrate that our approach effectively captures the correlation between music and dance and can be used to choreograph appropriate dance sequences.
Guofei Sun, Yongkang Wong, Zhiyong Cheng 0001, Mohan Kankanhalli, Weidong Geng
IEEE Trans. Multim.4
2021 A Matrix Factorization Based Framework for Fusion of Physical and Social Sensors
abstract
Our world is witnessing the on-going substantial increase in the number of multimodal physical and social sensors that are ubiquitously distributed and observing or reporting what is happening in their surroundings. These sensors provide massive amounts of spatio-temporal digital footprints which can be analyzed for various tasks such as event detection or situation awareness. However, inherent noise due to the nature of these sensors result in imprecise data and hence imprecise analysis. Also, the heterogeneous data from different modalities, formats and sources make interpreting different levels of information a big challenge. To overcome these limitations, we propose a novel unified matrix factorization-based model to fuse physical and social sensor signals for spatio-temporal analysis. Readings of physical sensor signals are represented by a spatio-temporal situation matrix, which then incorporates social content that can provide explanations for the signal strengths. We test our framework on large-scale real-world data including PSI stations data, traffic CCTV camera images, and tweets for situation prediction as well as for filtering noise to detect events of diverse situations. The experimental results suggest that the proposed matrix factorization approach can utilize the sources correlation, resulting in better performances in various situational understanding tasks.
Francesco Gelli, Christian von der Weth, Mohan Kankanhalli
IEEE Trans. Multim.4
2021 A New Foreground-Background based Method for Behavior-Oriented Social Media Image Classification
abstract
Due to various applications, research on personal traits using information on social media has become an important area. In this paper, a new method for the classification of behavior-oriented social images uploaded on various social media platforms is presented. The proposed method introduces a multimodality concept using skin of different parts of human body and background information, such as indoor and outdoor environments. For each image, the proposed method detects skin candidate components based on R, G, B color spaces and entropy features. The iterative mutual nearest neighbor approach is proposed to detect accurate skin candidate components, which result in foreground components. Next, the proposed method detects the remaining part (other than skin components) as background components based on structure tensor of R, G, B color spaces, and Maximally Stable Extremal Regions (MSER ) concept in the wavelet domain. We then explore Hanman Transform for extracting context features from foreground and background components through clustering and fusion operation. These features are then fed to an SVM classifier for the classification of behavior-oriented images. Comprehensive experiments on 10-class datasets of Normal Behavior-Oriented Social media Image (NBSI) and Abnormal Behavior-Oriented Social media Image (ABSI) show that the proposed method is effective and outperforms the existing methods in terms of average classification rate. Also, the results on the benchmark dataset of five classes of personality traits and two classes of emotions of different facial expressions (FERPlus dataset) demonstrated the robustness of the proposed method over the existing methods.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Divya Krishnani, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.7
2020 COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model Explanations
abstract
Interpretable machine learning models trade -off accuracy for simplicity to make explanations more readable and easier to comprehend. Drawing from cognitive psychology theories in graph comprehension, we formalize readability as visual cognitive chunks to measure and moderate the cognitive load in explanation visualizations. We present Cognitive-GAM (COGAM) to generate explanations with desired cognitive load and accuracy by combining the expressive nonlinear generalized additive models (GAM) with simpler sparse linear models. We calibrated visual cognitive chunks with reading time in a user study, characterized the trade-off between cognitive load and accuracy for four datasets in simulation studies, and evaluated COGAM against baselines with users. We found that COGAM can decrease cognitive load without decreasing accuracy and/or increase accuracy without increasing cognitive load. Our framework and empirical measurement instruments for cognitive load will enable more rigorous assessment of the human interpretability of explainable AI.
Ashraf M. Abdul, Christian von der Weth, Mohan Kankanhalli, Brian Y. Lim
CHI3
2020 n-Reference Transfer Learning for Saliency Prediction
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
ECCV (8)3
2020 Inferring DQN structure for high-dimensional continuous control
abstract
Despite recent advancements in the field of Deep Reinforcement Learning, Deep Q-network (DQN) models still show lackluster performance on problems with high-dimensional action spaces. The problem is even more pronounced for cases with high-dimensional continuous action spaces due to a combinatorial increase in the number of the outputs. Recent works approach the problem by dividing the network into multiple parallel or sequential (action) modules responsible for different discretized actions. However, there are drawbacks to both the parallel and the sequential approaches. Parallel module architectures lack coordination between action modules, leading to extra complexity in the task, while a sequential structure can result in the vanishing gradients problem and exploding parameter space. In this work, we show that the compositional structure of the action modules has a significant impact on model performance. We propose a novel approach to infer the network structure for DQN models operating with high-dimensional continuous actions. Our method is based on the uncertainty estimation techniques introduced in the paper. Our approach achieves state-of-the-art performance on MuJoCo environments with high-dimensional continuous action spaces. Furthermore, we demonstrate the improvement of the introduced approach on a realistic AAA sailing simulator game.
Andrey Sakryukin, Chedy Raïssi, Mohan Kankanhalli
ICML3
2020 Attacks Which Do Not Kill Training Make Adversarial Learning Stronger
abstract
Adversarial training based on the minimax formulation is necessary for obtaining adversarial robustness of trained models. However, it is conservative or even pessimistic so that it sometimes hurts the natural generalization. In this paper, we raise a fundamental question{—}do we have to trade off natural generalization for adversarial robustness? We argue that adversarial training is to employ confident adversarial data for updating the current model. We propose a novel formulation of friendly adversarial training (FAT): rather than employing most adversarial data maximizing the loss, we search for least adversarial data (i.e., friendly adversarial data) minimizing the loss, among the adversarial data that are confidently misclassified. Our novel formulation is easy to implement by just stopping the most adversarial data searching algorithms such as PGD (projected gradient descent) early, which we call early-stopped PGD. Theoretically, FAT is justified by an upper bound of the adversarial risk. Empirically, early-stopped PGD allows us to answer the earlier question negatively{—}adversarial robustness can indeed be achieved without compromising the natural generalization.
Jingfeng Zhang, Xilie Xu, Bo Han 0003, Gang Niu 0001, Masashi Sugiyama, Mohan Kankanhalli
ICML7
2020 The World has Changed - The World Needs to Change. What Multimedia has to Offer for Our Common Digital Future
abstract
Not only the current coronavirus is holding the world in breath. Beyond this current health crisis the world is facing several global challenges from climate change and environmental damage, access to clean water and food, socio-economic inequalities to name a few. The United have very well framed these global challenges in their 17 Sustainability Goals for a future in prosperity and equal opportunities for all, to be achieved by 2030. There is no one simple solution, no one easy cure in sight to address these pressing challenges of our days. Rather a collective approach of all of us is needed which in sum will be contributing to these. Obviously, the field of multimedia has contributed to many tools and applications that are so much in demand these days to stay connected while keeping the distance. But there is much more we can offer to our common digital future. Our future health system, global access to education, decent work, and reducing inequalities are just some of these goals where we our field can contribute. In this panel we will discuss which path we could follow.
Susanne Boll, Hari Sundaram, Svetha Venkatesh, Martha A. Larson, Mohan Kankanhalli
ACM Multimedia5
2020 Helping Users Tackle Algorithmic Threats on Social Media: A Multimedia Research Agenda
abstract
Participation on social media platforms has many benefits but also poses substantial threats. Users often face an unintended loss of privacy, are bombarded with mis-/disinformation, or are trapped in filter bubbles due to over-personalized content. These threats are further exacerbated by the rise of hidden AI-driven algorithms working behind the scenes to shape users' thoughts, attitudes, and behaviour. We investigate how multimedia researchers can help tackle these problems to level the playing field for social media users. We perform a comprehensive survey of algorithmic threats on social media and use it as a lens to set a challenging but important research agenda for effective and real-time user nudging. We further implement a conceptual prototype and evaluate it with experts to supplement our research agenda. This paper calls for solutions that combat the algorithmic threats on social media by utilizing machine learning and multimedia content analysis techniques but in a transparent manner and for the benefit of the users.
Christian von der Weth, Ashraf M. Abdul, Shaojing Fan, Mohan Kankanhalli
ACM Multimedia4
2020 Who You Are Decides How You Tell
abstract
Image captioning is gaining significance in multiple applications such as content-based visual search and chat-bots. Much of the recent progress in this field embraces a data-driven approach without deep consideration of human behavioural characteristics. In this paper, we focus on human-centered automatic image captioning. Our study is based on the intuition that different people will generate a variety of image captions for the same scene, as their knowledge and opinion about the scene may differ. In particular, we first perform a series of human studies to investigate what influences human description of a visual scene. We identify three main factors: a person's knowledge level of the scene, opinion on the scene, and gender. Based on our human study findings, we propose a novel human-centered algorithm that is able to generate human-like image captions. We evaluate the proposed model through traditional evaluation metrics, diversity metrics, and human-based evaluation. Experimental results demonstrate the superiority of our proposed model on generating diverse human-like image captions.
Shuang Wu 0002, Shaojing Fan, Zhiqi Shen 0002, Mohan Kankanhalli, Anthony K. H. Tung
ACM Multimedia4
2020 Locality-Sensitive Hashing Scheme based on Longest Circular Co-Substring
abstract
Locality-Sensitive Hashing (LSH) is one of the most popular methods for c-Approximate Nearest Neighbor Search (c-ANNS) in high-dimensional spaces. In this paper, we propose a novel LSH scheme based on the Longest Circular Co-Substring (LCCS) search framework (LCCS-LSH) with a theoretical guarantee. We introduce a novel concept of LCCS and a new data structure named Circular Shift Array (CSA) for k-LCCS search. The insight of LCCS search framework is that close data objects will have a longer LCCS than the far-apart ones with high probability. LCCS-LSH is LSH-family-independent, and it supports c-ANNS with different kinds of distance metrics. We also introduce a multi-probe version of LCCS-LSH and conduct extensive experiments over five real-life datasets. The experimental results demonstrate that LCCS-LSH outperforms state-of-the-art LSH schemes.
Yifan Lei, Mohan Kankanhalli, Anthony K. H. Tung
SIGMOD Conference3
2020 Weakly-Supervised Multi-Person Action Recognition in 360° Videos
abstract
The recent development of commodity 360° cameras have enabled a single video to capture an entire scene, which endows promising potentials in surveillance scenarios. However, research in omnidirectional video analysis has lagged behind the hardware advances. In this work, we address the important problem of action recognition in topview 360° videos. Due to the wide filed-of-view, 360° videos usually capture multiple people performing actions at the same time. Furthermore, the appearance of people are deformed. The proposed framework first transforms top-view omnidirectional videos into panoramic videos using a calibrationfree method. Then spatial-temporal features are extracted using region-based 3D CNNs for action recognition. We propose a weakly-supervised method based on multiinstance multi-label learning, which trains the model to recognize and localize multiple actions in a video using only video-level action labels as supervision. We perform experiments to quantitatively validate the efficacy of the proposed method over state-of-the-art baselines and variants of our model, and qualitatively demonstrate action localization results. To enable research in this direction, we introduce the 360Action dataset. It is the first omnidirectional video dataset for multi-person action recognition with a diverse set of scenes, actors and actions. The dataset is available at https://github.com/ryukenzen/360action.
Junnan Li 0001, Jianquan Liu, Yongkang Wong, Shoji Nishimura, Mohan Kankanhalli
WACV5
2020 GradMix: Multi-source Transfer across Domains and Tasks
abstract
The computer vision community is witnessing an unprecedented rate of new tasks being proposed and addressed, thanks to the deep convolutional networks' capability to find complex mappings from X to Y. The advent of each task often accompanies the release of a large-scale annotated dataset, for supervised training of deep network. However, it is expensive and time-consuming to manually label sufficient amount of training data. Therefore, it is important to develop algorithms that can leverage off-the-shelf labeled dataset to learn useful knowledge for the target task. While previous works mostly focus on transfer learning from a single source, we study multi-source transfer across domains and tasks (MS-DTT), in a semi-supervised setting. We propose GradMix, a model-agnostic method applicable to any model trained with gradient-based learning rule, to transfer knowledge via gradient descent by weighting and mixing the gradients from all sources during training. GradMix follows a meta-learning objective, which assigns layer-wise weights to the source gradients, such that the combined gradient follows the direction that minimize the loss for a small set of samples from the target dataset. In addition, we propose to adaptively adjust the learning rate for each mini-batch based on its importance to the target task, and a pseudo-labeling method to leverage the unlabeled samples in the target domain. We conduct MS-DTT experiments on two tasks: digit recognition and action recognition, and demonstrate the advantageous performance of the proposed method against multiple baselines.
Junnan Li 0001, Ziwei Xu 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
WACV5
2020 Protecting sensitive place visits in privacy-preserving trajectory publishing
Mohan Kankanhalli
Comput. Secur.2
2020 Evaluating salient object detection in natural images with multiple objects having multi-level saliency
abstract
Salient object detection is evaluated using binary ground truth (GT) with the labels being salient object class and background. In this study, the authors corroborate based on three subjective experiments on a novel image dataset that objects in natural images are inherently perceived to have varying levels of importance. The authors' dataset, named SalMoN (saliency in multi‐object natural images), has 588 images containing multiple objects. The subjective experiments performed record spontaneous attention and perception through eye fixation duration, point clicking and rectangle drawing. As object saliency in a multi‐object image is inherently multi‐level, they propose that salient object detection must be evaluated for the capability to detect all multi‐level salient objects apart from the salient object class detection capability. For this purpose, they generate multi‐level maps as GT corresponding to all the dataset images using the results of the subjective experiments, with the labels being multi‐level salient objects and background. They then propose the use of mean absolute error, Kendall's rank correlation and average area under precision–recall curve to evaluate existing salient object detection methods on their multi‐level saliency GT dataset. Approaches that represent saliency detection on images as local‐global hierarchical processing of a graph perform well in their dataset.
Gökhan Yildirim 0001, Debashis Sen, Mohan Kankanhalli, Sabine Süsstrunk
IET Image Process.3
2020 Visual Social Relationship Recognition
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
Int. J. Comput. Vis.4
2020 Egocentric Analysis of Dash-Cam Videos for Vehicle Forensics
abstract
Video acquisition using dashboard-mounted cameras has recently achieved massive popularity around the world. One of the major developments following the dash-cam's popularity is that videos captured by them can be used as testimony during scenarios, like traffic violations and accidents. The widespread deployment of dash-cams brings new problems ranging from the compromise of privacy by uploading these videos on public websites using videos captured from other cars for making fraudulent claims. Therefore, there is a compelling need to address the problems associated with the usage of dash-cam videos. In this paper, we discuss and highlight the importance of the emerging area of multimedia vehicle forensics. We propose an algorithm for linking a dash-cam video to a specific car. The proposed algorithm is useful for various applications, for example, insurance companies can authenticate the origin of video before processing the claim. In a different scenario of illegitimate video upload on the Web, the video can be traced back to the car it originated from. To this end, we make use of motion blur extracted from dash-cam videos for generating a discriminative feature. We observe that the subtle motion pattern of every vehicle can serve as its unique signature. We extract motion blur from dash-cam videos and use random forest trees for classifying the vehicle correctly. The experimental results on thousands of frames obtained from dash-cam videos of several cars show the effectiveness of our approach. We further investigate the process of forging the signature of a car and propose a counter forensics method to detect such forgery. Also, we discuss the application of our technique to other potential platforms where the camera can be mounted, for example, on the chest of a person. We believe that ours is the first work that describes this new area of research.
Ambuj Mehrish, Puneet Jain, A. Venkata Subramanyam, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.5
2020 Photography and Exploration of Tourist Locations Based on Optimal Foraging Theory
abstract
Animals search for food in their environment with a decision strategy which keeps them fit. Optimal foraging theory models this foraging behavior to determine the optimal decision strategy followed by animals. This theory has been successfully applied for humans as they search for information and is termed as information foraging. When people visit a tourist location, they follow a similar strategy to move from one spot to another, and collect information by capturing photographs. This behavior has similarities with the foraging behavior of animals which has been widely studied by the researchers. In this paper, we propose to employ optimal foraging theory to help tourists explore a location and capture photographs in an optimal way. We determine a decision strategy for tourist which provides a list of interesting spots to visit in a tourist location along with corresponding stay time. Finally, we solve an optimization problem to find a path through these spots which can be followed by tourists. The experimental results on a public dataset demonstrate the effectiveness of the proposed method.1Code available at https://github.com/vyzuer/foraging_theory.
Yogesh S. Rawat, Mubarak Shah, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.3
2020 Unsupervised Online Video Object Segmentation With Motion Property Understanding
abstract
Unsupervised video object segmentation aims to automatically segment moving objects over an unconstrained video without any user annotation. So far, only few unsupervised online methods have been reported in the literature, and their performance is still far from satisfactory because the complementary information from future frames cannot be processed under online setting. To solve this challenging problem, in this paper, we propose a novel unsupervised online video object segmentation (UOVOS) framework by construing the motion property to mean moving in concurrence with a generic object for segmented regions. By incorporating the salient motion detection and the object proposal, a pixel-wise fusion strategy is developed to effectively remove detection noises, such as dynamic background and stationary objects. Furthermore, by leveraging the obtained segmentation from immediately preceding frames, a forward propagation algorithm is employed to deal with unreliable motion detection and object proposals. Experimental results on several benchmark datasets demonstrate the efficacy of the proposed method. Compared to state-of-the-art unsupervised online segmentation algorithms, the proposed method achieves an absolute gain of 6.2%. Moreover, our method achieves better performance than the best unsupervised offline algorithm on the DAVIS-2016 benchmark dataset. Our code is available on the project website: https://www.github.com/visiontao/uovos.
Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Image Process.5
2020 Video Storytelling: Textual Summaries for Events
abstract
Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph generation. In this paper, we introduce the problem of video storytelling, which aims at generating coherent and succinct stories for long videos. Video storytelling introduces new challenges, mainly due to the diversity of the story and the length and complexity of the video. We propose novel methods to address the challenges. First, we propose a context-aware framework for multimodal embedding learning, where we design a residual bidirectional recurrent neural network to leverage contextual information from past and future. The multimodal embedding is then used to retrieve sentences for video clips. Second, we propose a Narrator model to select clips that are representative of the underlying storyline. The Narrator is formulated as a reinforcement learning agent, which is trained by directly optimizing the textual metric of the generated story. We evaluate our method on the video story dataset, a new dataset that we have collected to enable the study. We compare our method with multiple state-of-the-art baselines and show that our method achieves better performance, in terms of quantitative measures and user study.
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
IEEE Trans. Multim.4
2020 Interact as You Intend: Intention-Driven Human-Object Interaction Detection
abstract
The recent advances in instance-level detection tasks lay strong foundation for genuine comprehension of the visual scenes. However, the ability to fully comprehend a social scene is still in its preliminary stage. In this work, we focus on detecting human-object interactions (HOIs) in social scene images, which is demanding in terms of research and increasingly useful for practical applications. To undertake social tasks interacting with objects, humans direct their attention and move their body based on their intention. Based on this observation, we provide a unique computational perspective to explore human intention in HOI detection. Specifically, the proposed human intention-driven HOI detection (iHOI) framework models human pose with the relative distances from body joints to the object instances. It also utilizes human gaze to guide the attended contextual regions in a weakly-supervised setting. In addition, we propose a hard negative sampling strategy to address the problem of mis-grouping. We perform extensive experiments on two benchmark datasets, namely V-COCO and HICO-DET. The efficacy of each proposed component has also been validated.
Bingjie Xu 0002, Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
IEEE Trans. Multim.5
2020 G-Softmax: Improving Intraclass Compactness and Interclass Separability of Features
abstract
Intraclass compactness and interclass separability are crucial indicators to measure the effectiveness of a model to produce discriminative features, where intraclass compactness indicates how close the features with the same label are to each other and interclass separability indicates how far away the features with different labels are. In this paper, we investigate intraclass compactness and interclass separability of features learned by convolutional networks and propose a Gaussian-based softmax ( G -softmax) function that can effectively improve intraclass compactness and interclass separability. The proposed function is simple to implement and can easily replace the softmax function. We evaluate the proposed G -softmax function on classification data sets (i.e., CIFAR-10, CIFAR-100, and Tiny ImageNet) and on multilabel classification data sets (i.e., MS COCO and NUS-WIDE). The experimental results show that the proposed G -softmax function improves the state-of-the-art models across all evaluated data sets. In addition, the analysis of the intraclass compactness and interclass separability demonstrates the advantages of the proposed function over the softmax function, which is consistent with the performance improvement. More importantly, we observe that high intraclass compactness and interclass separability are linearly correlated with average precision on MS COCO and NUS-WIDE. This implies that the improvement of intraclass compactness and interclass separability would lead to the improvement of average precision.
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Emotion-Aware Human Attention Prediction
abstract
Despite the recent success in face recognition and object classification, in the field of human gaze prediction, computer models are still struggling to accurately mimic human attention. One main reason is that visual attention is a complex human behavior influenced by multiple factors, ranging from low-level features (e.g., color, contrast) to high-level human perception (e.g., objects interactions, object sentiment), making it difficult to model computationally. In this work, we investigate the relation between object sentiment and human attention. We first introduce a new evaluation metric (AttI) for measuring human attention that focuses on human fixation consensus. A series of empirical data analyses with AttI indicate that emotion-evoking objects receive attention favor, especially when they co-occur with emotionally-neutral objects, and this favor varies with different image complexity. Based on the empirical analyses, we design a deep neural network for human attention prediction which allows the attention bias on emotion-evoking objects to be encoded in its feature space. Experiments on two benchmark datasets demonstrate its superior performance, especially on metrics that evaluate relative importance of salient regions. This research provides the clearest picture to date on how object sentiments influence human attention, and it makes one of the first attempts to model this phenomenon computationally.
Macario O. Cordel II, Shaojing Fan, Zhiqi Shen 0002, Mohan Kankanhalli
CVPR4
2019 Learning to Learn From Noisy Labeled Data
abstract
Despite the success of deep neural networks (DNNs) in image classification tasks, the human-level performance relies on massive training data with high-quality manual annotations, which are expensive and time-consuming to collect. There exist many inexpensive data sources on the web, but they tend to contain inaccurate labels. Training on noisy labeled datasets causes performance degradation because DNNs can easily overfit to the label noise. To overcome this problem, we propose a noise-tolerant training algorithm, where a meta-learning update is performed prior to conventional gradient update. The proposed meta-learning method simulates actual training by generating synthetic noisy labels, and train the model such that after one gradient update using each set of synthetic noisy labels, the model does not overfit to the specific noise. We conduct extensive experiments on the noisy CIFAR-10 dataset and the Clothing1M dataset. The results demonstrate the advantageous performance of the proposed method compared to several state-of-the-art baselines.
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
CVPR4
2019 Learning to Detect Human-Object Interactions With Knowledge
abstract
The recent advances in instance-level detection tasks lay a strong foundation for automated visual scenes understanding. However, the ability to fully comprehend a social scene still eludes us. In this work, we focus on detecting human-object interactions (HOIs) in images, an essential step towards deeper scene understanding. HOI detection aims to localize human and objects, as well as to identify the complex interactions between them. Innate in practical problems with large label space, HOI categories exhibit a long-tail distribution, i.e., there exist some rare categories with very few training samples. Given the key observation that HOIs contain intrinsic semantic regularities despite they are visually diverse, we tackle the challenge of long-tail HOI categories by modeling the underlying regularities among verbs and objects in HOIs as well as general relationships. In particular, we construct a knowledge graph based on the ground-truth annotations of training dataset and external source. In contrast to direct knowledge incorporation, we address the necessity of dynamic image-specific knowledge retrieval by multi-modal learning, which leads to an enhanced semantic embedding space for HOI comprehension. The proposed method shows improved performance on V-COCO and HICO-DET benchmarks, especially when predicting the rare HOI categories.
Bingjie Xu 0002, Yongkang Wong, Junnan Li 0001, Qi Zhao 0001, Mohan Kankanhalli
CVPR5
2019 CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon Images
abstract
The primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets.
Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli
ICDAR6
2019 Sublinear Time Nearest Neighbor Search over Generalized Weighted Space
abstract
Nearest Neighbor Search (NNS) over generalized weighted space is a fundamental problem which has many applications in various fields. However, to the best of our knowledge, there is no sublinear time solution to this problem. Based on the idea of Asymmetric Locality-Sensitive Hashing (ALSH), we introduce a novel spherical asymmetric transformation and propose the first two novel weight-oblivious hashing schemes SL-ALSH and S2-ALSH accordingly. We further show that both schemes enjoy a quality guarantee and can answer the NNS queries in sublinear time. Evaluations over three real datasets demonstrate the superior performance of the two proposed schemes.
Yifan Lei, Mohan Kankanhalli, Anthony K. H. Tung
ICML3
2019 Towards Robust ResNet: A Small Step but a Giant Leap
abstract
This paper presents a simple yet principled approach to boosting the robustness of the residual network (ResNet) that is motivated by a dynamical systems perspective. Namely, a deep neural network can be interpreted using a partial differential equation, which naturally inspires us to characterize ResNet based on an explicit Euler method. This consequently allows us to exploit the step factor h in the Euler method to control the robustness of ResNet in both its training and generalization. In particular, we prove that a small step factor h can benefit its training and generalization robustness during backpropagation and forward propagation, respectively. Empirical evaluation on real-world datasets corroborates our analytical findings that a small h can indeed improve both its training and generalization robustness.
Jingfeng Zhang, Bo Han 0003, Laura Wynter, Kian Hsiang Low, Mohan Kankanhalli
IJCAI5
2019 LiveSense: Contextual Advertising in Live Streaming Videos
abstract
Live streaming has become a new form of entertainment, which attracts hundreds of millions of users worldwide. The huge amount of multimedia data in live streaming platforms creates tremendous opportunities for online advertising. However, existing state-of-the-art video advertising strategies (e.g., pre-roll and contextual mid-roll advertising) that rely on analyzing the whole video, are not applicable to live streaming videos. This paper describes a novel monetization framework, named LiveSense, for live streaming videos, which is able to display a contextually relevant ad at a suitable timestamp in a non-intrusive way. Specifically, given a live streaming video, we first employ a deep neural network to determine whether the current moment is appropriate for displaying an ad using the historical streaming data. Then, we detect a set of candidate ad insertion areas by incorporating image saliency, background map, and location priorities, so that the ad is displayed over the non-important area. We introduce three types of relevance metrics including textual relevance, global visual relevance and local visual relevance to select the contextually relevant ad. To minimize user intrusiveness, we initially display the ad at a non-important area. If the user is interested in the ad, we will show the ad in an overlaid window with a translucent background. Empirical evaluation on a real-world dataset demonstrates that our proposed framework is able to effectively display ads in live streaming videos while maintaining users' online experience.
Xiang Chen 0010, Tam V. Nguyen 0002, Zhiqi Shen 0002, Mohan Kankanhalli
ACM Multimedia4
2019 Self-supervised Representation Learning Using 360° Data
abstract
The amount of 360-degree panoramas shared online has been rapidly increasing due to the availability of affordable and compact omnidirectional cameras, which offers huge amount of new information unavailable before. In this paper, we present the first work to exploit unlabeled 360-degree data for image representation learning. We propose middle-out, a new self-supervised learning task, which leverages the spatial configuration of normal field-of-view images sampled from a 360-degree image as supervisory signal. We train a Siamese ConvNet model to identify the middle image among three shuffled images sampled from a panorama by perspective projection. Compared to previous self-supervised methods that train models using image patches or video frames with limited field-of-view, our method leverages the rich semantic information contained in 360-degree images and enforces the model to not only learn about objects, but also develop a higher-level understanding about object relationships and scene structures. We quantitatively demonstrate that the feature representation learned using the proposed task is useful for a wide range of vision tasks including object classification, object detection, scene classification, semantic segmentation, and geometry estimation. We also qualitatively show that the proposed method can enforce the ConvNet to extract high-level semantic concepts, an ability which previous self-supervised learning methods have not acquired.
Junnan Li 0001, Jianquan Liu, Yongkang Wong, Shoji Nishimura, Mohan Kankanhalli
ACM Multimedia5
2019 User Diverse Preference Modeling by Multimodal Attentive Metric Learning
abstract
Most existing recommender systems represent a user's preference with a feature vector, which is assumed to be fixed when predicting this user's preferences for different items. However, the same vector cannot accurately capture a user's varying preferences on all items, especially when considering the diverse characteristics of various items. To tackle this problem, in this paper, we propose a novel Multimodal Attentive Metric Learning (MAML) method to model user diverse preferences for various items. In particular, for each user-item pair, we propose an attention neural network, which exploits the item's multimodal features to estimate the user's special attention to different aspects of this item. The obtained attention is then integrated into a metric-based learning method to predict the user preference on this item. The advantage of metric learning is that it can naturally overcome the problem of dot product similarity, which is adopted by matrix factorization (MF) based recommendation models but does not satisfy the triangle inequality property. In addition, it is worth mentioning that the attention mechanism cannot only help model user's diverse preferences towards different items, but also overcome the geometrically restrictive problem caused by collaborative metric learning. Extensive experiments on large-scale real-world datasets show that our model can substantially outperform the state-of-the-art baselines, demonstrating the potential of modeling user diverse preference for recommendation.
Fan Liu 0008, Zhiyong Cheng 0001, Changchang Sun, Yinglong Wang 0001, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia6
2019 Human-imperceptible Privacy Protection Against Machines
abstract
Privacy concerns with social media have recently been under the spotlight, due to a few incidents on user data leakage on social networking platforms. With the current advances in machine learning and big data, computer algorithms often act as a first-step filter for privacy breaches, by automatically selecting content with sensitive information, such as photos that contain faces or vehicle license plate. In this paper we propose a novel algorithm to protect the sensitive attributes against machines, meanwhile keeping the changes imperceptible to humans. In particular, we first conducted a series of human studies to investigate multiple factors that influence human sensitivity to the visual changes. We discover that human sensitivity is influenced by multiple factors, from low-level features such as illumination, texture, to high-level attributes like object sentiment and semantics. Based on our human data, we propose for the first time the concept of human sensitivity map. With the sensitivity map, we design a human-sensitivity-aware image perturbation model, which is able to modify the computational classification results of sensitive attributes while preserving the remaining attributes. Experiments on real world data demonstrate the superior performance of the proposed model on human-imperceptible privacy protection.
Zhiqi Shen 0002, Shaojing Fan, Yongkang Wong, Tian-Tsong Ng, Mohan Kankanhalli
ACM Multimedia5
2019 Unsupervised Domain Adaptation for 3D Human Pose Estimation
abstract
Training an accurate 3D human pose estimator often requires a large amount of 3D ground-truth data which is inefficient and costly to collect. Previous methods have either resorted to weakly supervised methods to reduce the demand of ground-truth data for training, or using synthetically-generated but photo-realistic samples to enlarge the training data pool. Nevertheless, the former methods mainly require either additional supervision, such as unpaired 3D ground-truth data, or the camera parameters in multiview settings. On the other hand, the latter methods require accurately textured models, illumination configurations and background which need careful engineering. To address these problems, we propose a domain adaptation framework with unsupervised knowledge transfer, which aims at leveraging the knowledge in multi-modality data of the easy-to-get synthetic depth datasets to better train a pose estimator on the real-world datasets. Specifically, the framework first trains two pose estimators on synthetically-generated depth images and human body segmentation masks with full supervision, while jointly learning a human body segmentation module from the predicted 2D poses. Subsequently, the learned pose estimator and the segmentation module are applied to the real-world dataset to unsupervisedly learn a new RGB image based 2D/3D human pose estimator. Here, the knowledge encoded in the supervised learning modules are used to regularize a pose estimator without ground-truth annotations. Comprehensive experiments demonstrate significant improvements over weakly supervised methods when no ground-truth annotations are available. Further experiments with ground-truth annotations show that the proposed framework can outperform state-of-the-art fully supervised methods. In addition, we conducted ablation studies to examine the impact of each loss term, as well as with different amount of supervisions signal.
Xiheng Zhang, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
ACM Multimedia3
2019 Explainable Video Action Reasoning via Prior Knowledge and State Transitions
abstract
Human action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel action reasoning framework that uses prior knowledge to explain semantic-level observations of video state changes. Our method takes advantage of both classical reasoning and modern deep learning approaches. Specifically, prior knowledge is defined as the information of a target video domain, including a set of objects, attributes and relationships in the target video domain, as well as relevant actions defined by the temporal attribute and relationship changes (i.e. state transitions). Given a video sequence, we first generate a scene graph on each frame to represent concerned objects, attributes and relationships. Then those scene graphs are associated by tracking objects across frames to form a spatio-temporal graph (also called video graph), which represents semantic-level video states. Finally, by sequentially examining each state transition in the video graph, our method can detect and explain how those actions are executed with prior knowledge, just like the logical manner of thinking by humans. Compared to previous works, the action reasoning results of our method can be explained by both logical rules and semantic-level observations of video content changes. Besides, the proposed method can be used to detect multiple concurrent actions with detailed information, such as who (particular objects), when (time), where (object locations) and how (what kind of changes). Experiments on a re-annotated dataset CAD-120 show the effectiveness of our method.
Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli
ACM Multimedia5
2019 Embedding Symbolic Knowledge into Deep Networks
abstract
In this work, we aim to leverage prior symbolic knowledge to improve the performance of deep models. We propose a graph embedding network that projects propositional formulae (and assignments) onto a manifold via an augmented Graph Convolutional Network (GCN). To generate semantically-faithful embeddings, we develop techniques to recognize node heterogeneity, and semantic regularization that incorporate structural constraints into the embedding. Experiments show that our approach improves the performance of models trained to perform entailment checking and visual relation prediction. Interestingly, we observe a connection between the tractability of the propositional theory representation and the ease of embedding. Future exploration of this connection may elucidate the relationship between knowledge compilation and vector representation learning.
Yaqi Xie 0001, Ziwei Xu 0001, Kuldeep S. Meel, Mohan Kankanhalli, Harold Soh
NeurIPS4
2019 Augmenting Physiological Time Series Data: A Case Study for Sleep Apnea Detection
abstract
Supervised machine learning applications in the health domain often face the problem of insufficient training datasets. The quantity of labelled data is small due to privacy concerns and the cost of data acquisition and labelling by a medical expert. Furthermore, it is quite common that collected data are unbalanced and getting enough data to personalize models for individuals is very expensive or even infeasible. This paper addresses these problems by (1) designing a recurrent Generative Adversarial Network to generate realistic synthetic data and to augment the original dataset, (2) enabling the generation of balanced datasets based on heavily unbalanced dataset, and (3) to control the data generation in such a way that the generated data resembles data from specific individuals. We apply these solutions for sleep apnea detection and study in the evaluation the performance of four well-known techniques, i.e., K-Nearest Neighbour, Random Forest, Multi-Layer Perceptron, and Support Vector Machine. All classifiers exhibit in the experiments a consistent increase in sensitivity and a kappa statistic increase by between 0.007 and 0.182.
Konstantinos Nikolaidis, Stein Kristiansen, Vera Goebel, Thomas Plagemann, Knut Liestøl, Mohan Kankanhalli
ECML/PKDD (3)6
2019 Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
abstract
Benefiting from the advancement of computer vision, natural language processing and information retrieval techniques, visual question answering (VQA), which aims to answer questions about an image or a video, has received lots of attentions over the past few years. Although some progress has been achieved so far, several studies have pointed out that current VQA models are heavily affected by the language prior problem, which means they tend to answer questions based on the co-occurrence patterns of question keywords (e.g., how many) and answers (e.g., 2) instead of understanding images and questions. Existing methods attempt to solve this problem by either balancing the biased datasets or forcing models to better understand images. However, only marginal effects and even performance deterioration are observed for the first and second solution, respectively. In addition, another important issue is the lack of measurement to quantitatively measure the extent of the language prior effect, which severely hinders the advancement of related techniques.
Zhiyong Cheng 0001, Liqiang Nie, Yibing Liu, Yinglong Wang 0001, Mohan Kankanhalli
SIGIR6
2019 LSTM-based multi-label video event detection
Anan Liu, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli
Multim. Tools Appl.6
2019 Music auto-tagging based on the unified latent semantic modeling
Xi Shao, Zhiyong Cheng 0001, Mohan Kankanhalli
Multim. Tools Appl.3
2019 A multi-stream convolutional neural network for sEMG-based gesture recognition in muscle-computer interface
Yongkang Wong, Yu Du 0016, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng
Pattern Recognit. Lett.5
2019 Dual-Stream Recurrent Neural Network for Video Captioning
abstract
Recent progress in using recurrent neural networks (RNNs) for video description has attracted an increasing interest, due to its capability to encode a sequence of frames for caption generation. While existing methods have studied various features (e.g., CNN, 3D CNN, and semantic attributes) for visual encoding, the representation and fusion of heterogeneous information from multi-modal spaces have not fully explored. Consider that different modalities are often asynchronous, frame-level multi-modal fusion (e.g., concatenation and linear fusion) will negatively influence each modality. In this paper, we propose a dual-stream RNN (DS-RNN) framework to jointly discover and integrate the hidden states of both visual and semantic streams for video caption generation. First, an encoding RNN is used for each stream to flexibly exploit the hidden states of respective modality. Specifically, we proposed an attentive multi-grained encoder module to enhance the local feature learning with global semantics feature. Then, a dual-stream decoder is deployed to integrate the asynchronous yet complementary sequential hidden states from both streams for caption generation. Extensive experiments on three benchmark datasets, namely, MSVD, MSR-VTT, and MPII-MD, show that DS-RNN achieves competitive performance against the state-of-the-art. Additional ablation studies were conducted on various variants of the proposed DS-RNN.
Ning Xu 0003, Anan Liu, Yongkang Wong, Yongdong Zhang 0001, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.7
2019 Pricing Average Price Advertising Options When Underlying Spot Market Prices Are Discontinuous
abstract
Advertising options have been recently studied as a special type of guaranteed contracts in online advertising, which are an alternative sales mechanism to real-time auctions. An advertising option is a contract which gives its buyer a right but not obligation to enter into transactions to purchase page views or link clicks at one or multiple pre-specified prices in a specific future period. Different from typical guaranteed contracts, the option buyer pays a lower upfront fee but can have greater flexibility and more control of advertising. Many studies on advertising options so far have been restricted to the situations where the option payoff is determined by the underlying spot market price at a specific time point and the price evolution over time is assumed to be continuous. The former leads to a biased calculation of option payoff and the latter is invalid empirically for many online advertising slots. This paper addresses these two limitations by proposing a new advertising option pricing framework. First, the option payoff is calculated based on an average price over a specific future period. Therefore, the option becomes path-dependent. The average price is measured by the power mean, which contains several existing option payoff functions as its special cases. Second, jump-diffusion stochastic models are used to describe the movement of the underlying spot market price, which incorporate several important statistical properties including jumps and spikes, non-normality, and absence of autocorrelations. A general option pricing algorithm is obtained based on Monte Carlo simulation. In addition, an explicit pricing formula is derived for the case when the option payoff is based on the geometric mean. This pricing formula is also a generalized version of several other option pricing models discussed in related studies [1], [2], [3], [4], [5], [6].
Bowei Chen 0001, Mohan Kankanhalli
IEEE Trans. Knowl. Data Eng.2
2019 Multi-Modal and Multi-Domain Embedding Learning for Fashion Retrieval and Analysis
abstract
Big data analytics has been revolutionizing the fashion industry in recent years. This is evidenced by the fact that popular fashion brands and designers have relied on big data analytics to trace fashion trends and predict market patterns. In this paper, we propose learning a common latent feature representation from heterogeneous fashion data. Specifically, we design a multi-modal and multi-domain embedding learning framework for fashion analysis and data retrieval. Unlike most of the existing multi-view embedding methods, which only consider the heterogeneous similarity constraint, our proposed framework jointly considers both the homogeneous and heterogeneous similarity constraints to capture cross-view similarity and preserve the similarity of the same view. The proposed framework is comprised of two projection steps. In the first projection, a quintuplet-based ranking loss is proposed for multi-domain fashion data to preserve the homogeneous similarity. In the second projection, a cross-view similarity ranking loss is designed for multi-modal fashion data to capture heterogeneous similarity. By utilizing the learned common latent feature representation, the distance between any vector pairs from same or different modalities can reflect its semantic similarity. Quantitative evaluation on a new large-scale dataset and a fashion analysis case study demonstrate the effectiveness of our proposed method.
Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli
IEEE Trans. Multim.6
2019 MMALFM: Explainable Recommendation by Leveraging Reviews and Images
abstract
Personalized rating prediction is an important research problem in recommender systems. Although the latent factor model (e.g., matrix factorization) achieves good accuracy in rating prediction, it suffers from many problems including cold-start, non-transparency, and suboptimal results for individual user-item pairs. In this article, we exploit textual reviews and item images together with ratings to tackle these limitations. Specifically, we first apply a proposed multi-modal aspect-aware topic model (MATM) on text reviews and item images to model users’ preferences and items’ features from different aspects , and also estimate the aspect importance of a user toward an item. Then, the aspect importance is integrated into a novel aspect-aware latent factor model (ALFM), which learns user’s and item’s latent factors based on ratings. In particular, ALFM introduces a weight matrix to associate those latent factors with the same set of aspects in MATM, such that the latent factors could be used to estimate aspect ratings. Finally, the overall rating is computed via a linear combination of the aspect ratings, which are weighted by the corresponding aspect importance. To this end, our model could alleviate the data sparsity problem and gain good interpretability for recommendation. Besides, every aspect rating is weighted by its aspect importance, which is dependent on the targeted user’s preferences and the targeted item’s features. Therefore, it is expected that the proposed method can model a user’s preferences on an item more accurately for each user-item pair. Comprehensive experimental studies have been conducted on the Yelp 2017 Challenge dataset and Amazon product datasets. Results show that (1) our method achieves significant improvement compared to strong baseline methods, especially for users with only few ratings; (2) item visual features can improve the prediction performance—the effects of item image features on improving the prediction results depend on the importance of the visual features for the items; and (3) our model can explicitly interpret the predicted results in great detail.
Zhiyong Cheng 0001, Xiaojun Chang, Lei Zhu 0002, Rose Catherine, Mohan Kankanhalli
ACM Trans. Inf. Syst.5
2019 Attentive Long Short-Term Preference Modeling for Personalized Product Search
abstract
E-commerce users may expect different products even for the same query, due to their diverse personal preferences. It is well known that there are two types of preferences: long-term ones and short-term ones. The former refers to users’ inherent purchasing bias and evolves slowly. By contrast, the latter reflects users’ purchasing inclination in a relatively short period. They both affect users’ current purchasing intentions. However, few research efforts have been dedicated to jointly model them for the personalized product search. To this end, we propose a novel Attentive Long Short-Term Preference model, dubbed as ALSTP, for personalized product search. Our model adopts the neural networks approach to learn and integrate the long- and short-term user preferences with the current query for the personalized product search. In particular, two attention networks are designed to distinguish which factors in the short-term as well as long-term user preferences are more relevant to the current query. This unique design enables our model to capture users’ current search intentions more accurately. Our work is the first to apply attention mechanisms to integrate both long- and short-term user preferences with the given query for the personalized search. Extensive experiments over four Amazon product datasets show that our model significantly outperforms several state-of-the-art product search methods in terms of different evaluation metrics.
Zhiyong Cheng 0001, Liqiang Nie, Yinglong Wang 0001, Jun Ma 0001, Mohan Kankanhalli
ACM Trans. Inf. Syst.6
2019 CloseUp - A Community-Driven Live Online Search Engine
abstract
Search engines are still the most common way of finding information on the Web. However, they are largely unable to provide satisfactory answers to time- and location-specific queries. Such queries can best and often only be answered by humans that are currently on-site. Although online platforms for community question answering are very popular, very few exceptions consider the notion of users’ current physical locations. In this article, we present CloseUp, our prototype for the seamless integration of community-driven live search into a Google-like search experience. Our efforts focus on overcoming the defining differences between traditional Web search and community question answering, namely the formulation of search requests (keyword-based queries vs. well-formed questions) and the expected response times (milliseconds vs. minutes/hours). To this end, the system features a deep learning pipeline to analyze submitted queries and translate relevant queries into questions. Searching users can submit suggested questions to a community of mobile users. CloseUp provides a stand-alone mobile application for submitting, browsing, and replying to questions. Replies from mobile users are presented as live results in the search interface. Using a field study, we evaluated the feasibility and practicability of our approach.
Christian von der Weth, Ashraf M. Abdul, Abhinav Ramesh Kashyap, Mohan Kankanhalli
ACM Trans. Internet Techn.4
2019 A Multi-sensor Framework for Personal Presentation Analytics
abstract
Presentation has been an effective method for delivering information to an audience for many years. Over the past few decades, technological advancements have revolutionized the way humans deliver presentation. Conventionally, the quality of a presentation is usually evaluated through painstaking manual analysis with experts. Although the expert feedback is effective in assisting users to improve their presentation skills, manual evaluation suffers from high cost and is often not available to most individuals. In this work, we propose a novel multi-sensor self-quantification system for presentations, which is designed based on a new proposed assessment rubric. We present our analytics model with conventional ambient sensors (i.e., static cameras and Kinect sensor) and the emerging wearable egocentric sensors (i.e., Google Glass). In addition, we performed a cross-correlation analysis of speaker’s vocal behavior and body language. The proposed framework is evaluated on a new presentation dataset, namely, NUS Multi-Sensor Presentation dataset, which consists of 51 presentations covering a diverse range of topics. To validate the efficacy of the proposed system, we have conducted a series of user studies with the speakers and an interview with an English communication expert, which reveals positive and promising feedback.
Tian Gan 0002, Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.4
2018 Trends and Trajectories for Explainable, Accountable and Intelligible Systems: An HCI Research Agenda
abstract
Advances in artificial intelligence, sensors and big data management have far-reaching societal impacts. As these systems augment our everyday lives, it becomes increasing-ly important for people to understand them and remain in control. We investigate how HCI researchers can help to develop accountable systems by performing a literature analysis of 289 core papers on explanations and explaina-ble systems, as well as 12,412 citing papers. Using topic modeling, co-occurrence and network analysis, we mapped the research space from diverse domains, such as algorith-mic accountability, interpretable machine learning, context-awareness, cognitive psychology, and software learnability. We reveal fading and burgeoning trends in explainable systems, and identify domains that are closely connected or mostly isolated. The time is ripe for the HCI community to ensure that the powerful new autonomous systems have intelligible interfaces built-in. From our results, we propose several implications and directions for future research to-wards this goal.
Ashraf M. Abdul, Jo Vermeulen, Danding Wang, Brian Y. Lim, Mohan Kankanhalli
CHI5
2018 Emotional Attention: A Study of Image Sentiment and Visual Attention
abstract
Image sentiment influences visual perception. Emotion-eliciting stimuli such as happy faces and poisonous snakes are generally prioritized in human attention. However, little research has evaluated the interrelationships of image sentiment and visual saliency. In this paper, we present the first study to focus on the relation between emotional properties of an image and visual attention. We first create the EMOtional attention dataset (EMOd). It is a diverse set of emotion-eliciting images, and each image has (1) eye-tracking data collected from 16 subjects, (2) intensive image context labels including object contour, object sentiment, object semantic category, and high-level perceptual attributes such as image aesthetics and elicited emotions. We perform extensive analyses on EMOd to identify how image sentiment relates to human attention. We discover an emotion prioritization effect: for our images, emotion-eliciting content attracts human attention strongly, but such advantage diminishes dramatically after initial fixation. Aiming to model the human emotion prioritization computationally, we design a deep neural network for saliency prediction, which includes a novel subnetwork that learns the spatial and semantic context of the image scene. The proposed network outperforms the state-of-the-art on three benchmark datasets, by effectively capturing the relative importance of human attention within an image. The code, models, and dataset are available online at https://nus-sesame.top/emotionalattention/.
Shaojing Fan, Zhiqi Shen 0002, Ming Jiang 0019, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001
CVPR6
2018 EEG-based Evaluation of Cognitive Workload Induced by Acoustic Parameters for Data Sonification
abstract
Data Visualization has been receiving growing attention recently, with ubiquitous smart devices designed to render information in a variety of ways. However, while evaluations of visual tools for their interpretability and intuitiveness have been commonplace, not much research has been devoted to other forms of data rendering, \eg, sonification. This work is the first to automatically estimate the cognitive load induced by different acoustic parameters considered for sonification in prior studies~\citeferguson2017evaluation,ferguson2018investigating. We examine cognitive load via (a) perceptual data-sound mapping accuracies of users for the different acoustic parameters, (b) cognitive workload impressions explicitly reported by users, and (c) their implicit EEG responses compiled during the mapping task. Our main findings are that (i) low cognitive load-inducing (ıe, more intuitive) acoustic parameters correspond to higher mapping accuracies, (ii) EEG spectral power analysis reveals higher α band power for low cognitive load parameters, implying a congruent relationship between explicit and implicit user responses, and (iii) Cognitive load classification with EEG features achieves a peak F1-score of 0.64, confirming that reliable workload estimation is achievable with user EEG data compiled using wearable sensors.
Maneesh Bilalpur, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan
ICMI2
2018 Looking Beyond a Clever Narrative: Visual Context and Attention are Primary Drivers of Affect in Video Advertisements
abstract
Emotion evoked by an advertisement plays a key role in influencing brand recall and eventual consumer choices. Automatic ad affect recognition has several useful applications. However, the use of content-based feature representations does not give insights into how affect is modulated by aspects such as the ad scene setting, salient object attributes and their interactions. Neither do such approaches inform us on how humans prioritize visual information for ad understanding. Our work addresses these lacunae by decomposing video content into detected objects, coarse scene structure, object statistics and actively attended objects identified via eye-gaze. We measure the importance of each of these information channels by systematically incorporating related information into ad affect prediction models. Contrary to the popular notion that ad affect hinges on the narrative and the clever use of linguistic and social cues, we find that actively attended objects and the coarse scene structure better encode affective information as compared to individual scene objects or conspicuous background elements.
Abhinav Shukla, Harish Katti, Mohan Kankanhalli, Subramanian Ramanathan
ICMI3
2018 A^3NCF: An Adaptive Aspect Attention Model for Rating Prediction
abstract
Current recommender systems consider the various aspects of items for making accurate recommendations. Different users place different importance to these aspects which can be thought of as a preference/attention weight vector. Most existing recommender systems assume that for an individual, this vector is the same for all items. However, this assumption is often invalid, especially when considering a user's interactions with items of diverse characteristics. To tackle this problem, in this paper, we develop a novel aspect-aware recommender model named A$^3$NCF, which can capture the varying aspect attentions that a user pays to different items. Specifically, we design a new topic model to extract user preferences and item characteristics from review texts. They are then used to 1) guide the representation learning of users and items, and 2) capture a user's special attention on each aspect of the targeted item with an attention network. Through extensive experiments on several large-scale datasets, we demonstrate that our model outperforms the state-of-the-art review-aware recommender systems in the rating prediction task.
Zhiyong Cheng 0001, Xiangnan He 0001, Lei Zhu 0002, Xuemeng Song, Mohan Kankanhalli
IJCAI6
2018 AI + Multimedia Make Better Life?
abstract
No abstract available.
Wen-Huang Cheng, Jiaying Liu 0001, Mohan Kankanhalli, Abdulmotaleb El Saddik, Benoit Huet
ACM Multimedia3
2018 Multi-modal Preference Modeling for Product Search
abstract
The visual preference of users for products has been largely ignored by the existing product search methods. In this work, we propose a multi-modal personalized product search method, which aims to search products which not only are relevant to the submitted textual query, but also match the user preferences from both textual and visual modalities. To achieve the goal, we first leverage the also_view and buy_after_viewing products to construct the visual and textual latent spaces, which are expected to preserve the visual similarity and semantic similarity of products, respectively. We then propose a translation-based search model (TranSearch ) to 1) learn a multi-modal latent space based on the pre-trained visual and textual latent spaces; and 2) map the users, queries and products into this space for direct matching. The TranSearch model is trained based on a comparative learning strategy, such that the multi-modal latent space is oriented to personalized ranking in the training stage. Experiments have been conducted on real-world datasets to validate the effectiveness of our method. The results demonstrate that our method outperforms the state-of-the-art method by a large margin.
Zhiyong Cheng 0001, Liqiang Nie, Xin-Shun Xu, Mohan Kankanhalli
ACM Multimedia5
2018 Unsupervised Learning of View-invariant Action Representations
abstract
The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, label collection is an expensive and time-consuming process. In this work, we propose an unsupervised learning framework, which exploits unlabeled data to learn video representations. Different from previous works in video representation learning, our unsupervised learning task is to predict 3D motion in multiple target views using video representation from a source view. By learning to extrapolate cross-view motions, the representation can capture view-invariant motion dynamics which is discriminative for the action. In addition, we propose a view-adversarial training method to enhance learning of view-invariant features. We demonstrate the effectiveness of the learned representations for action recognition on multiple datasets.
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
NeurIPS4
2018 Aspect-Aware Latent Factor Model: Rating Prediction with Ratings and Reviews
abstract
Although latent factor models (e.g., matrix factorization) achieve good accuracy in rating prediction, they suffer from several problems including cold-start, non-transparency, and suboptimal recommendation for local users or items. In this paper, we employ textual review information with ratings to tackle these limitations. Firstly, we apply a proposed aspect-aware topic model (ATM) on the review text to model user preferences and item features from different aspects, and estimate the aspect importance of a user towards an item. The aspect importance is then integrated into a novel aspect-aware latent factor model (ALFM), which learns user's and item's latent factors based on ratings. In particular, ALFM introduces a weighted matrix to associate those latent factors with the same set of aspects discovered by ATM, such that the latent factors could be used to estimate aspect ratings. Finally, the overall rating is computed via a linear combination of the aspect ratings, which are weighted by the corresponding aspect importance. To this end, our model could alleviate the data sparsity problem and gain good interpretability for recommendation. Besides, an aspect rating is weighted by an aspect importance, which is dependent on the targeted user's preferences and targeted item's features. Therefore, it is expected that the proposed method can model a user's preferences on an item more accurately for each user-item pair locally. Comprehensive experimental studies have been conducted on 19 datasets from Amazon and Yelp 2017 Challenge dataset. Results show that our method achieves significant improvement compared with strong baseline methods, especially for users with only few ratings. Moreover, our model could interpret the recommendation results in depth.
Zhiyong Cheng 0001, Lei Zhu 0002, Mohan Kankanhalli
WWW4
2018 Robust tracking based on H-CNN with low-resource sampling and scaling by frame-wise motion localization
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Kangli Chen, Mohan Kankanhalli
Multim. Tools Appl.6
2018 Saliency flow based video segmentation via motion guided contour refinement
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Mohan Kankanhalli
Signal Process.4
2018 A Spring-Electric Graph Model for Socialized Group Photography
abstract
Visual balance is considered as one of the important factors in defining the aesthetic quality of visual arts. In this paper, we propose a novel method to obtain visual balance in a layout with dynamic visual elements. We use the idea of a spring-electric graph model and augment it with the concept of color energy from the literature of visual arts. We also present an interesting application of the proposed model in photography assistance. We focus on group photography and utilize social media images along with the proposed spring-electric model for providing a recommendation to the user. The proposed method can provide real-time feedback to the user regarding the arrangement of people, their position, and relative size on the image frame. We conducted qualitative experiments along with user studies to evaluate the proposed method. Experimental results and user studies show the effectiveness of the proposed model in obtaining visual balance and group photography recommendation.
Yogesh S. Rawat, Mingli Song, Mohan Kankanhalli
IEEE Trans. Multim.3
2018 Multimodal Multiplatform Social Media Event Summarization
abstract
Social media platforms are turning into important news sources since they provide real-time information from different perspectives. However, high volume, dynamism, noise, and redundancy exhibited by social media data make it difficult to comprehend the entire content. Recent works emphasize on summarizing the content of either a single social media platform or of a single modality (either textual or visual). However, each platform has its own unique characteristics and user base, which brings to light different aspects of real-world events. This makes it critical as well as challenging to combine textual and visual data from different platforms. In this article, we propose summarization of real-world events with data stemming from different platforms and multiple modalities. We present the use of a Markov Random Fields based similarity measure to link content across multiple platforms. This measure also enables the linking of content across time, which is useful for tracking the evolution of long-running events. For the final content selection, summarization is modeled as a subset selection problem. To handle the complexity of the optimal subset selection, we propose the use of submodular objectives. Facets such as coverage, novelty, and significance are modeled as submodular objectives in a multimodal social media setting. We conduct a series of quantitative and qualitative experiments to illustrate the effectiveness of our approach compared to alternative methods.
Akanksha Tiwari, Christian von der Weth, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Dual-Glance Model for Deciphering Social Relationships
abstract
Since the beginning of early civilizations, social relationships derived from each individual fundamentally form the basis of social structure in our daily life. In the computer vision literature, much progress has been made in scene understanding, such as object detection and scene parsing. Recent research focuses on the relationship between objects based on its functionality and geometrical relations. In this work, we aim to study the problem of social relationship recognition, in still images. We have proposed a dual-glance model for social relationship recognition, where the first glance fixates at the individual pair of interest and the second glance deploys attention mechanism to explore contextual cues. We have also collected a new large scale People in Social Context (PISC) dataset, which comprises of 22,670 images and 76,568 annotated samples from 9 types of social relationship. We provide benchmark results on the PISC dataset, and qualitatively demonstrate the efficacy of the proposed model.
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
ICCV4
2017 Multimedia signatures for vehicle forensics
abstract
The use of dashboard-mounted video cameras is rapidly spreading in many countries around the world. Widespread usage of dash-cams brings new problems, for example, dash-cam videos are uploaded on public websites which contain footage of other cars with the number-plates visible. This can potentially compromise privacy. Further, dash-cam videos can be used as evidence in case of accidents. There have been even cases of usage of dash-cam videos for insurance claims. Not only genuine claims can be made but fraudulent claims using some other cars footage can be used. The use as well as misuse of dash-cam videos is going to be widely prevalent in the near future. In this paper, we present a solution to problem of identifying the vehicle in which the dashboard camera is mounted. Our technique can be used by insurance companies for authenticating the source of origin (the vehicle on which the camera is mounted) of video before processing the insurance claim. We make use of features extracted from motion blur as a feature which is generated due to the unique motion of a vehicle. We have found that the subtle motion pattern of every vehicle acts as unique signature. To the best of our knowledge, ours is the first work that describes this new area of research.
Ambuj Mehrish, A. Venkata Subramanyam, Mohan Kankanhalli
ICME3
2017 Evaluating content-centric vs. user-centric ad affect recognition
abstract
Despite the fact that advertisements (ads) often include strongly emotional content, very little work has been devoted to affect recognition (AR) from ads. This work explicitly compares content-centric and user-centric ad AR methodologies, and evaluates the impact of enhanced AR on computational advertising via a user study. Specifically, we (1) compile an affective ad dataset capable of evoking coherent emotions across users; (2) explore the efficacy of content-centric convolutional neural network (CNN) features for encoding emotions, and show that CNN features outperform low-level emotion descriptors; (3) examine user-centered ad AR by analyzing Electroencephalogram (EEG) responses acquired from eleven viewers, and find that EEG signals encode emotional information better than content descriptors; (4) investigate the relationship between objective AR and subjective viewer experience while watching an ad-embedded online video stream based on a study involving 12 users. To our knowledge, this is the first work to (a) expressly compare user vs content-centered AR for ads, and (b) study the relationship between modeling of ad emotions and its impact on a real-life advertising application.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Subramanian Ramanathan
ICMI5
2017 Exploiting Music Play Sequence for Music Recommendation
abstract
Users leave digital footprints when interacting with various music streaming services. Music play sequence, which contains rich information about personal music preference and song similarity, has been largely ignored in previous music recommender systems. In this paper, we explore the effects of music play sequence on developing effective personalized music recommender systems. Towards the goal, we propose to use word embedding techniques in music play sequences to estimate the similarity between songs. The learned similarity is then embedded into matrix factorization to boost the latent feature learning and discovery. Furthermore, the proposed method only considers the k-nearest songs (e.g., k = 5) in the learning process and thus avoids the increase of time complexity. Experimental results on two public datasets demonstrate that our methods could significantly improve the performance of both rating prediction and top-n recommendation tasks.
Zhiyong Cheng 0001, Jialie Shen 0001, Lei Zhu 0002, Mohan Kankanhalli, Liqiang Nie
IJCAI4
2017 Semi-Supervised Learning for Surface EMG-based Gesture Recognition
abstract
Conventionally, gesture recognition based on non-intrusive muscle-computer interfaces required a strongly-supervised learning algorithm and a large amount of labeled training signals of surface electromyography (sEMG). In this work, we show that temporal relationship of sEMG signals and data glove provides implicit supervisory signal for learning the gesture recognition model. To demonstrate this, we present a semi-supervised learning framework with a novel Siamese architecture for sEMG-based gesture recognition. Specifically, we employ auxiliary tasks to learn visual representation; predicting the temporal order of two consecutive sEMG frames; and, optionally, predicting the statistics of 3D hand pose with a sEMG frame. Experiments on the NinaPro, CapgMyo and csl-hdemg datasets validate the efficacy of our proposed approach, especially when the labeled samples are very scarce.
Yu Du 0016, Yongkang Wong, Wenguang Jin, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng
IJCAI6
2017 The Role of Visual Attention in Sentiment Prediction
abstract
Automated assessment of visual sentiment has many applications, such as monitoring social media and facilitating online advertising. In current research on automated visual sentiment assessment, images are mainly input and processed as a whole. However, human attention is biased, and a focal region with high acuity can disproportionately influence visual sentiment. To investigate how attention influences visual sentiment, we conducted experiments that reveal critical insights into human perception. We discover that negative sentiments are elicited by the focal region without a notable influence of contextual information, whereas positive sentiments are influenced by both focal and contextual information. Building on these insights, we create new deep convolutional neural networks for sentiment prediction that have additional channels devoted to encoding focal information. On two benchmark datasets, the proposed models demonstrate superior performance compared with the state-of-the-art methods. Extensive visualizations and statistical analyses indicate that the focal channels are more effective on images with focal objects, especially for images that also elicit negative sentiments.
Shaojing Fan, Ming Jiang 0019, Zhiqi Shen 0002, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001
ACM Multimedia5
2017 Understanding Fashion Trends from Street Photos via Neighbor-Constrained Embedding Learning
abstract
Driven by the increasing popular image-dominated social networks, such as Instagram, Pinterest and Chictopica, sharing of daily-life street photos now plays an influential role in fashion adoption between fashion trend-setters and followers. In this work, we propose a deep learning based fine-grained embedding learning approach for street fashion analysis by leveraging user-generated street fashion data. Specifically, we present QuadNet, an effective CNN based image embedding network driven by both multi-task classification loss and neighbor-constrained similarity loss. The latter loss function is computed with a novel quadruplet loss function, which considers both hard and soft positive neighbors as well as a negative neighbor for each anchor image. The embedded feature learned from co-optimization is effective for both fine-grained classification task and image retrieval task. Quantitative evaluation on a newly collected large-scale multi-task street photo dataset shows that our QuadNet outperforms the state-of-the-art triplet network by a significant margin. In order to further evaluate the effectiveness of the learned embedding, we analyze and trace the fashion trends of New York City from 2011 to 2016. In our analysis, we are able to identify some short-term and long-term fashion styles.
Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli
ACM Multimedia6
2017 Attention Transfer from Web Images for Video Recognition
abstract
Training deep learning based video classifiers for action recognition requires a large amount of labeled videos. The labeling process is labor-intensive and time-consuming. On the other hand, large amount of weakly-labeled images are uploaded to the Internet by users everyday. To harness the rich and highly diverse set of Web images, a scalable approach is to crawl these images to train deep learning based classifier, such as Convolutional Neural Networks (CNN). However, due to the domain shift problem, the performance of Web images trained deep classifiers tend to degrade when directly deployed to videos. One way to address this problem is to fine-tune the trained models on videos, but sufficient amount of annotated videos are still required. In this work, we propose a novel approach to transfer knowledge from image domain to video domain. The proposed method can adapt to the target domain (i.e. video data) with limited amount of training data. Our method maps the video frames into a low-dimensional feature space using the class-discriminative spatial attention map for CNNs. We design a novel Siamese EnergyNet structure to learn energy functions on the attention maps by jointly optimizing two loss functions, such that the attention map corresponding to a ground truth concept would have higher energy. We conduct extensive experiments on two challenging video recognition datasets (i.e. TVHI and UCF101), and demonstrate the efficacy of our proposed method.
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli
ACM Multimedia4
2017 Affect Recognition in Ads with Application to Computational Advertising
abstract
Advertisements (ads) often include strongly emotional content to leave a lasting impression on the viewer. This work (i) compiles an affective ad dataset capable of evoking coherent emotions across users, as determined from the affective opinions of five experts and 14 annotators; (ii) explores the efficacy of convolutional neural network (CNN) features for encoding emotions, and observes that CNN features outperform low-level audio-visual emotion descriptors[9] upon extensive experimentation; and (iii) demonstrates how enhanced affect prediction facilitates computational advertising, and leads to better viewing experience while watching an online video stream embedded with ads based on a study involving 17 users. We model ad emotions based on subjective human opinions as well as objective multimodal features, and show how effectively modeling ad emotions can positively impact a real-life application.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Subramanian Ramanathan
ACM Multimedia5
2017 Optimizing Trade-offs Among Stakeholders in Real-Time Bidding by Incorporating Multimedia Metrics
abstract
Displaying banner advertisements (in short, ads) on webpages has usually been discussed as an Internet economics topic where a publisher uses auction models to sell an online user's page view to advertisers and the one with the highest bid can have her ad displayed to the user. This is also called real-time bidding (RTB) and the ad displaying process ensures that the publisher's benefit is maximized or there is an equilibrium in ad auctions. However, the benefits of the other two stakeholders - the advertiser and the user - have been rarely discussed. In this paper, we propose a two-stage computational framework that selects a banner ad based on the optimized trade-offs among all stakeholders. The first stage is still auction based and the second stage re-ranks ads by considering the benefits of all stakeholders. Our metric variables are: the publisher's revenue, the advertiser's utility, the ad memorability, the ad click-through rate (CTR), the contextual relevance, and the visual saliency. To the best of our knowledge, this is the first work that optimizes trade-offs among all stakeholders in RTB by incorporating multimedia metrics. An algorithm is also proposed to determine the optimal weights of the metric variables. We use both ad auction datasets and multimedia datasets to validate the proposed framework. Our experimental results show that the publisher can significantly improve the other stakeholders' benefits by slightly reducing her revenue in the short-term. In the long run, advertisers and users will be more engaged, the increased demand of advertising and the increased supply of page views can then boost the publisher's revenue.
Xiang Chen 0010, Bowei Chen 0001, Mohan Kankanhalli
SIGIR3
2017 Exploring User-Specific Information in Music Retrieval
abstract
With the advancement of mobile computing technology and cloud-based streaming music service, user-centered music retrieval has become increasingly important. User-specific information has a fundamental impact on personal music preferences and interests. However, existing research pays little attention to the modeling and integration of user-specific information in music retrieval algorithms/models to facilitate music search. In this paper, we propose a novel model, named User-Information-Aware Music Interest Topic (UIA-MIT) model. The model is able to effectively capture the influence of user-specific information on music preferences, and further associate users' music preferences and search terms under the same latent space. Based on this model, a user information aware retrieval system is developed, which can search and re-rank the results based on age- and/or gender-specific music preferences. A comprehensive experimental study demonstrates that our methods can significantly improve the search accuracy over existing text-based music retrieval methods.
Zhiyong Cheng 0001, Jialie Shen 0001, Liqiang Nie, Tat-Seng Chua, Mohan Kankanhalli
SIGIR5
2017 Multi-Camera Action Dataset for Cross-Camera Action Recognition Benchmarking
abstract
Action recognition has received increasing attention from the computer vision and machine learning communities in the last decade. To enable the study of this problem, there exist a vast number of action datasets, which are recorded under controlled laboratory settings, real-world surveillance environments, or crawled from the Internet. Apart from the "in-the-wild" datasets, the training and test split of conventional datasets often possess similar environments conditions, which leads to close to perfect performance on constrained datasets. In this paper, we introduce a new dataset, namely Multi-Camera Action Dataset (MCAD), which is designed to evaluate the open view classification problem under the surveillance environment. In total, MCAD contains 14,298 action samples from 18 action categories, which are performed by 20 subjects and independently recorded with 5 cameras. Inspired by the well received evaluation approach on the LFW dataset, we designed a standard evaluation protocol and benchmarked MCAD under several scenarios. The benchmark shows that while an average of 85% accuracy is achieved under the closed-view scenario, the performance suffers from a significant drop under the cross-view scenario. In the worst case scenario, the performance of 10-fold cross validation drops from 87.0% to 47.4%.
Wenhui Li 0001, Yongkang Wong, Anan Liu, Yang Li 0108, Yuting Su 0001, Mohan Kankanhalli
WACV6
2017 Hierarchical & multimodal video captioning: Discovering and transferring multimodal knowledge for vision to language
Anan Liu, Ning Xu 0003, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli
Comput. Vis. Image Underst.6
2017 Online object tracking based on CNN with spatial-temporal saliency guided sampling
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Kangli Chen, Mohan Kankanhalli
Neurocomputing5
2017 As-similar-as-possible saliency fusion
Tam V. Nguyen 0002, Mohan Kankanhalli
Multim. Tools Appl.2
2017 Content based authentication of visual cryptography
Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.3
2017 Hierarchical Clustering Multi-Task Learning for Joint Human Action Grouping and Recognition
abstract
This paper proposes a hierarchical clustering multi-task learning (HC-MTL) method for joint human action grouping and recognition. Specifically, we formulate the objective function into the group-wise least square loss regularized by low rank and sparsity with respect to two latent variables, model parameters and grouping information, for joint optimization. To handle this non-convex optimization, we decompose it into two sub-tasks, multi-task learning and task relatedness discovery. First, we convert this non-convex objective function into the convex formulation by fixing the latent grouping information. This new objective function focuses on multi-task learning by strengthening the shared-action relationship and action-specific feature learning. Second, we leverage the learned model parameters for the task relatedness measure and clustering. In this way, HC-MTL can attain both optimal action models and group discovery by alternating iteratively. The proposed method is validated on three kinds of challenging datasets, including six realistic action datasets (Hollywood2, YouTube, UCF Sports, UCF50, HMDB51 & UCF101), two constrained datasets (KTH & TJU), and two multi-view datasets (MV-TJU & IXMAS). The extensive experimental results show that: 1) HC-MTL can produce competing performances to the state of the arts for action recognition and grouping; 2) HC-MTL can overcome the difficulty in heuristic action grouping simply based on human knowledge; 3) HC-MTL can avoid the possible inconsistency between the subjective action grouping depending on human knowledge and objective action grouping based on the feature subspace distributions of multiple actions. Comparison with the popular clustered multi-task learning further reveals that the discovered latent relatedness by HC-MTL aids inducing the group-wise multi-task learning and boosts the performance. To the best of our knowledge, ours is the first work that breaks the assumption that all actions are either independent for individual learning or correlated for joint modeling and proposes HC-MTL for automated, joint action grouping and modeling.
Anan Liu, Yuting Su 0001, Weizhi Nie, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 ClickSmart: A Context-Aware Viewpoint Recommendation System for Mobile Photography
abstract
In this paper, we propose ClickSmart, a viewpoint recommendation system that can assist a user in capturing high-quality photographs at well-known tourist locations. ClickSmart can provide real-time viewpoint recommendation based on the preview on the user's camera, current time, and user's geolocation. It makes use of publicly available geotagged images along with the associated metadata for learning a recommendation model. We define view-cells, macroblocks in geospace, and propose the concepts of popularity, quality, and uniqueness of view-cells from the viewpoint perspective. Viewpoint recommendation is generated at the granularity of a view-cell and is based on its popularity, quality, and uniqueness, which are estimated using social media cues associated with images. We further observe that contextual information such as time and weather conditions play an important role in photography, and therefore augment the recommendation system with the associated context. ClickSmart also takes into account the presence of people in the view for making the recommendation. It can provide two kinds of recommendations, quality based and uniqueness based. Although both were found effective in the experimental evaluation, our user study showed that uniqueness-based recommendation was preferred more by skilled photographers compared with amateurs.
Yogesh S. Rawat, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.2
2017 Benchmarking a Multimodal and Multiview and Interactive Dataset for Human Action Recognition
abstract
Human action recognition is an active research area in both computer vision and machine learning communities. In the past decades, the machine learning problem has evolved from conventional single-view learning problem, to cross-view learning, cross-domain learning and multitask learning, where a large number of algorithms have been proposed in the literature. Despite having large number of action recognition datasets, most of them are designed for a subset of the four learning problems, where the comparisons between algorithms can further limited by variances within datasets, experimental configurations, and other factors. To the best of our knowledge, there exists no dataset that allows concurrent analysis on the four learning problems. In this paper, we introduce a novel multimodal and multiview and interactive (M2I) dataset, which is designed for the evaluation of human action recognition methods under all four scenarios. This dataset consists of 1760 action samples from 22 action categories, including nine person-person interactive actions and 13 person-object interactive actions. We systematically benchmark state-of-the-art approaches on M2I dataset on all four learning problems. Overall, we evaluated 13 approaches with nine popular feature and descriptor combinations. Our comprehensive analysis demonstrates that M2I dataset is challenging due to significant intraclass and view variations, and multiple similar action categories, as well as provides solid foundation for the evaluation of existing state-of-the-art algorithms.
Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Cybern.6
2017 Cyber-Physical Social Networks
abstract
In the offline world, getting to know new people is heavily influenced by people’s physical context, that is, their current geolocation. People meet in classes, bars, clubs, public transport, and so on. In contrast, first-generation online social networks such as Facebook or Google+ do not consider users’ context and thus mainly reflect real-world relationships (e.g., family, friends, colleagues). Location-based social networks, or second-generation social networks, such as Foursquare or Facebook Places, take the physical location of users into account to find new friends. However, with the increasing number and wide range of popular platforms and services on the Web, people spend a considerable time moving through the online worlds. In this article, we introduce cyber-physical social networks (CPSN) as the third generation of online social networks. Beside their physical locations, CPSN consider also users’ virtual locations for connecting to new friends. In a nutshell, we regard a web page as a place where people can meet and interact. The intuition is that a web page is a good indicator for a user’s current interest, likings, or information needs. Moreover, we link virtual and physical locations, allowing for users to socialize across the online and offline world. Our main contributions focus on the two fundamental tasks of creating meaningful virtual locations as well as creating meaningful links between virtual and physical locations, where “meaningful” depends on the application scenario. To this end, we present OneSpace, our prototypical implementation of a cyber-physical social network. OneSpace provides a live and social recommendation service for touristic venues (e.g., hotels, restaurants, attractions). It allows mobile users close to a venue and web users browsing online content about the venue to connect and interact in an ad hoc manner. Connecting users based on their shared virtual and physical locations gives way to a plethora of novel use cases for social computing, as we will illustrate. We evaluate our proposed methods for constructing and linking locations and present the results of a first user study investigating the potential impact of cyber-physical social networks.
Christian von der Weth, Ashraf M. Abdul, Mohan Kankanhalli
ACM Trans. Internet Techn.3
2016 Near-Optimal Active Learning of Multi-Output Gaussian Processes
abstract
This paper addresses the problem of active learning of a multi-output Gaussian process (MOGP) model representing multiple types of coexisting correlated environmental phenomena. In contrast to existing works, our active learning problem involves selecting not just the most informative sampling locations to be observed but also the types of measurements at each selected location for minimizing the predictive uncertainty (i.e., posterior joint entropy) of a target phenomenon of interest given a sampling budget. Unfortunately, such an entropy criterion scales poorly in the numbers of candidate sampling locations and selected observations when optimized. To resolve this issue, we first exploit a structure common to sparse MOGP models for deriving a novel active learning criterion. Then, we exploit a relaxed form of submodularity property of our new criterion for devising a polynomial-time approximation algorithm that guarantees a constant-factor approximation of that achieved by the optimal set of selected observations. Empirical evaluation on real-world datasets shows that our proposed approach outperforms existing algorithms for active learning of MOGP and single-output GP models.
Yehong Zhang, Trong Nghia Hoang, Kian Hsiang Low, Mohan Kankanhalli
AAAI4
2016 Marker-Less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps
Yu Du 0016, Yongkang Wong, Feilin Han, Yilin Gui, Zhen Wang 0003, Mohan Kankanhalli, Weidong Geng
ECCV (4)7
2016 Multi-stream Deep Learning Framework for Automated Presentation Assessment
abstract
Presentation is one of the most effective methods to disseminate information. Traditional methods to evaluate the quality of a presentation generally involves a human instructor, which is infeasible in many scenarios. Recent studies have focused on the automated assessment of presentations. A variety of systems have been developed that focus on analyzing various aspects of presentations. However, those systems are mainly limited by their performance, as they mostly adopt hand-crafted features and ad-hoc algorithms. In this work, we propose a multi-stream deep learning framework customized for presentation assessment. The framework uses Bidirectional Long Short-Term Memory with attention mechanism for temporal modeling, and fuses information from multiple modalities for the final decision. We also design a novel assessment rubric based on input from a domain expert. Experimental results on the NUS Multi-Sensor Presentation (NUSMAP) dataset show that the proposed framework is computationally efficient and achieves significant improvement in classification accuracy.
Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli
ISM3
2016 Demo Paper: PreSense - An Assistive Presentation Self-Quantification System
abstract
This technical demo presents PreSense, an Assistive Presentation Self-Quantification System. Oral presentation is traditionally evaluated by a human instructor, which is cost-ineffective and time-consuming. PreSense allows individuals to self-evaluate their presentation skills. The system is designed to receive multimodal inputs from multiple sources, including webcam, Kinect sensor and Google Glass. The multimodal data is processed by a deep assessment framework, which outputs the evaluation results based on a carefully designed assessment rubric. In order to systematically visualize the individual assessment results, we develop an interactive graphical interface and demonstrate its efficacy using presentation data from the NUS Multi-Sensor Presentation (NUSMAP) dataset.
Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli
ISM3
2016 Concept Based Hybrid Fusion of Multimodal Event Signals
abstract
Recent years have seen a significant increase in the number of sensors and resulting event related sensor data, allowing for a better monitoring and understanding of real-world events and situations. Event-related data come from not only physical sensors (e.g., CCTV cameras, webcams) but also from social or microblogging platforms (e.g., Twitter). Given the wide-spread availability of sensors, we observe that sensors of different modalities often independently observe the same events. We argue that fusing multimodal data about an event can be helpful for more accurate detection, localization and detailed description of events of interest. However, multimodal data often include noisy observations, varying information densities and heterogeneous representations, which makes the fusion a challenging task. In this paper, we propose a hybrid fusion approach that takes the spatial and semantic characteristics of sensor signals about events into account. For this, we first adopt the concept of an image-based representation that expresses the situation of particular visual concepts (e.g. "crowdedness", "people marching") called Cmage for both physical and social sensor data. Based on this Cmage representation, we model sparse sensor information using a Gaussian process, fuse multimodal event signals with a Bayesian approach, and incorporate spatial relations between the sensor and social observations. We demonstrate the effectiveness of our approach as a proof-of-concept over real-world data. Our early results show that the proposed approach can reliably reduce the sensor-related noise, locate the event place, improve event detection reliability, and add semantic context so that the fused data provides a better picture of the observed events.
Christian von der Weth, Yehong Zhang, Kian Hsiang Low, Vivek K. Singh 0001, Mohan Kankanhalli
ISM6
2016 ConTagNet: Exploiting User Context for Image Tag Recommendation
abstract
In recent years, deep convolutional neural networks have shown great success in single-label image classification. However, images usually have multiple labels associated with them which may correspond to different objects or actions present in the image. In addition, a user assigns tags to a photo not merely based on the visual content but also the context in which the photo has been captured. Inspired by this, we propose a deep neural network which can predict multiple tags for an image based on the content as well as the context in which the image is captured. The proposed model can be trained end-to-end and solves a multi-label classification problem. We evaluate the model on a dataset of 1,965,232 images which is drawn from the YFCC100M dataset provided by the organizers of Yahoo-Flickr Grand Challenge. We observe a significant improvement in the prediction accuracy after integrating user-context and the proposed model performs very well in the Grand Challenge.
Yogesh S. Rawat, Mohan Kankanhalli
ACM Multimedia2
2016 Visible watermarking based on importance and just noticeable distortion of image regions
Himanshu Agarwal, Debashis Sen, Balasubramanian Raman, Mohan Kankanhalli
Multim. Tools Appl.4
2015 SalAd: A Multimodal Approach for Contextual Video Advertising
abstract
The explosive growth of multimedia data on Internet has created huge opportunities for online video advertising. In this paper, we propose a novel advertising technique called SalAd, which utilizes textual information, visual content and the webpage saliency, to automatically associate the most suitable companion ads with online videos. Unlike most existing approaches that only focus on selecting the most relevant ads, SalAd further considers the saliency of selected ads to reduce intentional ignorance. SalAd consists of three basic steps. Given an online video and a set of advertisements, we first roughly identify a set of relevant ads based on the textual information matching. We then carefully select a sub-set of candidates based on visual content matching. In this regard, our selected ads are contextually relevant to online video content in terms of both textual information and visual content. We finally select the most salient ad among the relevant ads as the most appropriate one. To demonstrate the effectiveness of our method, we have conducted a rigorous eye-tracking experiment on two ad-datasets. The experimental results show that our method enhances the user engagement with the ad content while maintaining users' quality of video viewing experience.
Chen Xiang, Tam V. Nguyen 0002, Mohan Kankanhalli
ISM3
2015 Multi-sensor Self-Quantification of Presentations
abstract
Presentations have been an effective means of delivering information to groups for ages. Over the past few decades, technological advancements have revolutionized the way humans deliver presentations. Despite that, the quality of presentations can be varied and affected by a variety of reasons. Conventional presentation evaluation usually requires painstaking manual analysis by experts. Although the expert feedback can definitely assist users in improving their presentation skills, manual evaluation suffers from high cost and is often not accessible to most people. In this work, we propose a novel multi-sensor self-quantification framework for presentations. Utilizing conventional ambient sensors (i.e., static cameras, Kinect sensor) and the emerging wearable egocentric sensors (i.e., Google Glass), we first analyze the efficacy of each type of sensor with various nonverbal assessment rubrics, which is followed by our proposed multi-sensor presentation analytics framework. The proposed framework is evaluated on a new presentation dataset, namely NUS Multi-Sensor Presentation (NUSMSP) dataset, which consists of 51 presentations covering a diverse set of topics. The dataset was recorded with ambient static cameras, Kinect sensor, and Google Glass. In addition to multi-sensor analytics, we have conducted a user study with the speakers to verify the effectiveness of our system generated analytics, which has received positive and promising feedback.
Tian Gan 0002, Yongkang Wong, Bappaditya Mandal, Vijay Chandrasekhar 0001, Mohan Kankanhalli
ACM Multimedia5
2015 Face Search in Encrypted Domain
Wei Qi Yan 0001, Mohan Kankanhalli
PSIVT2
2015 Tweeting Cameras for Event Detection
abstract
We are living in a world of big sensor data. Due to the widespread prevalence of visual sensors (e.g. surveillance cameras) and social sensors (e.g. Twitter feeds), many events are implicitly captured in real-time by such heterogeneous "sensors". Combining these two complementary sensor streams can significantly improve the task of event detection and aid in comprehending evolving situations. However, the different characteristics of these social and sensor data make such information fusion for event detection a challenging problem. To tackle this problem, we propose an innovative multi-layer tweeting camera framework integrating both physical sensors and social sensors to detect various concepts of real-world events. In this framework, visual concept detectors are applied on camera video frames and these concepts can be construed as "camera tweets" posted regularly. These tweets are represented by a unified probabilistic spatio-temporal (PST) data structure which is then aggregated to a concept-based image (Cmage) as the common representation for visualization. To facilitate event analysis, we define a set of operators and analytic functions that can be applied on the PST data by the user to discover occurrences of events and to analyse evolving situations. We further leverage on geo-located social media data by mining current topics discussed on Twitter to obtain the high-level semantic meaning of detected events in images. We quantitatively evaluate our framework with a large-scale dataset containing images from 150 New York real-time traffic CCTV cameras, university foodcourt camera feeds and Twitter data, which demonstrates the feasibility and effectiveness of the proposed framework. Results of combining camera tweets and social tweets are shown to be promising for detecting real-world events.
Mohan Kankanhalli
WWW2
2015 A bio-inspired center-surround model for salience computation in images
Debashis Sen, Mohan Kankanhalli
J. Vis. Commun. Image Represent.2
2015 Currency security and forensics: a survey
Jarrett Chambers, Wei Qi Yan 0001, Abhimanyu Singh Garhwal, Mohan Kankanhalli
Multim. Tools Appl.4
2015 Salience computation in images based on perceptual distinctness
Debashis Sen, Mohan Kankanhalli
Signal Process. Image Commun.2
2015 Multi-Keyword Multi-Click Advertisement Option Contracts for Sponsored Search
abstract
In sponsored search, advertisement (abbreviated ad) slots are usually sold by a search engine to an advertiser through an auction mechanism in which advertisers bid on keywords. In theory, auction mechanisms have many desirable economic properties. However, keyword auctions have a number of limitations including: the uncertainty in payment prices for advertisers; the volatility in the search engine’s revenue; and the weak loyalty between advertiser and search engine. In this article, we propose a special ad option that alleviates these problems. In our proposal, an advertiser can purchase an option from a search engine in advance by paying an upfront fee, known as the option price. The advertiser then has the right, but no obligation, to purchase among the prespecified set of keywords at the fixed cost-per-clicks (CPCs) for a specified number of clicks in a specified period of time. The proposed option is closely related to a special exotic option in finance that contains multiple underlying assets (multi-keyword) and is also multi-exercisable (multi-click). This novel structure has many benefits: advertisers can have reduced uncertainty in advertising; the search engine can improve the advertisers’ loyalty as well as obtain a stable and increased expected revenue over time. Since the proposed ad option can be implemented in conjunction with the existing keyword auctions, the option price and corresponding fixed CPCs must be set such that there is no arbitrage between the two markets. Option pricing methods are discussed and our experimental results validate the development. Compared to keyword auctions, a search engine can have an increased expected revenue by selling an ad option.
Bowei Chen 0001, Jun Wang 0012, Ingemar J. Cox, Mohan Kankanhalli
ACM Trans. Intell. Syst. Technol.4
2015 Competence-Based Song Recommendation: Matching Songs to One's Singing Skill
abstract
Singing is a popular social activity and a pleasant way of expressing one’s feelings. One important reason for unsuccessful singing performance is because the singer fails to choose a suitable song. In this paper, we propose a novel competence-based song recommendation framework for the purpose of singing. It is distinguished from most existing music recommendation systems which rely on the computation of listeners’ interests or similarity. We model a singer’s vocal competence as a singer profile, which takes voice pitch, intensity, and quality into consideration. Then we propose techniques to acquire singer profiles. We also present a song profile model which is used to construct a human annotated song database. Then we propose a learning-to-rank scheme for recommending songs by a singer profile. Finally, we introduce a reduced singer profile which can greatly simplify the vocal competence modelling process. The experimental study on real singers demonstrates the effectiveness of our approach and its advantages over two baseline methods.
Kuang Mao, Lidan Shou, Ju Fan, Gang Chen 0001, Mohan Kankanhalli
IEEE Trans. Multim.5
2015 Multi-Camera Coordination and Control in Surveillance Systems: A Survey
abstract
The use of multiple heterogeneous cameras is becoming more common in today's surveillance systems. In order to perform surveillance tasks, effective coordination and control in multi-camera systems is very important, and is catching significant research attention these days. This survey aims to provide researchers with a state-of-the-art overview of various techniques for multi-camera coordination and control ( MC 3 ) that have been adopted in surveillance systems. The existing literature on MC 3 is presented through several classifications based on the applicable architectures, frameworks and the associated surveillance tasks. Finally, a discussion on the open problems in surveillance area that can be solved effectively using MC 3 and the future directions in MC 3 research is presented
Prabhu Natarajan, Pradeep K. Atrey, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Context-Aware Photography Learning for Smart Mobile Devices
abstract
In this work we have developed a photography model based on machine learning which can assist a user in capturing high quality photographs. As scene composition and camera parameters play a vital role in aesthetics of a captured image, the proposed method addresses the problem of learning photographic composition and camera parameters. Further, we observe that context is an important factor from a photography perspective, we therefore augment the learning with associated contextual information. The proposed method utilizes publicly available photographs along with social media cues and associated metainformation in photography learning. We define context features based on factors such as time, geolocation, environmental conditions and type of image, which have an impact on photography. We also propose the idea of computing the photographic composition basis, eigenrules and baserules , to support our composition learning. The proposed system can be used to provide feedback to the user regarding scene composition and camera parameters while the scene is being captured. It can also recommend position in the frame where people should stand for better composition. Moreover, it also provides camera motion guidance for pan, tilt and zoom to the user for improving scene composition.
Yogesh S. Rawat, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.2
2014 No One is Left "Unwatched": Fairness in Observation of Crowds of Mobile Targets in Active Camera Surveillance
abstract
Central to the problem of active multi-camera surveillance is the fundamental issue of fairness in the observation of crowds of targets such that no target is “starved” of observation by the cameras for a long time. This paper presents a principled decision-theoretic multi-camera coordination and control (MC2) algorithm called fair-MC2that can coordinate and control the active cameras to achieve max-min fairness in the observation of crowds of targets moving stochastically. Our fair-MC2algorithm is novel in demonstrating how (a) the uncertainty in the locations, directions, speeds, and observation times of the targets arising from the stochasticity of their motion can be modeled probabilistically, (b) the notion of fairness in observing targets can be formally realized in the domain of multi-camera surveillance for the first time by exploiting the max-min fairness metric to formalize our surveillance objective, that is, to maximize the expected minimum observation time over all targets while guaranteeing a predefined image resolution of observing them, and (c) a structural assumption in the state transition dynamics of a surveillance environment can be exploited to improve its scalability to linear time in the number of targets to be observed during surveillance. Empirical evaluation through extensive simulations in realistic surveillance environments shows that fair-MC2outperforms the state-of-the-art and baseline MC2algorithms. We have also demonstrated the feasibility of deploying our fair-MC2algorithm on real AXIS 214 PTZ cameras.
Prabhu Natarajan, Kian Hsiang Low, Mohan Kankanhalli
ECAI3
2014 Nonmyopic \(\epsilon\)-Bayes-Optimal Active Learning of Gaussian Processes
Trong Nghia Hoang, Kian Hsiang Low, Patrick Jaillet, Mohan Kankanhalli
ICML4
2014 Song Recommendation for Social Singing Community
abstract
Nowadays, an increasing number of singing enthusiasts upload their cover songs and share their performances in online social singing communities. They can also listen and rate other users' song renderings. An important feature of the social singing communities is to recommend appropriate singing-songs which users are able to perform excellently.
Kuang Mao, Ju Fan, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli
ACM Multimedia5
2014 Context-Based Photography Learning using Crowdsourced Images and Social Media
abstract
This paper presents a photography model based on machine learning which utilizes crowd-sourced images along with social media cues. As scene composition and camera parameters play a vital role in aesthetics of a captured image, the proposed system addresses the problem of learning photographic composition and camera parameters. Further, we observe that context is an important factor from a photography perspective, we therefore augment the learning with associated contextual information. We define context features based on factors such as time, geo-location, environmental conditions and type of image, which have an impact on photography. The meta information available with crowd-sourced images is utilized for context identification and social media cues are used for photo quality evaluation. We also propose the idea of computing the photographic composition basis, eigenrules and baserules, to support our composition learning method. The trained photography model can provide assistance to the user in determining image composition and camera parameters.
Yogesh S. Rawat, Mohan Kankanhalli
ACM Multimedia2
2014 View-invariant feature discovering for multi-camera human action recognition
abstract
Intelligent video surveillance system is built to automatically detect events of interest, especially on object tracking and behavior understanding. In this paper, we focus on the task of human action recognition under surveillance environment, specifically in a multi-camera monitoring scene. Despite many approaches have achieved success in recognizing human action from video sequences, they are designed for single view and generally not robust against viewpoint invariant. Human action recognition across different views remains challenging due to the large variations from one view to another. We present a framework to solve the problem of transferring action models learned in one view (source view) to another view (target view). First, local space-time interest point feature and global shape-flow feature are extracted as low-level feature, followed by building the hybrid Bag-of-Words model for each action sequence. The data distribution of relevant actions from source view and target view are linked via a cross-view discriminative dictionary learning method. Through the view-adaptive dictionary pair learned by the method, the data from source and target view can be respectively mapped into a common space which is view-invariant. Furthermore, We extend our framework to transfer action models from multiple views to one view when there are multiple source views available. Experiments on the IXMAS human action dataset, which contains videos captured with five viewpoints, show the efficacy of our framework.
Lekha Chaisorn, Yongkang Wong, Anan Liu, Yuting Su 0001, Mohan Kankanhalli
MMSP6
2014 Multi-view action recognition by cross-domain learning
abstract
This paper proposes a novel multi-view human action recognition method by discovering and sharing common knowledge among different video sets captured in multiple viewpoints. To our knowledge, we are the first to treat a specific view as target domain and the others as source domains and consequently formulate the multi-view action recognition into the cross-domain learning framework. First, the classic bag-of-visual word framework is implemented for visual feature extraction in individual viewpoints. Then, we propose a cross-domain learning method with block-wise weighted kernel function matrix to highlight the saliency components and consequently augment the discriminative ability of the model. Extensive experiments are implemented on IXMAS, the popular multi-view action dataset. The experimental results demonstrate that the proposed method can consistently outperform the state of the arts.
Weizhi Nie, Anan Liu, Yuting Su 0001, Lekha Chaisorn, Yongkang Wong, Mohan Kankanhalli
MMSP7
2014 Active Learning Is Planning: Nonmyopic ε-Bayes-Optimal Active Learning of Gaussian Processes
Trong Nghia Hoang, Kian Hsiang Low, Patrick Jaillet, Mohan Kankanhalli
ECML/PKDD (3)4
2014 Guest editorial: Advances in multimedia surveillance
Pradeep K. Atrey, M. Anwar Hossain 0001, Mohan Kankanhalli
Multim. Tools Appl.3
2014 Real-life events in multimedia: detection, representation, retrieval, and applications
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli
Multim. Tools Appl.4
2014 W3-privacy: understanding what, when, and where inference channels in multi-camera surveillance video
Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Mohan Kankanhalli
Multim. Tools Appl.4
2014 Audio Matters in Visual Attention
abstract
There is a dearth of information on how perceived auditory information guides image-viewing behavior. To investigate auditory-driven visual attention, we first generated a human eye-fixation database from a pool of 200 static images and 400 image-audio pairs viewed by 48 subjects. The eye tracking data for the image-audio pairs were captured while participants viewed images, which took place immediately after exposure to coherent/incoherent audio samples. The database was analyzed in terms of time to first fixation, fixation durations on the target object, entropy, AUC, and saliency ratio. It was found that coherent audio information is an important cue for enhancing the feature-specific response to the target object. Conversely, incoherent audio information attenuates this response. Finally, a system predicting the image-viewing with the influence of different audio sources was developed. The detailedly discussed top-down module in the system is composed of auditory estimation based on Gaussian mixture model-maximum a posteriori algorithm-universal background model structure, as well as visual estimation based on the conditional random field model and sparse latent variables. The evaluation experiments show that the proposed models in the system exhibit strong consistency with eye fixations.
Tam V. Nguyen 0002, Mohan Kankanhalli, Shuicheng Yan, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2014 CAVVA: Computational Affective Video-in-Video Advertising
abstract
Advertising is ubiquitous in the online community and more so in the ever-growing and popular online video delivery websites (e.g., YouTube). Video advertising is becoming increasingly popular on these websites. In addition to the existing pre-roll/post-roll advertising and contextual advertising, this paper proposes an in-stream video advertising strategy-Computational Affective Video-in-Video Advertising (CAVVA). Humans being emotional creatures are driven by emotions as well as rational thought. We believe that emotions play a major role in influencing the buying behavior of users and hence propose a video advertising strategy which takes into account the emotional impact of the videos as well as advertisements. Given a video and a set of advertisements, we identify candidate advertisement insertion points (step 1) and also identify the suitable advertisements (step 2) according to theories from marketing and consumer psychology. We formulate this two part problem as a single optimization function in a non-linear 0-1 integer programming framework and provide a genetic algorithm based solution. We evaluate CAVVA using a subjective user-study and eye-tracking experiment. Through these experiments, we demonstrate that CAVVA achieves a good balance between the following seemingly conflicting goals of (a) minimizing the user disturbance because of advertisement insertion while (b) enhancing the user engagement with the advertising content. We compare our method with existing advertising strategies and show that CAVVA can enhance the user's experience and also help increase the monetization potential of the advertising content.
Karthik Yadati, Harish Katti, Mohan Kankanhalli
IEEE Trans. Multim.3
2014 Online Estimation of Evolving Human Visual Interest
abstract
Regions in video streams attracting human interest contribute significantly to human understanding of the video. Being able to predict salient and informative Regions of Interest (ROIs) through a sequence of eye movements is a challenging problem. Applications such as content-aware retargeting of videos to different aspect ratios while preserving informative regions and smart insertion of dialog (closed-caption text) 1 into the video stream can significantly be improved using the predicted ROIs. We propose an interactive human-in-the-loop framework to model eye movements and predict visual saliency into yet-unseen frames. Eye tracking and video content are used to model visual attention in a manner that accounts for important eye-gaze characteristics such as temporal discontinuities due to sudden eye movements, noise, and behavioral artifacts. A novel statistical- and algorithm-based method gaze buffering is proposed for eye-gaze analysis and its fusion with content-based features. Our robust saliency prediction is instantiated for two challenging and exciting applications. The first application alters video aspect ratios on-the-fly using content-aware video retargeting, thus making them suitable for a variety of display sizes. The second application dynamically localizes active speakers and places dialog captions on-the-fly in the video stream. Our method ensures that dialogs are faithful to active speaker locations and do not interfere with salient content in the video stream. Our framework naturally accommodates personalisation of the application to suit biases and preferences of individual users.
Harish Katti, Anoop Kolar Rajagopal, Mohan Kankanhalli, Kalpathi Ramakrishnan
ACM Trans. Multim. Comput. Commun. Appl.3
2014 Up-Fusion: An Evolving Multimedia Fusion Method
abstract
The amount of multimedia data on the Internet has increased exponentially in the past few decades and this trend is likely to continue. Multimedia content inherently has multiple information sources, therefore effective fusion methods are critical for data analysis and understanding. So far, most of the existing fusion methods are static with respect to time, making it difficult for them to handle the evolving multimedia content. To address this issue, in recent years, several evolving fusion methods were proposed, however, their requirements are difficult to meet, making them useful only in limited applications. In this article, we propose a novel evolving fusion method based on the online portfolio selection theory. The proposed method takes into account the correlation among different information sources and evolves the fusion model when new multimedia data is added. It performs effectively on both crisp and soft decisions without requiring additional context information. Extensive experiments on concept detection and human detection tasks over the TRECVID dataset and surveillance data have been conducted and significantly better performance has been obtained.
Xiangyu Wang 0002, Yong Rui, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2013 Seeing through the fence: Image de-fencing using a video sequence
abstract
Tourists and amateur photographers are often hindered in capturing their cherished images/videos by a fence/occlusion that limits accessibility to the scene of interest. The situation has been exacerbated by growing concerns of security at public places and a need exists to provide a tool that can be used for post-processing such “fenced videos” to produce a “de-fenced” image. There are several challenges in this problem and in this work, we identify them as 1. Robust detection of the fence/occlusions. 2. Estimating pixel motion of background scene. 3. Filling in the fence/occlusions by utilizing information in multiple frames of the input video. We use a video captured by a camera panning the scene containing a fence and obtain a “de-fenced” image. Our method can effectively remove fences from images as demonstrated for several synthetic and real-world cases.
Vrushali S. Khasare, Rajiv Ranjan Sahay, Mohan Kankanhalli
ICIP3
2013 Temporal encoded F-formation system for social interaction detection
abstract
In the context of a social gathering, such as a cocktail party, the memorable moments are generally captured by professional photographers or by the participants. The latter case is often undesirable because many participants would rather enjoy the event instead of being occupied by the photo-taking task. Motivated by this scenario, we propose the use of a set of cameras to automatically take photos. Instead of performing dense analysis on all cameras for photo capturing, we first detect the occurrence and location of social interactions via F-formation detection. In the sociology literature, F-formation is a concept used to define social interactions, where each detection only requires the spatial location and orientation of each participant. This information can be robustly obtained with additional Kinect depth sensors. In this paper, we propose an extended F-formation system for robust detection of interactions and interactants. The extended F-formation system employs a heat-map based feature representation for each individual, namely Interaction Space (IS), to model their location, orientation, and temporal information. Using the temporally encoded IS for each detected interactant, we propose a best-view camera selection framework to detect the corresponding best view camera for each detected social interaction. The extended F-formation system is evaluated with synthetic data on multiple scenarios. To demonstrate the effectiveness of the proposed system, we conducted a user study to compare our best view camera ranking with human's ranking using real-world data.
Tian Gan 0002, Yongkang Wong, Daqing Zhang 0001, Mohan Kankanhalli
ACM Multimedia4
2013 Static saliency vs. dynamic saliency: a comparative study
abstract
Recently visual saliency has attracted wide attention of researchers in the computer vision and multimedia field. However, most of the visual saliency-related research was conducted on still images for studying static saliency. In this paper, we give a comprehensive comparative study for the first time of dynamic saliency (video shots) and static saliency (key frames of the corresponding video shots), and two key observations are obtained: 1) video saliency is often different from, yet quite related with, image saliency, and 2) camera motions, such as tilting, panning or zooming, affect dynamic saliency significantly. Motivated by these observations, we propose a novel camera motion and image saliency aware model for dynamic saliency prediction. The extensive experiments on two static-vs-dynamic saliency datasets collected by us show that our proposed method outperforms the state-of-the-art methods for dynamic saliency prediction. Finally, we also introduce the application of dynamic saliency prediction for dynamic video captioning, assisting people with hearing impairments to better entertain videos with only off-screen voices, e.g., documentary films, news videos and sports videos.
Tam V. Nguyen 0002, Mengdi Xu, Guangyu Gao, Mohan Kankanhalli, Qi Tian 0001, Shuicheng Yan
ACM Multimedia4
2013 Interactive Video Advertising: A Multimodal Affective Approach
Karthik Yadati, Harish Katti, Mohan Kankanhalli
MMM (1)3
2013 Image Re-Attentionizing
abstract
In this paper, we propose a computational framework, called Image Re-Attentionizing, to endow the target region in an image with the ability of attracting human visual attention. In particular, the objective is to recolor the target patches by color transfer with naturalness and smoothness preserved yet visual attention augmented. We propose to approach this objective within the Markov Random Field (MRF) framework and an extended graph cuts method is developed to pursue the solution. The input image is first over-segmented into patches, and the patches within the target region as well as their neighbors are used to construct the consistency graphs. Within the MRF framework, the unitary potentials are defined to encourage each target patch to match the patches with similar shapes and textures from a large salient patch database, each of which corresponds to a high-saliency region in one image, while the spatial and color coherence is reinforced as pairwise potentials. We evaluate the proposed method on the direct human fixation data. The results demonstrate that the target region(s) successfully attract human attention and in the meantime both spatial and color coherence is well preserved.
Tam V. Nguyen 0002, Bingbing Ni, Hairong Liu, Jiebo Luo 0001, Mohan Kankanhalli, Shuicheng Yan
IEEE Trans. Multim.6
2013 Multimedia Fusion With Mean-Covariance Analysis
abstract
The number of multimedia applications has been increasing over the past two decades. Multimedia information fusion has therefore attracted significant attention with many techniques having been proposed. However, the uncertainty and correlation among different information sources have not been fully considered in the existing fusion methods. In general, the predictions of individual information source have uncertainty. Furthermore, many information sources in the multimedia systems are correlated with each other. In this paper, we propose a novel multimedia fusion method based on the portfolio theory. Portfolio theory is a widely used financial investment theory dealing with how to allocate funds across securities. The key idea is to maximize the performance of the allocated portfolio while minimize the risk in returns. We adapt this approach to multimedia fusion to derive optimal weights that can achieve good fusion results. The optimization is formulated as a quadratic programming problem. Experimental results with both simulation and real data confirm the theoretical insights and show promising results.
Xiangyu Wang 0002, Mohan Kankanhalli
IEEE Trans. Multim.2
2013 A reward-and-punishment-based approach for concept detection using adaptive ontology rules
abstract
Despite the fact that performance improvements have been reported in the last years, semantic concept detection in video remains a challenging problem. Existing concept detection techniques, with ontology rules, exploit the static correlations among primitive concepts but not the dynamic spatiotemporal correlations. The proposed method rewards (or punishes) detected primitive concepts using dynamic spatiotemporal correlations of the given ontology rules and updates these ontology rules based on the accuracy of detection. Adaptively learned ontology rules significantly help in improving the overall accuracy of concept detection as shown in the experimental result.
Chidansh Amitkumar Bhatt, Pradeep K. Atrey, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2012 Depth Matters: Influence of Depth Cues on Visual Saliency
Congyan Lang, Tam V. Nguyen 0002, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Shuicheng Yan
ECCV (2)5
2012 A Synaesthetic Approach for Image Slideshow Generation
abstract
In this paper, we present a novel automatic image slideshow system that explores a new medium between images and music. It can be regarded as a new image selection and slideshow composition criterion. Based on the idea of ``hearing colors, seeing sounds" from the art of music visualization, equal importance is assigned to image features and audio properties for better synchronization. We minimize the aesthetic energy distance between visual and audio features. Given a set of images, a subset is selected by correlating image features with the input audio properties. The selected images are then synchronized with the music subclips by their audio-visual distance. The inductive image displaying approach has been introduced for common displaying devices.
Yang Yang Xiang, Mohan Kankanhalli
ICME2
2012 Guest editorial: Privacy-aware multimedia surveillance systems
Pradeep K. Atrey, Sabu Emmanuel, Sharad Mehrotra, Mohan Kankanhalli
Multim. Syst.4
2012 Concept-based near-duplicate video clip detection for novelty re-ranking of web video search results
Chidansh Amitkumar Bhatt, Pradeep K. Atrey, Mohan Kankanhalli
Multim. Syst.3
2012 Introduction to the special issue of the multimedia tools and applications journal on events in multimedia
Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli
Multim. Tools Appl.3
2012 Robust Watermarking of Compressed and Encrypted JPEG2000 Images
abstract
Digital asset management systems (DAMS) generally handle media data in a compressed and encrypted form. It is sometimes necessary to watermark these compressed encrypted media items in the compressed-encrypted domain itself for tamper detection or ownership declaration or copyright management purposes. It is a challenge to watermark these compressed encrypted streams as the compression process would have packed the information of raw media into a low number of bits and encryption would have randomized the compressed bit stream. Attempting to watermark such a randomized bit stream can cause a dramatic degradation of the media quality. Thus it is necessary to choose an encryption scheme that is both secure and will allow watermarking in a predictable manner in the compressed encrypted domain. In this paper, we propose a robust watermarking algorithm to watermark JPEG2000 compressed and encrypted images. The encryption algorithm we propose to use is a stream cipher. While the proposed technique embeds watermark in the compressed-encrypted domain, the extraction of watermark can be done in the decrypted domain. We investigate in detail the embedding capacity, robustness, perceptual quality and security of the proposed algorithm, using these watermarking schemes: Spread Spectrum (SS), Scalar Costa Scheme Quantization Index Modulation (SCS-QIM), and Rational Dither Modulation (RDM).
A. Venkata Subramanyam, Sabu Emmanuel, Mohan Kankanhalli
IEEE Trans. Multim.3
2012 Adaptive Workload Equalization in Multi-Camera Surveillance Systems
abstract
Surveillance and monitoring systems generally employ a large number of cameras to capture people's activities in the environment. These activities are analyzed by hosts (human operators and/or computers) for threat detection. Threat detection is a target centric task in which the behavior of each target is analyzed separately, which requires a significant amount of human attention and is a computationally intensive task for automatic analysis. In order to meet the real-time requirements of surveillance, it is necessary to distribute the video processing load over multiple hosts. In general, cameras are statically assigned to the hosts; we show that this is not a desirable solution as the workload for a particular camera may vary over time depending on the number of targets in its view. In the future, this uneven distribution of workload will become more critical as the sensing infrastructures are being deployed on the cloud. In this paper, we model the camera workload as a function of the number of targets, and use that to dynamically assign video feeds to the hosts. Experimental results show that the proposed model successfully captures the variability of the workload, and that the dynamic workload assignment provides better results than a static assignment.
Mukesh Saini, Xiangyu Wang 0002, Pradeep K. Atrey, Mohan Kankanhalli
IEEE Trans. Multim.4
2012 Introduction to special issue on multimedia security
abstract
No abstract available.
Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Aggregate licenses validation for digital rights violation detection
abstract
Digital Rights Management (DRM) is the term associated with the set of technologies to prevent illegal multimedia content distribution and consumption. DRM systems generally involve multiple parties such as owner, distributors, and consumers. The owner issues redistribution licenses to its distributors. The distributors in turn using their received redistribution licenses can generate and issue new redistribution licenses to other distributors and new usage licenses to consumers. As a part of rights violation detection, these newly generated licenses must be validated by a validation authority against the redistribution license used to generate them. The validation of these newly generated licenses becomes quite complex when there exist multiple redistribution licenses for a media with the distributors. In such cases, the validation process requires validation using an exponential number (to the number of redistribution licenses) of validation inequalities and each validation inequality may contain up to an exponential number of summation terms. This makes the validation process computationally intensive and necessitates to do the validation efficiently. To overcome this, we propose validation tree , a prefix-tree-based validation method to do the validation efficiently. Theoretical analysis and experimental results show that our proposed technique reduces the validation time significantly.
Amit Sachan, Sabu Emmanuel, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2012 Image hatching for visual cryptography
abstract
Image hatching (or nonphotorealistic line-art) is a technique widely used in the printing or engraving of currency. Diverse styles of brush strokes have previously been adopted for different areas of an image to create aesthetically pleasing textures and shading. Because there is no continuous tone within these types of images, a multilevel scheme is proposed, which uses different textures based on a threshold level. These textures are then applied to the different levels and are then combined to build up the final hatched image. The proposed technique allows a secret to be hidden using Visual Cryptography (VC) within the hatched images. Visual cryptography provides a very powerful means by which one secret can be distributed into two or more pieces known as shares. When the shares are superimposed exactly together, the original secret can be recovered without computation. Also provided is a comparison between the original grayscale images and the resulting hatched images that are generated by the proposed algorithm. This reinforces that the overall quality of the hatched scheme is sufficient. The Structural SIMilarity index (SSIM) is used to perform this comparison.
Jonathan Weir, Wei Qi Yan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2011 Anonymous surveillance
abstract
Video surveillance is a very effective tool of surveillance that enables a single security agent to monitor wide areas. However, it compromises the privacy of the individuals. There have been attempts to obfuscate face and silhouette regions of the images to hide the identity of individuals. We recognize that in traditional surveillance systems, the viewer generally has sufficient contextual knowledge about location of the camera, time, and activity patterns; which can lead to identity leakage even when the visual cues (face and appearance) are not present. In this way, the viewer can relate the identity of individuals to the sensitive information in the video causing privacy loss. In order to provide robust privacy preservation, the context knowledge needs to be decoupled from the video; however, human monitoring of the videos is also necessary for the assessment of the situation. In this paper we propose anonymous surveillance framework that decouples the contextual knowledge and video to the minimal extent required for situation assessment. The experimental results confirm that the proposed framework is very effective in protecting the privacy, yet does not affect much of the surveillance utility of the data.
Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Mohan Kankanhalli
ICME4
2011 Dynamic workload assignment in video surveillance systems
abstract
Current surveillance systems consist of large numbers of cameras. The video feeds from cameras are automatically processed for threat detection, which is a computationally intensive task. In order to meet the real-time requirements of surveillance, we need to distribute the video processing over multiple computers. Generally the cameras are statically assigned to the processors; we show that this is not a desirable solution as the workload for a particular camera may vary over time depending on the number of the targets in its view. In future, this uneven distribution of workload will become more critical as the sensing infrastructures are being deployed on the cloud. In this work, we model the camera workload as a function of the number of targets, and use that to dynamically assign video feeds to the processors. Experimental results show that the proposed model successfully captures the variability of the workload, and that dynamic workload assignment provides better results than a static assignment.
Mukesh Saini, Xiangyu Wang 0002, Pradeep K. Atrey, Mohan Kankanhalli
ICME4
2011 Affective Video Summarization and Story Board Generation Using Pupillary Dilation and Eye Gaze
abstract
We propose a semi-automated, eye-gaze based method for affective analysis of videos. Pupillary Dilation (PD) is introduced as a valuable behavioural signal for assessment of subject arousal and engagement. We use PD information for computationally inexpensive, arousal based composition of video summaries and descriptive story-boards. Video summarization and story-board generation is done offline, subsequent to a subject viewing the video. The method also includes novel eye-gaze analysis and fusion with content based features to discover affective segments of videos and Regions of interest (ROIs) contained therein. Effectiveness of the framework is evaluated using experiments over a diverse set of clips, significant pool of subjects and comparison with a fully automated state-of-art affective video summarization algorithm. Acquisition and analysis of PD information is demonstrated and used as a proxy for human visual attention and arousal based video summarization and story-board generation. An important contribution is to demonstrate usefulness of PD information in identifying affective video segments with abstract semantics or affective elements of discourse and story-telling, that are likely to be missed by automated methods. Another contribution is the use of eye-fixations in the close temporal proximity of PD based events for key frame extraction and subsequent story board generation. We also show how PD based video summarization can to generate either a personalized video summary or to represent a consensus over affective preferences of a larger group or community.
Harish Katti, Karthik Yadati, Mohan Kankanhalli, Tat-Seng Chua
ISM3
2011 Eye-tracking methodology and applications to images and video
abstract
Our tutorial introduces eye-tracking as an exciting, non-intrusive method of capturing user attention during human interaction with digital images and videos. We believe eye-gaze can play a valuable role in understanding and processing (a) huge volumes of image and video content generated as a result of human experiences and interaction with the environment (b) Personalization and in human-media interaction, having access to individual preferences and behavioral patterns would be a key component of such a system. Recent possibilities to seamlessly integrate eye-tracking into laptops and mobile devices opens up a plethora of possibilities for applications that can respond to user's visual attention strategies. (c) Visual content design such as in advertising often employs techniques that guide user attention to produce visual impact and elements of surprise and emotion. Eye-gaze has been used as a tool to evaluate different choices of visual elements and their placement. (d) Affective analysis of images and videos is an ongoing and challenging area in multimedia research, show recent results on how eye-gaze and accompanying pupillary dilation information can aid affective analysis.
Harish Katti, Mohan Kankanhalli
ACM Multimedia2
2011 Modeling and representing events in multimedia
abstract
This paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia.
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang
ACM Multimedia4
2011 Up-fusion: an evolving multimedia decision fusion method
abstract
The amount of multimedia data available on the Internet has increased exponentially in the past few decades and is likely to keep on increasing. Given that a multimedia system has multiple information sources, fusion methods are critical for its analysis and understanding. However, most of the traditional fusion methods are static with respect to time. To address this, in recent years, several evolving fusion methods have been proposed. However, they can only be used in limited scenarios. For example, the context aware fusion methods need the context information to update the fusion model, but the context may not always be available in many applications. In this paper, a new evolving fusion method is proposed based on the online portfolio selection theory. The proposed method takes the correlation among different information sources into account, and evolves the fusion model when new multimedia data is added. It can deal with either crisp or soft decisions without requiring additional context information. Extensive experiments on concept detection task over TRECVID dataset have been conducted, and promising results have been obtained.
Xiangyu Wang 0002, Yong Rui, Mohan Kankanhalli
ACM Multimedia3
2011 Affect-based adaptive presentation of home videos
abstract
In recent times, the proliferation of multimedia devices and reduced costs of data storage have enabled people to easily record and collect a large number of home videos; furthermore, this collection is growing with time. With the popularity of participatory media such as YouTube and facebook, problems are encountered when people intend to share their home videos with others. The first problem is that different people might be interested in different video content. Given the numbers of home videos, it is a time-consuming and hard task to manually select proper content for people with different interests. Secondly, as short videos are becoming more and more popular in media sharing applications, people need to manually cut and edit home videos which is again a tedious task. In this paper, we propose a method that employs affective analysis to automatically create video presentations from home videos. Our novel method adaptively creates presentations based on three properties: emotional tone, local main character and global main character. A novel sparsity-based affective labeling method is proposed to identify the emotional content of the videos. The local and global main characters are determined by applying face recognition in each shot. To demonstrate the proposed method, three kinds of presentations are created for family, acquaintance and outsider. Experimental results show that our method is very effective in video sharing and the users are satisfied with the videos generated by our method.
Xiaohong Xiang, Mohan Kankanhalli
ACM Multimedia2
2011 Pedestrian Tracking Based on Hidden-Latent Temporal Markov Chain
Peng Zhang 0005, Sabu Emmanuel, Mohan Kankanhalli
MMM (2)3
2011 Effective multimedia surveillance using a human-centric approach
Pradeep K. Atrey, Abdulmotaleb El Saddik, Mohan Kankanhalli
Multim. Tools Appl.3
2011 Multimedia data mining: state of the art and challenges
Chidansh Amitkumar Bhatt, Mohan Kankanhalli
Multim. Tools Appl.2
2011 Probabilistic temporal multimedia data mining
abstract
Existing sequence pattern mining techniques assume that the obtained events from event detectors are accurate. However, in reality, event detectors label the events from different modalities with a certain probability over a time-interval. In this article, we consider for the first time Probabilistic Temporal Multimedia (PTM) Event data to discover accurate sequence patterns. PTM event data considers the start time, end time, event label and associated probability for the sequence pattern discovery. As the existing sequence pattern mining techniques cannot work on such realistic data, we have developed a novel framework for performing sequence pattern mining on probabilistic temporal multimedia event data. We perform probability fusion to resolve the redundancy among detected events from different modalities, considering their cross-modal correlation. We propose a novel sequence pattern mining algorithm called Probabilistic Interval based Event Miner (PIE-Miner) for discovering frequent sequence patterns from interval based events. PIE-Miner has a new support counting mechanism developed for PTM data. Existing sequence pattern mining algorithms have event label level support counting mechanism, whereas we have developed event cluster level support counting mechanism. We discover the complete set of all possible temporal relationships based on Allen's interval algebra. The experimental results showed that the discovered sequence patterns are more useful than the patterns discovered with state-of-the-art sequence pattern mining algorithms.
Chidansh Amitkumar Bhatt, Mohan Kankanhalli
ACM Trans. Intell. Syst. Technol.2
2010 Functionality Delegation in Distributed Surveillance Systems
abstract
The utilization of multimedia devices is growing rapidly in surveillance and monitoring applications. These multimedia surveillance systems need to process large amounts of multimodal sensor data in order to detect events and objects. While processing this large amount of data, the system faces many processing and network bottlenecks. The design of efficient multimedia surveillance system requires intelligent architectural decisions and performance evaluation to cope with these resource demands. One critical issue among all these architectures is task assignment among processing units. To study the effect of this task assignment on system performance with quantifiable performance measures is very useful and challenging. We define a Functionality Delegation Coefficient which abstracts the delegation of functionality among processing units of a distributed surveillance system and show its effect on event blocking probability and response time. Simulation and real implementation results are provided to validate the model.
Mukesh Saini, Pradeep K. Atrey, Sabu Emmanuel, Mohan Kankanhalli
AVSS4
2010 An Authentication Mechanism Using Chinese Remainder Theorem for Efficient Surveillance Video Transmission
abstract
Now-a-days, surveillance cameras have been widely deployed in various security applications. In many surveillance applications, the background changes very slowly and the foreground objects occupy only a relatively small portion of a video frame. In these type of applications, an efficient solution for transmissions over bandwidth-limited networks is to send only the foreground objects for every frame in real time while the background is sent occasionally. At the receiving end of the transmission, the objects and the most recent background can be fused together and the original frame can be reconstructed. However, protecting the authenticity of the video becomes more challenging in this case as a malicious entity can modify/replace/remove the individual foreground objects and background in the video. In this paper, we propose a Chinese remainder theorem based watermarking mechanism for protecting the authenticity of videos transmitted or stored as objects and background. Our mechanism ensures the authenticity between video objects and their associated background.
Tony Thomas, Sabu Emmanuel, Peng Zhang 0005, Mohan Kankanhalli
AVSS4
2010 Efficient Aggregate Licenses Validation in DRM
Amit Sachan, Sabu Emmanuel, Mohan Kankanhalli
DASFAA (2)3
2010 An Eye Fixation Database for Saliency Detection in Images
Subramanian Ramanathan, Harish Katti, Nicu Sebe, Mohan Kankanhalli, Tat-Seng Chua
ECCV (4)4
2010 Privacy modeling for video data publication
abstract
Video cameras are being extensively used in many applications. Huge amounts of video are being recorded and stored everyday by surveillance systems. Any proposed application of this data raises severe privacy concerns. An assessment of privacy loss is necessary before any potential application of the data. In traditional methods of privacy modeling, researchers have focused on explicit means of identity leakage like facial information, etc. However, other implicit inference channels through which individual's an identity can be learned have not been considered. For example, an adversary can observe the behavior, look at the places visited and combine that with the temporal information to infer the identity of the person in the video. In this work, we thoroughly investigate privacy issues involved with the video data considering both implicit and explicit channels. We first establish an analogy with the statistical databases and then propose a model to calculate the privacy loss that might occur due to publication of the video data. The experimental results demonstrate the utility of the proposed model.
Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Sabu Emmanuel, Mohan Kankanhalli
ICME5
2010 Compressed-encrypted domain JPEG2000 image watermarking
abstract
In digital rights management (DRM) systems, digital media is often distributed by multiple levels of distributors in a compressed and encrypted format. The distributors in the chain face the problem of embedding their watermark in compressed, encrypted domain for copyright violation detection purpose. In this paper, we propose a robust watermark embedding technique for JPEG2000 compressed and encrypted images. While the proposed technique embeds watermark in the compressed-encrypted domain, the extraction of watermark can be done either in decrypted domain or in encrypted domain.
A. Venkata Subramanyam, Sabu Emmanuel, Mohan Kankanhalli
ICME3
2010 EMD and psychoacoustic model based watermarking for audio
abstract
The audio watermarking method proposed in this paper offers the copyright protection to an audio without the use of the original signal for watermark detection. The analysis filterbank decomposition, the psychoacoustic model and the empirical mode decomposition (EMD) are the three key techniques used in the novel audio watermarking method. Unlike the traditional audio watermarking algorithms where the watermark bits are embedded directly in the signal either by time domain or transform domain processing, the novel blind audio watermarking algorithm proposed in this paper embeds the watermark bits in the final residue of the subbands in the transform domain. Four watermark messages are embedded into the proposed audio watermarking system. The inaudibility, capacity and robustness of the audio watermarking system are evaluated, in order to optimize the system performance. The experimental results show that the proposed blind watermarking scheme is robust against MP3 compression and adding Gaussian noise attacks.
Liang Wang 0001, Sabu Emmanuel, Mohan Kankanhalli
ICME3
2010 Making computers look the way we look: exploiting visual attention for image understanding
abstract
Human Visual attention (HVA) is an important strategy to focus on specific information while observing and understanding visual stimuli. HVA involves making a series of fixations on select locations while performing tasks such as object recognition, scene understanding, etc. We present one of the first works that combines fixation information with automated concept detectors to (i) infer abstract image semantics, and (ii) enhance performance of object detectors.
Harish Katti, Subramanian Ramanathan, Mohan Kankanhalli, Nicu Sebe, Tat-Seng Chua, K. R. Ramakrishnan
ACM Multimedia3
2010 Modeling, detecting, and processing events in multimedia
abstract
No abstract available.
Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Vasileios Mezaris
ACM Multimedia3
2010 Portfolio theory of multimedia fusion
abstract
The number of multimedia applications has been increasing over the past two decades. Multimedia information fusion has therefore attracted significant attention with many techniques having been proposed. However, the uncertainty and correlation among different modalities have not been fully considered in the existing fusion methods. In general, the predictions of individual modality have uncertainty, furthermore, many modalities are correlated with each other. In this paper, we propose a novel multimedia fusion method based on the Portfolio theory. Portfolio theory is a widely used financial investment theory dealing with how to allocate funds across assets. The key idea is to maximize the performance of the allocated portfolio while minimize the risk in returns. We adapt this approach to multimodal fusion to derive optimal weights that can achieve good fusion results. The optimization is formulated as a quadratic programming problem. Experimental results with both simulated data and real data confirm the theoretical insights and show promising results.
Xiangyu Wang 0002, Mohan Kankanhalli
ACM Multimedia2
2010 Automated aesthetic enhancement of videos
abstract
In this paper, we present a content based single-shot video editing scheme. We follow the classic long take directing and editing schemes. This system automatically adjusts the projection velocity of raw video clips to enhance the aesthetic interest. We build up the mathematical model for projection rhythm manipulation based on film theories. The system segments interesting sub-shots and ordinary sub-shots within the single video clip. Different sub-shots are projected to different duration to maximize the video interest. The output video is rendered according to adjusted projection duration. Within this framework, we transform the screen rhythm and camera motion of a given single video shot. Motion interests of frames are re-distributed in projection duration modification and certain special projection patterns are introduced to enhance the aesthetic interest of original video. The user study shows that our scheme is very effective.
Yang Yang Xiang, Mohan Kankanhalli
ACM Multimedia2
2010 Video retargeting for aesthetic enhancement
abstract
In this paper, we present a post-editing scheme for camera-work. It is based on video retargeting, but aims to enhance the aesthetic interest of home produced video sequences. The essential part of video clips are emphasized by automatically zooming in. The camera pans to preserve the important features within the frame while zooming in. Different from traditional video retargeting schemes, we use a variable zooming factor which is based on the motion saliency of frames.
Yang Yang Xiang, Mohan Kankanhalli
ACM Multimedia2
2010 Multimodal fusion for multimedia analysis: a survey
Pradeep K. Atrey, M. Anwar Hossain 0001, Abdulmotaleb El Saddik, Mohan Kankanhalli
Multim. Syst.4
2010 MultiFusion: A boosting approach for multimedia fusion
abstract
The multimodal data usually contain complementary, correlated and redundant information. Thus, multimodal fusion is useful for many multisensor applications. Here, a novel multimodal fusion algorithm is proposed, which is referred to as “MultiFusion.” The approach adopts a boosting structure where the atomic event is considered as the fusion unit. The correlation of multimodal data is used to form an overall classifier in each iteration. Moreover, by adopting the Adaboost-like structure, the overall fusion performance is improved. Both the simulation experiment and the real application show the effectiveness of the MultiFusion approach. Our approach can be applied in different multimodal applications to exploit the multimedia data characteristics and improve the performance.
Xiangyu Wang 0002, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.2
2009 Context-Based Multimedia Sensor Selection Method
abstract
Modern multimedia systems have large number of sensors spread across a wide area. In a time-shared multimedia system, many people will be making queries to the system simultaneously which requires sharing of computing resources. In such scenarios, processing information from all the sensors for each query will make the system inefficient. Considering the fact that only few sensors provide information relevant to the query, we can reduce the cost incurred in query evaluation by efficiently selecting a subset of sensors to be processed without compromising the system performance. This paper demonstrates a two-stage sensor selection method which uses contextual information and confidence in individual sensors to select sensors which provide more reliable answers to the queries.
Mukesh Saini, Mohan Kankanhalli
AVSS2
2009 A Flexible Surveillance System Architecture
abstract
Traditional multimedia surveillance systems are task specific and tightly coupled to the environment. Moreover, system designs generally start with the assumption that the environment, context, and sensors always remain static. With such a tight coupling, it becomes very difficult to port the system to new environments. Furthermore, for most of the systems, there is no straightforward way to upgrade the existing system to incorporate technological advancements such as new sensors or novel feature extraction techniques. We propose a flexible surveillance system architecture which can be easily ported in different environments, is dynamic without any significant compromise in system performance, and can be extended to integrate newer technological developments. We also introduce the notion of environment model (EM), which completely defines the coupling between system and the physical environment. The isolation of environment specific variables in EM makes the system easily portable in different environments. We present results of a prototype implementation of the system that highlights our design goals.
Mukesh Saini, Mohan Kankanhalli, Ramesh Jain 0001
AVSS2
2009 Privacy Preserving Multiparty Multilevel DRM Architecture
abstract
Traditional digital rights management (DRM) systems are only two party systems, involving the owner and consumers. However, for scalability of business it is often necessary to involve additional levels of distributors and sub-distributors, who can promote and distribute the content in regions unknown to the owner. Thus, we propose an architecture for multiparty multilevel DRM system. The term 'multiparty' refers to involvement of many parties such as the owner, distributors, sub-distributors and consumers and the term 'multilevel' refers to multiple levels of distributors/sub-distributors. The architecture also supports the log files based violation detection, in case of violation of DRM system by any party. However, violation detection imposes a problem of preserving privacy of consumers. So, in the architecture, a provision is made to preserve their privacy.
Amit Sachan, Sabu Emmanuel, Amitabha Das, Mohan Kankanhalli
CCNC4
2009 Efficient license validation in MPML DRM architecture
abstract
Multiparty multilevel DRM architecture (MPML-DRM-A) involves multiple parties such as owner, multiple levels of distributors and consumers. The owner issues redistribution licenses to its distributors, who in turn generate and issue variations of these redistribution licenses to their sub-distributors. Also the distributors generate and issue usage licenses to the consumers to consume the contents. But, these variations of the redistribution licenses and usage licenses generated and issued by each distributor must be validated by a validation authority against the redistribution licenses that it has received. In MPML-DRM-A, there may exist multiple, different types of redistribution licenses for a content. Validation using multiple redistribution licenses may become difficult in real time. Further, storage of multiple redistribution licenses for validation presents a challenge of reducing storage space requirements. Hence, in this paper we propose a bit-vector transform based license organizing structure, and present a method to do the validation of issued licenses in the bit-vector transform domain efficiently. Experimental results show that our license organization structure helps to achieve low validation time and storage space complexity.
Amit Sachan, Sabu Emmanuel, Mohan Kankanhalli
Digital Rights Management Workshop3
2009 Spatiotemporal latent semantic cues for moving people tracking
abstract
Effective and robust visual tracking is one of the most important tasks for the intelligent visual surveillance. In this paper, we proposed a novel method for detecting and tracking moving people using the spatiotemporal latent semantic cues and the incremental eigenspace tracking techniques. During tracking process, the target appearance model is incrementally learned in low dimensional tensor eigenspace by adaptively updating the eigenbasis and sample mean. At the same time, the spatiotemporal latent semantic cues calibrate the estimation of tracking and detect new moving people coming in the same surveillance scene. Experiment results show that with the calibration based on spatiotemporal latent semantic cues, the proposed method can track the moving people automatically and effectively.
Peng Zhang 0005, Sabu Emmanuel, Pradeep K. Atrey, Mohan Kankanhalli
ICASSP4
2009 A robust framework for aligning lecture slides with video
abstract
We propose a robust approach for aligning lecture slides with lecture videos using a combination of Hough transform, optical flow and Gabor analysis. A Markov Decision Process model is used to incorporate prior knowledge for enhanced recognition. We demonstrate synchronization of slides with videos containing de-focused slide content, speaker occlusion as well as camera pan, tilt and zoom sequences. Experimental results confirm the effectiveness of our approach for multimedia indexing applications.
Xiangyu Wang 0002, Subramanian Ramanathan, Mohan Kankanhalli
ICIP3
2009 A CRT based watermark for multiparty multilevel DRM architecture
abstract
In this paper, we propose a joint digital watermarking protocol for the multiparty multilevel DRM architecture using Garner's algorithm for the Chinese remainder theorem (CRT). Our protocol exploits the incremental nature of the computation of CRT by the Garner's algorithm. The proposed joint watermarking protocol embeds a single watermark signal into the content while taking care of the various security concerns such as proof of involvement in the distribution chain, nonrepudiation of the involvement and protection against false framing of the different parties involved. Further, in the event of finding an illegal copy of the content, the identities of all the parties involved in that content distribution chain can be traced back by extracting the watermark information.
Tony Thomas, Sabu Emmanuel, Amitabha Das, Mohan Kankanhalli
ICME4
2009 Performance Modeling of Multimedia Surveillance Systems
abstract
Automated surveillance is critically important in the current scenario of heightened security concerns. Therefore, there has been a surge in the development of surveillance systems. Surveillance systems employ sensors to capture various environmental aspects in order to reason about the dynamically changing situation. Due to their cheap availability, the number of sensors used in modern systems is quite large. While processing the large amount of sensor data, the system faces many bottlenecks in terms of processor and memory requirements. Despite many efforts to propose efficient system architectures, they all ignore the study of the dynamic behavior of the system and the impact of various factors on system performance. In this work, we develop an analytical model to evaluate the performance of surveillance systems. Using the proposed model, we obtain closed form equations for event miss probability and response time as a function of system parameters. The results obtained from the model are validated with those obtained from the simulator. Finally we explore the different trade-offs among the system parameters and performance metrics.
Mukesh Saini, Yashas Natraj, Mohan Kankanhalli
ISM3
2009 Automated localization of affective objects and actions in images via caption text-cum-eye gaze analysis
abstract
We propose a novel framework to localize and label affective objects and actions in images through a combination of text, visual and gaze-based analysis. Human gaze provides useful cues to infer locations and interactions of affective objects. While concepts (labels) associated with an image can be determined from its caption, we demonstrate localization of these concepts upon learning from a statistical affect model for world concepts. The affect model is derived from non-invasively acquired fixation patterns on labeled images, and guides localization of affective objects (faces, reptiles) and actions (look, read) from fixations in unlabeled images. Experimental results obtained on a database of 500 images confirm the effectiveness and promise of the proposed approach.
Subramanian Ramanathan, Harish Katti, Raymond Huang, Tat-Seng Chua, Mohan Kankanhalli
ACM Multimedia5
2009 Events in multimedia
abstract
No abstract available.
Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli
ACM Multimedia3
2009 Secure multimedia content delivery with multiparty multilevel DRM architecture
abstract
For scalability of business, multiparty multilevel digital rights management (DRM) architecture, where a multimedia content is delivered by an owner to a consumer through several levels of distributors has been suggested as an alternative to the traditional two party (buyer-seller) DRM architecture.
Tony Thomas, Sabu Emmanuel, Amitabha Das, Mohan Kankanhalli
NOSSDAV4
2009 On the Security of an MPEG-Video Encryption Scheme Based on Secret Huffman Tables
Shujun Li 0001, Guanrong Chen, Albert Cheung, Kwok-Tung Lo, Mohan Kankanhalli
PSIVT5
2009 Adversary aware surveillance systems
abstract
We consider surveillance problems to be a set of system-adversary interaction problems in which an adversary can be modeled as a rational (selfish) agent trying to maximize his utility. We feel that appropriate adversary modeling can provide deep insights into the system performance and also clues for optimizing the system's performance against the adversary. Further, we propose that system designers should exploit the fact that they can impose certain restrictions on the intruders and the way they interact with the system. The system designers can analyze the scenario to determine conditions under which system outperforms the adversaries, and then suitably reengineer the environment under a "scenario engineering" approach to help the system outperform the adversary. We study the proposed enhancements using a game theoretic framework and present results of their adaptation to two significantly different surveillance scenarios. While the precise enforcements for the studied zero-sum ATM lobby monitoring scenario and the nonzero-sum traffic monitoring scenario were different, they lead to some useful generic guidelines for surveillance system designers.
Vivek K. Singh 0001, Mohan Kankanhalli
IEEE Trans. Inf. Forensics Secur.2
2009 Joint watermarking scheme for multiparty multilevel DRM architecture
abstract
Multiparty multilevel digital rights management (DRM) architecture involving several levels of distributors in between an owner and a consumer has been suggested as an alternative business model to the traditional two-party (buyer-seller) DRM architecture for digital content delivery. In the two-party DRM architecture, cryptographic techniques are used for secure delivery of the content, and watermarking techniques are used for protecting the rights of the seller and the buyer. The cryptographic protocols used in the two-party case for secure content delivery can be directly applied to the multiparty multilevel case. However, the watermarking protocols used in the two-party case may not directly carry over to the multiparty multilevel case, as it needs to address the simultaneous security concerns of multiple parties such as the owner, multiple levels of distributors, and consumers. Towards this, in this paper, we propose a joint digital watermarking scheme using Chinese remainder theorem for the multiparty multilevel DRM architecture. In the proposed scheme, watermark information is jointly created by all the parties involved; then a watermark signal is generated out of it and embedded into the content. This scheme takes care of the security concerns of all parties involved. Further, in the event of finding an illegal copy of the content, the violator(s) can be traced back.
Tony Thomas, Sabu Emmanuel, A. Venkata Subramanyam, Mohan Kankanhalli
IEEE Trans. Inf. Forensics Secur.4
2009 Design of multimedia surveillance systems
abstract
This article addresses the problem of how to select the optimal combination of sensors and how to determine their optimal placement in a surveillance region in order to meet the given performance requirements at a minimal cost for a multimedia surveillance system. We propose to solve this problem by obtaining a performance vector, with its elements representing the performances of subtasks, for a given input combination of sensors and their placement. Then we show that the optimal sensor selection problem can be converted into the form of Integer Linear Programming problem (ILP) by using a linear model for computing the optimal performance vector corresponding to a sensor combination. Optimal performance vector corresponding to a sensor combination refers to the performance vector corresponding to the optimal placement of a sensor combination. To demonstrate the utility of our technique, we design and build a surveillance system consisting of PTZ (Pan-Tilt-Zoom) cameras and active motion sensors for capturing faces. Finally, we show experimentally that optimal placement of sensors based on the design maximizes the system performance.
Garimella S. V. S. Sivaram, Mohan Kankanhalli, K. R. Ramakrishnan
ACM Trans. Multim. Comput. Commun. Appl.2
2008 Pre-attentive discrimination of interestingness in images
abstract
Interestingness is an important aesthetic property, which literally means something that arouses curiosity and is a precursor to attention. Aesthetics is becoming more important as multimedia systems become more human and content centric as opposed to technology centric. In this paper, we use insights from cognitive science, neurophysiology of the early visual system and a mix of human experiments and computational modeling for the purpose of investigating interestingness. Categories in image interestingness and their computational realization are explored through a nontrivial dataset and a real-world problem.
Harish Katti, Kwok Yang Bin, Tat-Seng Chua, Mohan Kankanhalli
ICME4
2008 Quality-aware GSM speech watermarking
abstract
Use of watermarking techniques to provide authentication and tamper proofing of speech in mobile environment is becoming important. However, the current efforts do not allow for user-specifiable quality for the watermarked speech. This paper proposes a watermarking algorithm that allows user-customizable quality for watermarked GSM (Global System for Mobile) speech. Sensitivity (in terms of quality degradation) of each GSM coefficient bits against bit watermark embedding was investigated first, which is then used to select the coefficient bits in a secure manner for watermarking. The proposed algorithm's execution time requirement was studied to draw conclusions on the real-time usability of the algorithm. The embedding capacity and the quality awareness of the algorithm were also investigated.
K. J.-L. Christabel, Sabu Emmanuel, Mohan Kankanhalli
ISCAS3
2008 Multimodal observation systems
abstract
In recent years, we have seen a significant research interest in a number of multimodal sensing applications like surveillance, video ethnography, tele-presence, assisted living, life blogging etc. However, these applications are currently evolving as separate silos with no interconnection. Further, the individual application-centric architectures typically tend to focus on specific sensors, specific (hardwired) queries and deal with specific environments. We present a generic sensing architecture 'Observation System', which allows multiple users to undertake different applications through abstracted interaction with a common set of sensors. The observation system observes behavior of various objects in an environment and keeps a record of important events and activities in an eventbase. In this system, multifarious data collected from disparate sensors and other sources are correlated to understand and gain insights in the environment. The observation system has applications in many areas including but not limited to surveillance, traffic monitoring, ethnography, marketing, and healthcare. In this paper, we present the architecture and functionality of such a system and present details of activity detection using multiple sensor streams in a distributed sensing environment. We also present results of such an approach and potential extensions to the analysis of more complex activities and events.
Mukesh Saini, Vivek K. Singh 0001, Ramesh Jain 0001, Mohan Kankanhalli
ACM Multimedia4
2008 Effectiveness of Signal Segmentation for Music Content Representation
Namunu Chinthaka Maddage, Mohan Kankanhalli, Haizhou Li 0001
MMM2
2008 A cross-modal approach for karaoke artifacts correction
Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.2
2008 Coopetitive multi-camera surveillance using model predictive control
Vivek K. Singh 0001, Pradeep K. Atrey, Mohan Kankanhalli
Mach. Vis. Appl.3
2008 Application Potential of Multimedia Information Retrieval
abstract
This paper will first briefly survey the existing impact of multimedia information retrieval (MIR) in applications. It will then analyze the current trends of MIR research which can have an influence on future applications. It will then detail the future possibilities and bottlenecks in applying the MIR research results in the main target application areas, such as the consumer (e.g., personal video recorders, web information retrieval), public safety (e.g., automated smart surveillance systems), and professional world (e.g., automated meeting capture and summarization). In particular, recommendations will be made to the research community regarding the challenges that need to be met to make the knowledge transfer towards the applications more efficient and effective. It will also attempt to study the trends in the applications which can inform the MIR community on directing intellectual resources towards MIR problems which can have a maximal real-world impact.
Mohan Kankanhalli, Yong Rui
Proc. IEEE1
2008 Progressive Audio Scrambling in Compressed Domain
abstract
Audio scrambling can be employed to ensure confidentiality in audio distribution. We first describe scrambling for raw audio using the discrete wavelet transform (DWT) first and then focus on MP3 audio scrambling. We perform scrambling based on a set of keys which allows for a set of audio outputs having different qualities. During descrambling, the number of keys provided and the number of rounds of descrambling performed will decide the audio output quality. We also perform scrambling by using multiple keys on the MP3 audio format. With a subset of keys, we can descramble to obtain a low quality audio. However, we can obtain the original quality audio by using all of the keys. Our experiments show that the proposed algorithms are effective, fast, simple to implement while providing flexible control over the progressive quality of the audio output. The security level provided by the scheme is sufficient for protecting MP3 music content.
Wei Qi Yan 0001, Wei-Gang Fu, Mohan Kankanhalli
IEEE Trans. Multim.3
2007 A Survey on Digital Camera Image Forensic Methods
abstract
There are two main interests in digital camera image forensics, namely source identification and forgery detection. In this paper, we first briefly provide an introduction to the major processing stages inside a digital camera and then review several methods for source digital camera identification and forgery detection. Existing methods for source identification explore the various processing stages inside a digital camera to derive the clues for distinguishing the source cameras while forgery detection checks for inconsistencies in image quality or for presence of certain characteristics as evidence of tampering.
Tran Van Lanh, Kai-Sen Chong, Sabu Emmanuel, Mohan Kankanhalli
ICME4
2007 Towards Adversary Aware Surveillance Systems
abstract
We consider surveillance problems to be a set of system-adversary interaction problems in which an adversary can be modeled as a rational (selfish) agent trying to maximize his utility. We feel that appropriate adversary modeling can provide deep insights into the system performance and also clues for optimizing the system's performance against the adversary. Further, we propose that system designers should exploit the fact that they can impose certain restrictions on the intruders and the way they interact with the system. The system designers can find the assumptions under which the surveillance system shall out-perform the intruder and then enforce those assumptions over the system-intruder interaction as part of a 'scenario engineering' approach. We study both these aspects using a game theoretic framework and undertake practical experiments to verify the proposed enhancements.
Vivek K. Singh 0001, Mohan Kankanhalli
ICME2
2007 Identifying Source Cell Phone using Chromatic Aberration
abstract
Chromatic aberration is the phenomenon where light of different wavelengths fail to converge at the same position on the focal plane. There are two kinds of chromatic aberration: longitudinal aberration causes different wavelengths to focus at different distances from the lens while lateral aberration is attributed to different wavelengths focusing at different positions on the sensor. In this paper, we estimate the parameters of lateral chromatic aberration by maximizing the mutual information between the corrected R and B channels with the G channel. The extracted parameters are then used as input features to a SVM classifier for identifying source cell phone of images. By considering only a certain part of the image when estimating the parameters, we reduce the runtime complexity of the algorithm dramatically while preserving the accuracy at a high level.
Tran Van Lanh, Sabu Emmanuel, Mohan Kankanhalli
ICME3
2007 Confidence Building Among Correlated Streams in Multimedia Surveillance Systems
Pradeep K. Atrey, Mohan Kankanhalli, Abdulmotaleb El Saddik
MMM (2)2
2007 Metadata Management, Reuse, Inference and Propagation in a Collection-Oriented Metadata Framework for Digital Images
William Ku, Mohan Kankanhalli, Joo-Hwee Lim
MMM (2)2
2007 Coopetitive Multimedia Surveillance
Vivek K. Singh 0001, Pradeep K. Atrey, Mohan Kankanhalli
MMM (2)3
2007 A scalable signature scheme for video authentication
Pradeep K. Atrey, Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.3
2007 Render Sequence Encoding for Document Protection
abstract
We present in this paper a novel electronic document watermarking method, render sequence encoding (RSE), and then further develop a RSE authentication method for electronic documents. RSE watermarks an electronic document by modulating the display sequences of words or characters. It features large information-carrying capacity and robustness over document format transcoding. The RSE authentication method is based on the NP-complete exact traveling salesman problem, which provides a rigorous foundation for security. The RSE authentication method is secure in the sense it is extremely difficult to forge the authentication process. RSE authentication process is also easy to operate, especially in comparison to digital signatures which requires public key infrastructure for its operation
Baoshi Zhu, Jian-Kang Wu, Mohan Kankanhalli
IEEE Trans. Multim.3
2007 Goal-oriented optimal subset selection of correlated multimedia streams
abstract
A multimedia analysis system utilizes a set of correlated media streams, each of which, we assume, has a confidence level and a cost associated with it, and each of which partially helps in achieving the system goal. However, the fact that at any instant, not all of the media streams contribute towards a system goal brings up the issue of finding the best subset from the available set of media streams. For example, a subset of two video cameras and two microphones could be better than any other subset of sensors at some time instance to achieve a surveillance goal (e.g. event detection). This article presents a novel framework that finds the optimal subset of media streams so as to achieve the system goal under specified constraints. The proposed framework uses a dynamic programming approach to find the optimal subset of media streams based on three different criteria: first, by maximizing the probability of achieving the goal under the specified cost and confidence; second, by maximizing the confidence in the achieved goal under the specified cost and probability with which the goal is achieved; and third, by minimizing the cost to achieve the goal with a specified probability and confidence. Each of these problems is proven to be NP-Complete. From an AI point of view, the solution we propose is heuristic-based, and for each criterion, utilizes a heuristic function which for a given problem, combines optimal solutions of small-sized subproblems to yield a potential near-optimal solution to the original problem. The proposed framework allows for a tradeoff among the aforementioned three criteria, and offers the flexibility to compare whether any one set of media streams of low cost would be better than any other set of higher cost, or whether any one set of media streams of high confidence would be better than any other set of low confidence. To show the utility of our framework, we provide the experimental results for event detection in a surveillance scenario.
Pradeep K. Atrey, Mohan Kankanhalli, B. John Oommen
ACM Trans. Multim. Comput. Commun. Appl.2
2007 Multimedia simplification for optimized MMS synthesis
abstract
We propose a novel transcoding technique called multimedia simplification which is based on experiential sampling. Multimedia simplification helps optimize the synthesis of MMS (multimedia messaging service) messages for mobile phones. Transcoding is useful in overcoming the limitations of these compact devices. The proposed approach aims at reducing the redundancy in the multimedia data captured by multiple types of media sensors. The simplified data is first stored into a gallery for further usage. Once a request for MMS is received, the MMS server makes use of the simplified media from the gallery. The multimedia data is aligned with respect to the timeline for MMS message synthesis. We demonstrate the use of the proposed techniques for two applications, namely, soccer video and home care monitoring video. The MMS sent to the receiver can basically reflect the gist of important events of interest to the user. Our technique is targeted towards users who are interested in obtaining salient multimedia information via mobile devices.
Wei Qi Yan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.2
2006 Multimedia Surveillance and Monitoring
abstract
In spite of the development of various media sensors, multimedia (and computer vision) researchers have mostly adopted a video-centric approach to solve the automated surveillance and monitoring related problems. We look at the monitoring/surveillance problem from an information-centric perspective and advocate the use of diverse sources of information which enables the use of multiple correlated media. We advocate a design methodology for building systems which can explicitly take performance into account. We then propose that the surveillance problem can be better posed as an "information-search" problem in which the user can query for the information of his/her interest. We will present a framework for multimedia monitoring that uses a domain-data transformation model based approach to map domain-events to their equivalent data-events. We will motivate the new approach, present the architecture and highlight the information assimilation aspects. We will also present some open problems and issues arising from the novel way of looking at monitoring.
Mohan Kankanhalli
AVSS1
2006 Audio Based Event Detection for Multimedia Surveillance
abstract
With the increasing use of audio sensors in surveillance and monitoring applications, event detection using audio streams has emerged as an important research problem. This paper presents a hierarchical approach for audio based event detection for surveillance. The proposed approach first classifies a given audio frame into vocal and nonvocal events, and then performs further classification into normal and excited events. We model the events using a Gaussian mixture model and optimize the parameters for four different audio features ZCR, LPC, LPCC and LFCC. Experiments have been performed to evaluate the effectiveness of the features for detecting various normal and the excited state human activities. The results show that the proposed top-down event detection approach works significantly better than the single level approach
Pradeep K. Atrey, Namunu Chinthaka Maddage, Mohan Kankanhalli
ICASSP (5)3
2006 Experiential Sampling based Foreground/Background Segmentation for Video Surveillance
abstract
Segmentation of foreground and background has been an important research problem arising out of many applications including video surveillance. A method commonly used for segmentation is "background subtraction" or thresholding the difference between the estimated background image and current image. Adaptive Gaussian mixture based background modelling has been proposed by many researchers for increasing the robustness against environmental changes. However, all these methods, being computationally intensive, need to be optimized for efficient and real-time performance especially at a higher image resolution. In this paper, we propose an improved foreground/background segmentation method which uses experiential sampling technique to restrict the computational efforts in the region of interest. We exploit the fact that the region of interest in general is present only in a small part of the image, therefore, the attention should only be focused in those regions. The proposed method shows a significant gain in processing speed at the expense of minor loss in accuracy. We provide experimental results and detailed analysis to show the utility of our method
Pradeep K. Atrey, Anurag Kumar 0001, Mohan Kankanhalli
ICME4
2006 A Collection-Oriented Metadata Framework for Digital Images
abstract
A digital photo can "tell a thousand words" through the use of its metadata and as it is usually part of a collection, metadata management, reuse, propagation&inference could be achieved via its association with a collection. However, there is not much work on metadata management, reuse, propagation&inference, particularly on a group basis. In this paper, we proposed a collection-oriented metadata framework which provides a basis for metadata management, reuse, propagation&inference and demonstrated the utility of such a framework
William Ku, Mohan Kankanhalli, Joo-Hwee Lim
ICME2
2006 A Hierarchical Approach for Music Chord Modeling Based on the Analysis of Tonal Characteristics
abstract
This paper first discusses how the signal segmentation and tonal characteristics of music notes effect in music chord detection. Two approaches, pitch class profile approach and psycho-acoustical approach, which differently represent these tonal characteristics, are examined for chord detection. The analysis of the tonal characteristics reveals that not only the fundamental frequency of music note but also its harmonics and sub-harmonies in different octaves contribute for detecting related music chord. A hierarchical approach, which transforms the music chord tonal characteristics in each octave onto probabilistic space, is then proposed for modeling the music chord. Our experimental results show that detection of chord type, major, minor, diminish, and augmented, and individual chords, 12 chords per chord type, are improved with the proposed hierarchical chord modeling approach. Experimental results also reveal that the tempo proportional signal segmentation is more effective extracting tonal characteristics than using fixed length segmentation
Namunu Chinthaka Maddage, Mohan Kankanhalli, Haizhou Li 0001
ICME2
2006 Predominant Vocal Pitch Detection in Polyphonic Music
abstract
We present a novel method for predominant vocal pitch detection in two-channel polyphonic music. The proposed method contains two stages. In the first stage, we apply the frequency domain independent component Analysis (FD-ICA) for the two-channel polyphonic music to separate the vocal content from the background music. Considering the vocal singing voice and background music are two heterogeneous signals, we employ a statistical learning based method to solve the permutation inconsistency problem in FD-ICA. In the second stage, a noise insensitive vocal pitch detection method is proposed, which is robust to noise and errors introduced by the separation process in the first stage. The proposed method has been tested on the two-channel polyphonic music signals, and experimental results show promising performance
Xi Shao, Changsheng Xu, Mohan Kankanhalli
ICME3
2006 An Anonymous Routing Protocol with The Local-repair Mechanism for Mobile Ad Hoc Networks
abstract
In this paper, we first define the requirements on anonymity and security properties of the routing protocol in mobile ad hoc networks, and then propose a new anonymous routing protocol with the local-repair mechanism. Detailed analysis shows that our protocol achieves both anonymity and security properties defined. A major challenge in designing anonymous routing protocols is to reduce computation and communication costs. To overcome this challenge, our protocol is design to require neither asymmetric nor symmetric encryption/decryption while updating the flooding route requests; more importantly, once a route is broken, instead of re-launching a new costly flooding route discovery process like previous work, our protocol provides a local-repair mechanism to fix broken parts of a route without compromising anonymity
Bo Zhu 0001, Sushil Jajodia, Mohan Kankanhalli, Feng Bao 0001, Robert H. Deng
SECON3
2006 Music structure based vector space retrieval
abstract
This paper proposes a novel framework for music content indexing and retrieval. The music structure information, i.e., timing, harmony and music region content, is represented by the layers of the music structure pyramid. We begin by extracting this layered structure information. We analyze the rhythm of the music and then segment the signal proportional to the inter-beat intervals. Thus, the timing information is incorporated in the segmentation process, which we call Beat Space Segmentation. To describe Harmony Events, we propose a two-layer hierarchical approach to model the music chords. We also model the progression of instrumental and vocal content as Acoustic Events. After information extraction, we propose a vector space modeling approach which uses these events as the indexing terms. In query-by-example music retrieval, a query is represented by a vector of the statistics of the n-gram events. We then propose two effective retrieval models, a hard-indexing scheme and a soft-indexing scheme. Experiments show that the vector space modeling is effective in representing the layered music information, achieving 82.5% top-5 retrieval accuracy using 15-sec music clips as the queries. The soft-indexing outperforms hard-indexing in general.
Namunu Chinthaka Maddage, Haizhou Li 0001, Mohan Kankanhalli
SIGIR3
2006 Information assimilation framework for event detection in multimedia surveillance systems
Pradeep K. Atrey, Mohan Kankanhalli, Ramesh Jain 0001
Multim. Syst.2
2006 Mask-based fingerprinting scheme for digital video broadcasting
Sabu Emmanuel, Mohan Kankanhalli
Multim. Tools Appl.2
2006 Experiential Sampling in Multimedia Systems
abstract
Multimedia systems must deal with multiple data streams. Each data stream usually contains significant volume of redundant noisy data. In many real-time applications, it is essential to focus the computing resources on a relevant subset of data streams at any given time instant and use it to build the model of the environment. We formulate this problem as an experiential sampling problem and propose an approach to utilize computing resources efficiently on the most informative subset of data streams. First, in this paper, we focus on theoretical background and develop a theoretical framework for a single data stream. We generalize the notion of static visual attention in a dynamical systems setting and propose a dynamical attention-orientated analysis method. This is achieved by a sampling representation that utilizes the current context and past experience for attention evolution. Hence, the multimedia analysis task at hand can select its data of interest while immediately discarding the irrelevant data to achieve efficiency and adaptability.
Mohan Kankanhalli, Jun Wang 0012, Ramesh Jain 0001
IEEE Trans. Multim.1
2006 Experiential Sampling on Multiple Data Streams
abstract
Multimedia systems must deal with multiple data streams. Each data stream usually contains significant volume of redundant noisy data. In many real-time applications, it is essential to focus the computing resources on a relevant subset of data streams at any given time instant and use it to build the model of the environment. We formulate this problem as an experiential sampling problem and propose an approach to utilize computing resources efficiently on the most informative subset of data streams. In this paper, we generalize our experiential sampling framework to multiple data streams and provide an evaluation measure for this technique. We have successfully applied this framework to the problems of traffic monitoring, face detection and monologue detection.
Mohan Kankanhalli, Jun Wang 0012, Ramesh Jain 0001
IEEE Trans. Multim.1
2006 Precise pitch profile feature extraction from musical audio for key detection
abstract
The majority of pieces of music, including classical and popular music,are composed using music scales, such as keys. The key or the scale information of a piece provides important clues on its high level musical content, like harmonic and melodic context. Automatic key detection from music data can be useful for music classification, retrieval or further content analysis. Many researchers have addressed key finding from symbolically encoded music(MIDI); however, works for key detection in musical audio is still limited. Techniques for key detection from musical audio mainly consist of two steps:pitch extraction and key detection. The pitch feature typically characterizes the weights of presence of particular pitch classes in the music audio. In the existing approaches to pitch extraction, little consideration has been taken on pitch mistuning and interference of noisy percussion sounds in the audio signals, which inevitably affects the accuracy of key detection. In this paper, we present a novel technique of precise pitch profile feature extraction, which deals with pitch mistuning and noisy percussive sounds. The extracted pitch profile feature can characterize the pitch content in the signal more accurately than the previous techniques, thus lead to a higher key detection accuracy. Experiments based on classical and popular music data were conducted. The results showed that the proposed method has higher key detection accuracy than previous methods, especially for popular music with a lot of noisy drum sounds.
Yongwei Zhu, Mohan Kankanhalli
IEEE Trans. Multim.2
2006 Metadata handling: A video perspective
abstract
This article addresses the problem of processing the annotations of preexisting video productions to enable reuse and repurposing of metadata. We introduce the concept of automatic content-based editing of preexisting semantic home video metadata. We propose a formal representation and implementation techniques for reusing and repurposing semantic video metadata in concordance with the actual video editing operations. A novel representation for metadata editing is proposed and an implementation framework for editing the metadata in accordance with the video editing operations is demonstrated. Conflict resolution and regularization operations are defined and implemented in the context of the video metadata editing operations.
Chitra L. Madhwacharyula, Marc Davis, Philippe Mulhem, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.4
2006 Automatic summarization of music videos
abstract
In this article, we propose a novel approach for automatic music video summarization. The proposed summarization scheme is different from the current methods used for video summarization. The music video is separated into the music track and video track. For the music track, a music summary is created by analyzing the music content using music features, an adaptive clustering algorithm, and music domain knowledge. Then, shots in the video track are detected and clustered. Finally, the music video summary is created by aligning the music summary and clustered video shots. Subjective studies by experienced users have been conducted to evaluate the quality of music summaries and effectiveness of the proposed summarization approach. Experiments are performed on different genres of music videos and comparisons are made with the summaries generated based on music track, video track, and manually. The evaluation results indicate that summaries generated using the proposed method are effective in helping realize users' expectations.
Xi Shao, Changsheng Xu, Namunu Chinthaka Maddage, Qi Tian 0002, Mohan Kankanhalli, Jesse S. Jin
ACM Trans. Multim. Comput. Commun. Appl.5
2005 Automatic music summarization based on music structure analysis
abstract
In this paper, we present a novel approach for music summarization based on music structure analysis. From the audio signal, we first extract the note onset representing the time tempo of the song and the music structure analysis can be performed based on this tempo information. After music content has been structured into different semantic regions such as introduction (intro), verse, chorus, ending (outro), etc., the final music summary can be created with chorus and music phrases which are included anterior or posterior to the selected chorus to get the desired length of the final summary. In this way, we can guarantee that the summaries begin and end at meaningful music phrase boundaries, which is a difficult problem for existing music summarization methods. Experiments show our proposed method can capture the main theme of the music compared to the ideal summaries selected by music experts and user subjective evaluation indicates our proposed method has a good performance.
Xi Shao, Namunu Chinthaka Maddage, Changsheng Xu, Mohan Kankanhalli
ICASSP (2)4
2005 Goal based optimal selection of media streams
abstract
A multimedia system utilizes a set of correlated media streams each of which partially help in achieving the system goal. However, since not all of the streams always contribute towards the goal, there is a need for determining the most informative subset from the available set of media streams at any instant. For example, a subset of two video cameras and two microphones could be better than any other subset of multimedia sensors at some time instance. This paper presents a novel framework to find the optimal subset of media streams that achieves the system goal under specified constraints. The proposed framework uses a dynamic programming approach to find the optimal subset of media streams based on two criteria; first, by maximizing the probability of achieving the goal under the specified maximum cost, and second by minimizing the cost of using the streams so that the goal is achieved with a specified minimum probability. To show the utility of our framework, we provide the simulation results for hypothesis testing.
Pradeep K. Atrey, Mohan Kankanhalli
ICME2
2005 Providing efficient certification services against active attacks in ad hoc networks
abstract
Most of previous research work in key management can only resist passive attacks, such as dropping the certificate request, and are vulnerable under active attacks, such as returning a fake reply to the node requesting the certification service. In this paper, we propose two algorithms to address both security and efficiency issues of certification services in ad hoc networks. Both of the algorithms can resist active attacks. In addition, simulation results show that, compared to the previous works, our second algorithm is not only much faster in a friendly environment, but it also works well in a hostile environment in which existing schemes work poorly. Furthermore, the process of generating partial certificates in our second algorithm is extremely fast. Such advantage is critical in ad hoc networks where by nature the less help a node requests from its neighbors, the higher is the chance of obtaining the help. Consequently, using our second algorithm, a node can easily find enough neighboring nodes which provide the certification service.
Bo Zhu 0001, Guilin Wang, Zhiguo Wan, Mohan Kankanhalli, Feng Bao 0001, Robert H. Deng
IPCCC4
2005 What is the state of our community?
abstract
10.1145/1101149.1101297
Yong Rui, Ramesh Jain 0001, Nicolas D. Georganas, HongJiang Zhang, Klara Nahrstedt, John R. Smith, Mohan Kankanhalli
ACM Multimedia7
2005 Music Key Detection for Musical Audio
abstract
The key or the scale information of a piece of music provides important clues on its high level musical content, like harmonic and melodic context, which can be useful for music classification, retrieval or further content analysis. Researchers have previously addressed the issue of finding the key for symbolically encoded music (MIDI); however, very little work has been done on key detection for acoustic music. In this paper, we present a method for estimating the root of diatonic scale and the key directly from acoustic signals (waveform) of popular and classical music. We propose a method to extract pitch profile features from the audio signal, which characterizes the tone distribution in the music. The diatonic scale root and key are estimated based on the extracted pitch profile by using a tone clustering algorithm and utilizing the tone structure of keys. Experiments on 72 music pieces have been conducted to evaluate the proposed techniques. The success rate of scale root detection for pop music pieces is above 90%.
Yongwei Zhu, Mohan Kankanhalli
MMM2
2005 Automatic music video summarization based on audio-visual-text analysis and alignment
abstract
In this paper, we propose a novel approach for automatic music video summarization based on audio-visual-text analysis and alignment. The music video is separated into the music and video tracks. For the music track, the chorus is detected based on music structure analysis. For the video track, we first segment the shots and classify the shots into close-up face shots and non-face shots, then we extract the lyrics and detect the most repeated lyrics from the shots. The music video summary is generated based on the alignment of boundaries of the detected chorus, shot class and the most repeated lyrics from the music video. The experiments on chorus detection, shot classification, and lyrics detection using 20 English music videos are described. Subjective user studies have been conducted to evaluate the quality and effectiveness of summary. The comparisons with the summaries based on our previous method and the manual method indicate that the results of summarization using the proposed method are better at meeting users' expectations.
Changsheng Xu, Xi Shao, Namunu Chinthaka Maddage, Mohan Kankanhalli
SIGIR4
2005 Efficient and robust key management for large mobile ad hoc networks
Bo Zhu 0001, Feng Bao 0001, Robert H. Deng, Mohan Kankanhalli, Guilin Wang
Comput. Networks4
2005 Analogies based video editing
Wei Qi Yan 0001, Mohan Kankanhalli, Jun Wang 0012
Multim. Syst.2
2005 Automatic video logo detection and removal
Wei Qi Yan 0001, Jun Wang 0012, Mohan Kankanhalli
Multim. Syst.3
2004 Automatic music summarization in compressed domain
abstract
A novel compressed domain automatic music summarization approach is presented in this paper. The proposed method works directly in the compressed domain. Only the encoded subband samples are extracted and processed for characterizing music content and discovering the music structure. The experimental results and the evaluation by a subjective study have shown that the summarization based on MPEG-1 Layer 3 (MP3) music is comparable to the summarization based on uncompressed PCM music samples.
Xi Shao, Changsheng Xu, Ye Wang 0007, Mohan Kankanhalli
ICASSP (4)4
2004 Harmonicity and dynamics-based features for audio
abstract
Features are very important for audio processing. Tasks like speech recognition and instrument identification are based on features. Most low-level features currently used are based on LPC and cepstral analysis. We propose a class of features based on dynamics and harmonicity. In particular, we define the notion of harmonic derivative. The efficacy of the features is demonstrated for music genre classification and instrument family classification. In particular, the features are shown to be cepstrum-equivalent.
S. H. Srinivasan, Mohan Kankanhalli
ICASSP (4)2
2004 Goal detection in soccer video using audio/visual keywords
Yu-Lin Kang, Joo-Hwee Lim, Mohan Kankanhalli, Changsheng Xu, Qi Tian 0002
ICIP3
2004 A new approch to automatic music video summarization
abstract
A new automatic summarization approach for music videos is presented. The proposed method detects and recognizes lyric captions appearing commonly in karaoke music videos and uses the captions to analyze music video structure and identify the most salient music part. The music video summary is created based on the salient part. Experimental results show our proposed method is promising.
Xi Shao, Changsheng Xu, Mohan Kankanhalli
ICIP3
2004 Content based editing of semantic video metadata
abstract
Bridging the 'signal-symbol gap' existing between multimedia signals generated through audio, video or other multimedia streams and the high level symbols (metadata) which describe them is presently one of the most vital areas of multimedia research. The paper attempts to bridge this significant gap by proposing a novel automatic mechanism for XML based video metadata editing, in tandem with video editing operations. An implementation framework for editing metadata in accordance with the video editing operations is demonstrated. Conflicting resolution and regularization operations are defined and implemented with respect to video metadata editing operations.
Chitra L. Madhwacharyula, Mohan Kankanhalli, Philippe Mulhem
ICME2
2004 Unsupervised classification of music genre using hidden Markov model
abstract
Music genre classification can be of great utility to musical database management. Most current classification methods are supervised and tend to be based on contrived taxonomies. However, due to the ambiguities and inconsistencies in the chosen taxonomies, these methods are not applicable for a much larger database. We proposed an unsupervised clustering method, based on a given measure of similarity which can be provided by hidden Markov models. In addition, in order to better characterize music content, a novel segmentation scheme is proposed, based on music intrinsic rhythmic structure analysis and features are extracted based on these segments. The performance of this feature segmentation scheme performs better than the traditional fixed-length method, according to experimental results. Our preliminary results also suggest that the proposed method is comparable to the supervised classification method.
Xi Shao, Changsheng Xu, Mohan Kankanhalli
ICME3
2004 Mosaic based view enlargement for moving objects in moving pictures
abstract
Conventional mosaicing techniques convert a video from frame-based representation to scene-based representation, but they usually lack dynamic information so that their mosaic is not complete. In this paper, we present a novel method to detect moving objects in the video sequences, then add them into the static background mosaic to represent the scene completely. This novel algorithm separates static and dynamic information in a video sequence, builds the background mosaic from static part and reconstructs moving objects on the static mosaic. We have implemented our techniques and the experimental results demonstrate the effectiveness of our approach
Mohan Kankanhalli, S. H. Srinivasan, Wei Qi Yan 0001
ICME2
2004 Video content representation on tiny devices
abstract
The perceptual satisfaction of a user watching video on a tiny mobile device is constrained by the display capability and network bandwidth. To maximize the user's perceptual satisfaction in this constrained environment, we propose a new method to represent the video content adaptively in real-time on tiny devices according to the user's attention. First, a sampling based dynamic attention model is proposed to obtain and maintain the user's attention in the video streams. Second, based on the most attended regions and sequences extracted, the attention based representation is introduced to achieve a higher perceptual satisfaction on a small display. Experiments with users show the effectiveness of our proposed method in a video surveillance application.
Jun Wang 0012, Marcel J. T. Reinders, Reginald L. Lagendijk, Jasper Lindenberg, Mohan Kankanhalli
ICME5
2004 Automatically summarize musical audio using adaptive clustering
abstract
Automatic music summarization is very useful for music indexing, content-based music retrieval and on-line music distribution, but it is a challenge to extract automatically the most common and salient themes from unstructured raw music data. We propose an effective approach to summarize music content automatically. First, a number of features are extracted to characterize the music content. Based on the extracted features, an adaptive clustering algorithm is then applied to structure the music content. Finally, the music summary is created in terms of the clustering results and domain-related music knowledge. A user study is conducted to evaluate the quality of summarization. The experiments on different genres of music illustrate the results of summarization are significant and effective to actual expectation.
Changsheng Xu, Xi Shao, Namunu Chinthaka Maddage, Mohan Kankanhalli, Qi Tian 0002
ICME4
2004 A method for solmization of melody
abstract
This work presents a novel method for the automatic solmization of a melody, by which a melody (a sequence of MIDI notes) can be transcribed to sol-fa syllables (i.e., do, re, me, fa, sol, la, ti). Automatic solmization can assist in music skill training, music notation and content-based music retrieval. The proposed method is based on an approach for estimating the music scale of a melody. The key of the major scale ("do") is estimated using music scale models. Due to the diversity of melody types, models for both diatonic and pentatonic scales are employed to avoid the possible key ambiguity for folk songs. The decision of the key of a melody is based on the scale estimation results for aggregating music notes, so that the method can work for both short and long melodies. Experiments have shown that the technique can achieve 95% correct solmization of the melodies of pop songs.
Yongwei Zhu, Mohan Kankanhalli
ICME2
2004 Anonymous Secure Routing in Mobile Ad-Hoc Networks
abstract
Although there are a large number of papers on secure routing in mobile ad-hoc networks, only a few consider the anonymity issue. We define more strict requirements on the anonymity and security properties of the routing protocol, and notice that previous research works only provide weak location privacy and route anonymity, and are vulnerable to specific attacks. Therefore, we propose the anonymous secure routing (ASR) protocol that can provide additional properties on anonymity, i.e. identity anonymity and strong location privacy, and at the same time ensure the security of discovered routes against various passive and active attacks. Detailed analysis shows that ASR can achieve both anonymity and security properties, as defined in the requirements, of the routing protocol in mobile ad-hoc networks.
Bo Zhu 0001, Zhiguo Wan, Mohan Kankanhalli, Feng Bao 0001, Robert H. Deng
LCN3
2004 Probability fusion for correlated multimedia streams
abstract
The fusion of multiple correlated observations of a multimedia system is a research problem arising in many multimedia applications. In this paper, we propose a novel framework for the probabilistic fusion of correlated multimedia observations. Assuming that each of the media stream has a priori probability of achieving the goal and their underlying correlations are known, our framework fuses the individual probabilities using the quantitative correlation based on a Bayesian approach. The simulation results show that fewer highly-positively-correlated observations better achieve a specified goal when compared to the use of a larger number of observations with low correlation.
Pradeep K. Atrey, Mohan Kankanhalli
ACM Multimedia2
2004 Content-based music structure analysis with applications to music semantics understanding
abstract
In this paper, we present a novel approach for music structure analysis. A new segmentation method, beat space segmentation, is proposed and used for music chord detection and vocal/instrumental boundary detection. The wrongly detected chords in the chord pattern sequence and the misclassified vocal/instrumental frames are corrected using heuristics derived from the domain knowledge of music composition. Melody-based similarity regions are detected by matching sub-chord patterns using dynamic programming. The vocal content of the melody-based similarity regions is further analyzed to detect the content-based similarity regions. Based on melody-based and content-based similarity regions, the music structure is identified. Experimental results are encouraging and indicate that the performance of the proposed approach is superior to that of the existing methods. We believe that music structure analysis can greatly help music semantics understanding which can aid music transcription, summarization, retrieval and streaming.
Namunu Chinthaka Maddage, Changsheng Xu, Mohan Kankanhalli, Xi Shao
ACM Multimedia3
2004 A Hierarchical Signature Scheme for Robust Video Authentication using Secret Sharing
abstract
Ensuring the integrity of a digital video is an important and challenging research problem arising out of many video applications. In this paper, we present a hierarchical framework for video authentication based on cryptographic secret sharing that protects a video from spatial cropping and temporal jittering, yet is robust against frame dropping in the streaming video scenario. Our algorithm provides a tradeoff between security and robustness by having configurable inputs. The authentication signature is compact and very sensitive against spatial attacks such as region tampering, and interframe attacks like frame replacement, major frame dropping, and frame reordering. Given a video, we identify the key frames based on different energy between the frames. Considering video frames as shares, we compute the secret at three hierarchical levels. The master secret is used as digital signature to authenticate the video. We present extensive experimental results which show the utility of our technique.
Pradeep K. Atrey, Wei Qi Yan 0001, Ee-Chien Chang, Mohan Kankanhalli
MMM4
2003 Print signatures for document authentication
abstract
We present a novel solution for authenticating printed paper documents by utilizing the inherent non--repeatable randomness existing in the printing process. For a document printed by a laser-printer, we extract the unique features of the non--repeatable print content for each copy. The shape profiles of this content are used as the feature to represent the uniqueness of that particular printed copy. These features along with some important document content is then captured as the print signature. We present theoretical and experimental details on how to register as well as authenticate this print signature. The security analysis of this technique is also presented. We finally provide experimental results to demonstrate the feasibility of the proposed method.
Baoshi Zhu, Jian-Kang Wu, Mohan Kankanhalli
CCS3
2003 Harmonicity and dynamics based audio separation
abstract
Audio signal source separation is an interesting task performed by humans. In this paper, we present a frequency grouping algorithm based on principles of harmonicity and dynamics: frequency components with a harmonic relation and similar dynamics belong to the same source. The grouping is demonstrated for a variety of sound mixtures.
S. H. Srinivasan, Mohan Kankanhalli
ICASSP (5)2
2003 A hierarchical framework for face tracking using state vector fusion for compressed video
abstract
Faces usually are the most interesting objects in certain categories of video, like home videos and news clips. A novel sensor fusion based face tracking system is presented that tracks faces in compressed video, and aids automatic video indexing. Tracking is done by fusing the measurements from three independent sensors - motion and colour based trackers (Achanta, R. et al., IEEE Int. Conf. on Multimedia and Expo, 2002) and a face detector (Wang, J. et al., Proc. Int. Workshop on Advanced Image Technology, 2002) using a novel hierarchical framework based on Kalman filter state vector fusion. The tracking results show that the fused results are better than those of any individual sensors or their mean.
Jun Wang 0012, Radhakrishna S. V. Achanta, Mohan Kankanhalli, Philippe Mulhem
ICASSP (3)3
2003 Automatically generating summaries for musical video
abstract
In this paper, we propose a novel approach to automatically summarize musical videos. The proposed summarization scheme is different from the current methods used for video summarization. The musical video is separated into the musical and visual tracks. A music summary is created by analyzing the music content based on music features, adaptive clustering algorithm and musical domain knowledge. Then, shots are detected and clustered in the visual track. Finally, the music video summary is created by aligning the music summary and clustered video shots. Subjective studies by experienced users have been conducted to evaluate the quality of summarization. The experiments on different genres of musical video and comparisons with the summaries only based on music track and video track indicate that the results of summarization using proposed method are significant and effective to help realize user's expectation.
Xi Shao, Changsheng Xu, Mohan Kankanhalli
ICIP (2)3
2003 Fractional scaling of image and video in DCT domain
abstract
An algorithm for scaling image and video with fractional factors of 1.25 and 1.50 directly in compressed (DCT) domain without explicit decompression and recompression is presented. It differs from those compressed-domain sampling methods that work only with integer factors, and provides more flexibility in changing image sizes. We employ a simple, consistent and extensible way to implement the algorithm. The resulting scheme ensures that the compressed domain algorithms always use fewer arithmetic operations than their conventional spatial domain counterparts. Experimental results in terms of visual quality and objective evaluation are provided.
Mohan Kankanhalli, Tat-Seng Chua
ICIP (1)2
2003 Wide baseline spectral matching
abstract
Wide baseline matching of images is an important problem in multimedia and computer vision. Recently a class of spectral methods have been proposed for tasks like segmentation. In this paper we propose a spectral algorithm for wide baseline matching.
S. H. Srinivasan, Mohan Kankanhalli
ICME2
2003 Creating audio keywords for event detection in soccer video
abstract
This paper presents a novel framework called audio keywords to assist event detection in soccer video. Audio keyword is a middle-level representation that can bridge the gap between low-level features and high-level semantics. Audio keywords are created from low-level audio features by using support vector machine learning. The created audio keywords can be used to detect semantic events in soccer video by applying a heuristic mapping. Experiments of audio keywords creation and event detection based on audio keywords have illustrated promising results. According to the experimental results, we believe that audio keyword is an effective representation that is able to achieve more intuitionistic result for event detection in sports video compared with the method of event detection directly based on low-level features.
Min Xu 0001, Namunu Chinthaka Maddage, Changsheng Xu, Mohan Kankanhalli, Qi Tian 0002
ICME4
2003 Colorizing infrared home videos
abstract
A color video always conveys more vivid sentiments than a grayscale one. Obtaining a grayscale video from a color video is almost trivial but the converse is known to be hard. Nowadays, digital camcorders come equipped with an infrared device for night shot that enables one to shoot home videos in the dark. Unfortunately, the infrared lighting device used generates a "green-scale" video which is akin to a grayscale video albeit possessing all tints of green. In this paper, we present a novel technique for colorizing infrared home videos. We first convert the green scale video into grayscale, afterwards our technique involves generating key-frames for every shot and then building up a one to one correspondence map between the key frames and the designated color images. These pairs are used to generate the color palette table for the video segment, which is then utilized to colorize that segment of the home video. Our novel technique could also be applied for colorizing X-ray videos generated by diagnostic imaging devices as well as surveillance videos generated by baggage scanners at airports.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME2
2003 Scrambling of engineering drawings
abstract
Engineering drawings are ubiquitously used for capturing, conveying and archiving innovative engineering designs. Many engineering companies' core intellectual property resides in their proprietary engineering drawings. Therefore, protection of such vital data is extremely important. This paper provides a swap-transformation matrix based approach to scramble engineering drawings in order to enable confidentiality. An engineering drawing involves the topological information and vertex information. The vertex information is more valuable than the topological information, since the vertices information primarily determines the content of engineering drawings. We argue that the vertex information is more valuable than the topological information, even if some topological information is lost, a drawing may be reconstructed from the vertex positions. We provide for three keys to ensure the security of the drawing. The technique can facilitate digital rights management of engineering drawings. The advantages of our technique are that scrambling is computationally less intensive than encryption and it allows for partial obfuscation.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME2
2003 Semantic video summarization in compressed domain MPEG video
abstract
In this paper, we present a semantic summarization algorithm that interfaces with the metadata and that works in compressed domain, in particular MPEG-1 and MPEG-2 videos. In enabling a summarization algorithm through high-level semantic content, we try to address two major problems. First, we present the facility provided in the DVA system that allows the semi-automatic creation of this metadata. Second, we address the main point of this system which is the utilization of this metadata to filter out frames, creating an abstract of a video summary quality survey indicates that the proposed method performs satisfactorily.
Jek Charlson So Yu, Mohan Kankanhalli, Philippe Mulhem
ICME2
2003 Lossless Watermarking Considering the Human Visual System
Mohammad Awrangjeb, Mohan Kankanhalli
IWDW2
2003 Experience based sampling technique for multimedia analysis
abstract
We present a novel experience based sampling or experiential sampling technique which has the ability to focus on the analysis's task by making use of the contextual information from the environment. In this technique, sensor samples are used to gather information about the current environment and attention samples are used to represent the current state of attention. The task-attended samples are inferred from experience and maintained by a sampling based dynamical system. The multimedia analysis task can then focus on the attention samples only. Moreover, past experiences and the current environment can be used to adaptively correct and tune the attention. Experimental results have been presented to demonstrate the efficacy of our technique.
Jun Wang 0012, Mohan Kankanhalli
ACM Multimedia2
2003 Music scale modeling for melody matching
abstract
Several time series matching techniques have been proposed for content-based music retrieval. These techniques have shown to be robust and effective for music retrieval by acoustic inputs, such as query-by-humming. However, due to the key transposition issue, all the current methods need to search a large space for the proper key in melody matching. This computation can be prohibitive for a practical music retrieval system with a large database.In this paper, we present a music scale modeling technique for melody matching. The root note of music scale (Major or Minor) of a melody is estimated by fitting the notes to a music scale model. The estimated root note can then be used as the key in melody matching. To the best of our knowledge, this is the first approach that utilizes music scale knowledge for retrieval. In our experiments, 96% of the songs in the database (3000 melodies) can fit into the music scale model. Promising results for query-by-humming retrieval have been obtained by using this novel approach.
Yongwei Zhu, Mohan Kankanhalli
ACM Multimedia2
2003 Semantic Video Annotation and Vague Query
Qiuying Zhang, Mohan Kankanhalli, Philippe Mulhem
MMM2
2003 Robust image authentication using content based compression
Ee-Chien Chang, Mohan Kankanhalli
Multim. Syst.2
2003 A digital rights management scheme for broadcast video
Sabu Emmanuel, Mohan Kankanhalli
Multim. Syst.2
2002 Compressed domain object tracking for automatic indexing of objects in MPEG home video
abstract
Object tracking is of utmost importance for automatic indexing of video content. This work presents an object tracker that operates directly on MPEG compressed data. Motion vectors and discrete cosine transform (DCT) coefficients directly available from the compressed video stream are exploited for the purpose of tracking. Tracking proceeds in two steps: motion vector based tracking in P and B frames within the groups of pictures (GOPs), and object identification in I frames. Colour, which is one of the strongest cues for tracking is used for the identification step. Such a system offers speed, simplicity and robustness against occlusion and camera motion, with good intra-shot tracking for shots in excess of 500 frames, as shown in the experimental results.
Radhakrishna S. V. Achanta, Mohan Kankanhalli, Philippe Mulhem
ICME (2)2
2002 Erasing video logos based on image inpainting
abstract
A video logo is usually a declaration of the video copyright. However it sometimes causes visual discomfort due to the presence of multiple logos in videos that have been filed and exchanged by different channels. We present an approach to erase logos from video clips. Based on the histogram energy analysis of the relevant video frames, we obtain the best quality logo frame that can be easily processed in the selected region of video frames. After that, we mark the logo area in the entire sequence of frames and inpaint each frame of the video logo based on color interpolation. We describe our technique and also provide experimental results.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME (2)2
2002 SmartAlbum: a multi-modal photo annotation system
abstract
This demonstration presents a novel application (called SmartAlbum) for photo indexing and retrieval that unifies two different image indexing approaches. The system uses two modalities to extract information about a digital photograph; i.e. content-based and speech annotation for image description. The result is a powerful image retrieval tool that has capabilities beyond what current single-mode retrieval systems can offer. We show on a corpus of 1200 images the interest of our approach.
Tele Tan, Philippe Mulhem, Mohan Kankanhalli
ACM Multimedia4
2002 Detection and removal of lighting & shaking artifacts in home videos
abstract
Many amateur videographers, like home video enthusiasts, may capture videos that are not of a professional quality. Many minor but visually annoying distortions like lighting imbalance and shaking artifacts could be introduced by the unskilled operations of the video camcorder. Since home videos constitute footage of great sentimental value, such videos cannot be summarily discarded. Unlike movies and sitcoms, shot re-takes of important events, such as wedding ceremonies are just not possible. Therefore, such distortions need to be corrected. In this paper, we present a novel method to detect segments of videos that have lighting and shaking artifacts. These segments can then be subjected to a restoration process that can remove these artifacts. We present techniques to correct lighting artifacts by appropriately adjusting the luminance. In order to remove the shaking artifact, image mosaicing is first employed to build a mosaic frame for the segment with the aid of edge blending techniques. Subsequently a Bezier-curve based blending of motion trajectory is employed to perform motion-compensated filtering of the shaking artifact. The restored video is then created by appropriately cropping the mosaic frame based on the compensated motion trajectory. We have implemented the developed techniques and the experimental results on home videos demonstrate the effectiveness of our approach. Detection and removal of artifacts are significant in other videos as well as those obtained from autonomous vehicles, robots and remote sensing.
Wei Qi Yan 0001, Mohan Kankanhalli
ACM Multimedia2
2002 Detection of human faces in a compressed domain for video stratification
Tat-Seng Chua, Mohan Kankanhalli
Vis. Comput.3
2001 Robust Invisible Watermarking of Volume Data Using the 3D DCT
abstract
Proposes a novel watermarking algorithm for 3D volume data based on the spread-spectrum communication technique, which is invisible and robust. "Invisible" means that the 2D rendered image of this watermarked volume is perceptually indistinguishable from that of the original volume. "Robust" watermarking implies that the watermark is resistant to most intentional or unintentional attacks. We have implemented the algorithm in a software package and have conducted experiments showing that the watermark is invisible in the volume-rendered 2D images. This is further confirmed by computing the signal-to-noise ratio (SNR) and the peak signal-to-noise ratio (PSNR) of the 3D watermarked volume data. We addressed different attacks, putting more emphasis on those attacks that are most commonly performed by attackers. The experiments show that the watermarking scheme is robust.
Mohan Kankanhalli
Computer Graphics International3
2001 Copyright Protection For Mpeg-2 Compressed Broadcast Video
abstract
We describe a novel compressed domain (MPEG-2) method to achieve video security and copyright protection in a broadcasting or a multicasting environment. The proposed method is an efficient solution for resolving the conflicting requirements of single copy transmission for broadcasting and multiple unique copies for copyright protection by the broadcaster. The new method supports dynamic join and leave and it relies upon robust invisible watermarking, masking and conditional access techniques to achieve the desired result.
Sabu Emmanuel, Mohan Kankanhalli
ICME2
2001 Melody Curve Processing For Music Retrieval
abstract
There have been several query-by-humming techniques developed for music retrieval. The techniques either are errorprone due to the inaccuracy of the hummed query or force the users to hum according to a metronome. This paper presents a new slope-based query-by-humming technique, in which the retrieval is robust to the inaccuracy in query and the use of metronome is eliminated. We use melody curve to represent the melodies of the original songs and the hummed query. And curve features: slope ranges and spans and note changes are extracted from the melody curves. Music retrieval is done by matching the curve features of query with those of the original musical songs. Results have shown the features and the algorithms are robust to humming inaccuracy.
Yongwei Zhu, Changsheng Xu, Mohan Kankanhalli
ICME3
2000 A caching and streaming framework for mulitmedia
abstract
In this paper, we explore the convergence of the caching and streaming technologies for Internet multimedia. The paper describes a design for a streaming and caching architecture to be deployed on broadband networks. The basis of the work is the proposed Internet standard, Real Time Streaming Protocol (RTSP), likely to be the de-facto standard for web-based A/V caching and streaming, in the near future. The proxies are all managed by an `Intelligent Agent' or `Broker' - this has been designed as an enhanced RTSP proxy server that maintains the state information that is so essential in streaming of media data. In addition, all the caching algorithms run on the broker. Having an intelligent agent or broker ensures that the `simple' caching servers can be easily embedded into the network. However, RTSP does nor have the right model for doing broker based streaming/caching architecture. The work reported here is an attempt to contribute towards that end.
Shantanu Paknikar, Mohan Kankanhalli, K. R. Ramakrishnan, S. H. Srinivasan, Lek Heng Ngoh
ACM Multimedia2
1999 A dual watermarking technique for images
abstract
Digital watermarking is the technique in which a visible/invisible signal (watermark) is embedded in a multimedia document for copyright protection. In this paper, we propose a watermarking scheme called "dual watermarking". Dual watermark is a combination of a visible watermark and an invisible watermark.
Saraju P. Mohanty, K. R. Ramakrishnan, Mohan Kankanhalli
ACM Multimedia (2)3
1999 Color and spatial feature for content-based image retrieval
Mohan Kankanhalli, B. M. Mehtre, Hock Yiung Huang
Pattern Recognit. Lett.1
1998 Content Based Watermarking of Images
Mohan Kankanhalli, K. R. Ramakrishnan, Rajmohan
ACM Multimedia1
1998 Content-Based Image Retrieval Using a Composite Color-Shape Approach
B. M. Mehtre, Mohan Kankanhalli, Wing Foon Lee
Inf. Process. Manag.2
1997 Shape Measures for Content Based Image Retrieval: A Comparison
B. M. Mehtre, Mohan Kankanhalli, Wing Foon Lee
Inf. Process. Manag.2
1997 Benchmarking Multimedia Databases
Arcot Desai Narasimhalu, Mohan Kankanhalli, Jian-Kang Wu
Multim. Tools Appl.2
1996 Cluster-based color matching for image retrieval
Mohan Kankanhalli, B. M. Mehtre, Ran Kang Wu
Pattern Recognit.1
1995 Selectively meshed surface representation
Renben Shu, Mohan Kankanhalli
Comput. Graph.3
1995 Area and Perimeter Computation of the Union of a Set of Iso-Rectangles in Parallel
Mohan Kankanhalli, W. Randolph Franklin
J. Parallel Distributed Comput.1
1995 Color Indexing for Efficient Image Retrieval
G. Phanendra Babu, B. M. Mehtre, Mohan Kankanhalli
Multim. Tools Appl.3
1995 Color matching for image retrieval
B. M. Mehtre, Mohan Kankanhalli, Arcot Desai Narasimhalu, Guo Chang Man
Pattern Recognit. Lett.2
1995 Adaptive marching cubes
Renben Shu, Mohan Kankanhalli
Vis. Comput.3
1994 Handling small features in isosurface generation using Marching Cubes
Renben Shu, Mohan Kankanhalli
Comput. Graph.3
1994 Efficient linear octree generation from voxels
Renben Shu, Mohan Kankanhalli
Image Vis. Comput.2
1993 Multidimensional On-Line Bin-Packing: An Algorithm and its Average-Case Analysis
Ee-Chien Chang, Weiguo Wang, Mohan Kankanhalli
Inf. Process. Lett.3
1993 An adaptive dominant point detection algorithm for digital curves
Mohan Kankanhalli
Pattern Recognit. Lett.1
1990 Parallel object-space hidden surface removal
abstract
A parallel object-space hidden surface removal algorithm for polyhedral scenes is presented. The uniform grid technique is used to achieve parallelism for the hidden line removal. A conflict-detection and back-off strategy is then used to obtain parallelism for the visible region reconstruction from the visible segments. The algorithm has been implemented on a Sequent Balance 21000 shared-memory parallel computer. An average speedup of 10 has been obtained using 15 processors.
W. Randolph Franklin, Mohan Kankanhalli
SIGGRAPH2