EDBT 2026 Demo / reviewers in the wild / expert
Yongkang Wong
dblp:89/7407 · also Yongkang Wang 0001
· DBLP profile ↗
77ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0002-1239-4428ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 14 since 2021Computer networks · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Object-Centric Framework for Video Moment RetrievalabstractMost existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks. Yongkang Wong, Satoshi Yamazaki, Jianquan Liu, Mohan Kankanhalli |
AAAI | 2 |
| 2026 | TailorEdit: An Adaptive Framework for Instruction-Guided Fashion Image EditingabstractFashion image editing has garnered significant attention due to its growing demand in e-commerce, social media, and virtual try-on applications. However, existing methods are typically designed for specific editing tasks in isolation, lacking a unified framework capable of handling diverse editing requirements. This work addresses this limitation from two critical perspectives. First, we constructInstructFashion, a large-scale, high-quality dataset specifically curated for instruction-guided fashion image editing. It is generated through carefully designed pipelines that cover four distinct editing tasks. Second, we proposeTailorEdit, an adaptive framework for instruction-guided fashion image editing. It integrates human segmentation map-based denoising guidance, modular LoRA-based editing experts, and a dynamic expert routing mechanism to enable precise and semantically coherent modifications. Extensive quantitative and qualitative evaluations demonstrate that TailorEdit consistently outperforms state-of-the-art methods in terms of realism, coherence, and instruction adherence. Our code is available at https://github.com/EndaJude/TailorEdit. Xiaoling Gu, Lingda Zhu, Yongkang Wong, Zhou Yu 0001, Huan Li 0003, Zizhao Wu, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Learning to Predict Gradients for Semi-Supervised Continual LearningabstractA key challenge for machine intelligence is to learn new visual concepts without forgetting the previously acquired knowledge. Continual learning (CL) is aimed toward addressing this challenge. However, there still exists a gap between CL and human learning. In particular, humans are able to continually learn from the samples associated with known or unknown labels in their daily lives, whereas existing CL and semi-supervised CL (SSCL) methods assume that the training samples are associated with known labels. Specifically, we are interested in two questions: 1) how to utilize unrelated unlabeled data for the SSCL task and 2) how unlabeled data affect learning and catastrophic forgetting in the CL task. To explore these issues, we formulate a new SSCL method, which can be generically applied to existing CL models. Furthermore, we propose a novel gradient learner to learn from labeled data to predict gradients on unlabeled data. In this way, the unlabeled data can fit into the supervised CL framework. We extensively evaluate the proposed method on mainstream CL methods, adversarial CL (ACL), and semi-supervised learning (SSL) tasks. The proposed method achieves state-of-the-art performance on classification accuracy and backward transfer (BWT) in the CL setting while achieving the desired performance on classification accuracy in the SSL setting. This implies that the unlabeled images can enhance the generalizability of CL models on the predictive ability of unseen data and significantly alleviate catastrophic forgetting. The code is available at https://github.com/luoyan407/grad_prediction.git. Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Implications of Privacy Regulations on Video Surveillance SystemsabstractAdvanced video surveillance systems (VSS), which collect information of every individual who passes through a surveilled area, have become ubiquitous due to its utility for security. However, such proactive monitoring threatens the individual's privacy due to the public's lack of control over personal data. Additionally, individuals or organizations may unethically misuse VSS for other purposes (e.g., individual profiling and unwarranted monitoring). To safeguard individual privacy, various governments have introduced mandatory information privacy regulations (e.g., GDPR, PDPA, and CCPA) to provide extensive guidelines for the purpose of achieving identity confidentiality. Currently, there is a gap between the information privacy regulations and VSS. This article aims to bridge this gap through four contributions. First, this article conceptualizes VSS as comprising various data stages based on the idea of data lifecycle and studies the implications of existing regulations on VSS. Second, we conducted a survey in ASEAN and European regions to understand the public perception of data risks at each data stage. Third, we review existing privacy-enhancing technologies and its relation to each data stage. Finally, we discuss open research problems in order to realize privacy-aware VSS. Kajal Kansal, Yongkang Wong, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Finetuning Text-to-Image Diffusion Models for FairnessabstractThe rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for minority groups. In this work, we frame fairness as a distributional alignment problem. Our solution consists of two main technical contributions: (1) a distributional alignment loss that steers specific characteristics of the generated images towards a user-defined target distribution, and (2) adjusted direct finetuning of diffusion model's sampling process (adjusted DFT), which leverages an adjusted gradient to directly optimize losses defined on the generated images. Empirically, our method markedly reduces gender, racial, and their intersectional biases for occupational prompts. Gender bias is significantly reduced even when finetuning just five soft tokens. Crucially, our method supports diverse perspectives of fairness beyond absolute equality, which is demonstrated by controlling age to a 75% young and 25% old distribution while simultaneously debiasing gender and race. Finally, our method is scalable: it can debias multiple concepts at once by simply including these prompts in the finetuning data. We share code and various fair diffusion model adaptors at https://sail-sg.github.io/finetune-fair-diffusion/. Tianyu Pang, Yongkang Wong, Mohan Kankanhalli |
ICLR | 5 |
| 2024 | Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition LearningabstractPrevious efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending the context of vision and language, we introduce Multimodal Composition Learning (MCL) for the purpose of mapping or aligning the vision and language input. In particular, we introduce two tasks: Multimodal-Context Captioning (MC-Cap) and Multimodal-Context Retrieval (MC-Ret) to guide a frozen LLM in comprehending the vision and language context. These specialized tasks are crafted to improve the LLM’s capacity for efficient processing and utilization of multimodal inputs, thereby enhancing its proficiency in generating more accurate text or visual representations. Extensive experiments on both retrieval tasks (i.e., zero-shot composed image retrieval, visual storytelling image retrieval and visual dialog image retrieval) and text generation tasks (i.e., visual question answering) demonstrate the effectiveness of the proposed method. The code is available at: https://github.com/dhg-wei/MCL. Hehe Fan, Yongkang Wong, Yi Yang 0001, Mohan Kankanhalli |
ICML | 3 |
| 2024 | MCM: Multi-condition Motion Synthesis Framework
Zeyu Ling, Bo Han 0003, Yongkang Wong, Mohan Kankanhalli, Weidong Geng |
IJCAI | 3 |
| 2024 | TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-AlignmentabstractRecent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inefficient language supervision in recent web-collected video-text datasets. In this paper, we introduce Text-Only Pre-Alignment (TOPA), a novel approach to extend large language models (LLMs) for video understanding, without the need for pre-training on real video data. Specifically, we first employ an advanced LLM to automatically generate Textual Videos comprising continuous textual frames, along with corresponding annotations to simulate real video-text data. Then, these annotated textual videos are used to pre-align a language-only LLM with the video modality. To bridge the gap between textual and real videos, we employ the CLIP model as the feature extractor to align image and text modalities. During text-only pre-alignment, the continuous textual frames, encoded as a sequence of CLIP text features, are analogous to continuous CLIP image features, thus aligning the LLM with real video representation. Extensive experiments, including zero-shot evaluation and finetuning on various video understanding tasks, demonstrate that TOPA is an effective and efficient framework for aligning video content with LLMs. In particular, without training on any video data, the TOPA-Llama2-13B model achieves a Top-1 accuracy of 51.0% on the challenging long-form video understanding benchmark, Egoschema. This performance surpasses previous video-text pre-training approaches and proves competitive with recent GPT-3.5 based video agents. Hehe Fan, Yongkang Wong, Mohan Kankanhalli, Yi Yang 0001 |
NeurIPS | 3 |
| 2024 | Privacy-Enhancing Person Re-identification Framework - A Dual-Stage ApproachabstractIn this work, we show that deep learning-based re-identification (Re-ID) models, albeit trained only with a Re-ID objective (i.e. if two samples belong to the same identity), encode personally identifiable information (PII) in the learned features that may lead to serious privacy concerns. In cognizance of the modern privacy regulations on protecting PII, we propose a novel dual-stage person Re-ID framework that (1) suppresses the PII from the discriminative features, and (2) introduces a controllable privacy mechanism through differential privacy. The former is achieved with a self-supervised de-identification (De-ID) decoder and an adversarial-identity (Adv-ID) module, whereas the latter mechanism leverages a controllable privacy budget to generate a privacy-protected gallery with a Gaussian noise generator. Furthermore, we introduce the notion of a privacy metric to quantify the privacy leakage in Re-ID features which is not explicitly examined in prior work. We demonstrate the feasibility of our approach in achieving a better trade-off between utility and privacy through rigorous experiments on person Re-ID benchmarks. Kajal Kansal, Yongkang Wong, Mohan Kankanhalli |
WACV | 2 |
| 2024 | Multi2Human: Controllable human image generation with multimodal controls
Xiaoling Gu, Shengwenzhuo Xu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
Neurocomputing | 3 |
| 2024 | Unsupervised Domain Adaptation by Causal Learning for Biometric Signal-based HCIabstractBiometric signal based human-computer interface (HCI) has attracted increasing attention due to its wide application in healthcare, entertainment, neurocomputing, and so on. In recent years, deep learning-based approaches have made great progress on biometric signal processing. However, the state-of-the-art (SOTA) approaches still suffer from model degradation across subjects or sessions. In this work, we propose a novel unsupervised domain adaptation approach for biometric signal-based HCI via causal representation learning. Specifically, three kinds of interventions on biometric signals (i.e., subjects, sessions, and trials) can be selected to generalize deep models across the selected intervention. In the proposed approach, a generative model is trained for producing intervened features that are subsequently used for learning transferable and causal relations with three modes. Experiments on the EEG-based emotion recognition task and sEMG-based gesture recognition task are conducted to confirm the superiority of our approach. An improvement of +0.21% on the task of inter-subject EEG-based emotion recognition is achieved using our approach. Besides, on the task of inter-session sEMG-based gesture recognition, our approach achieves improvements of +1.47%, +3.36%, +1.71%, and +1.01% on sEMG datasets including CSL-HDEMG, CapgMyo DB-b, 3DC, and Ninapro DB6, respectively. The proposed approach also works on the task of inter-trial sEMG-based gesture recognition and an average improvement of +0.66% on Ninapro databases is achieved. These experimental results show the superiority of the proposed approach compared with the SOTA unsupervised domain adaptation methods on HCIs based on biometric signal. Qingfeng Dai, Yongkang Wong, Guofei Sun, Zhou Zhou 0012, Mohan Kankanhalli, Weidong Geng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | PAINT: Photo-realistic Fashion Design SynthesisabstractIn this article, we investigate a new problem of generating a variety of multi-view fashion designs conditioned on a human pose and texture examples of arbitrary sizes, which can replace the repetitive and low-level design work for fashion designers. To solve this challenging multi-modal image translation problem, we propose a novel Photo-reAlistic fashIon desigN synThesis (PAINT) framework, which decomposes the framework into three manageable stages. In the first stage, we employ a Layout Generative Network (LGN) to transform an input human pose into a series of person semantic layouts. In the second stage, we propose a Texture Synthesis Network (TSN) to synthesize textures on all transformed semantic layouts. Specifically, we design a novel attentive texture transfer mechanism for precisely expanding texture patches to the irregular clothing regions of the target fashion designs. In the third stage, we leverage an Appearance Flow Network (AFN) to generate the fashion design images of other viewpoints from a single-view observation by learning 2D multi-scale appearance flow fields. Experimental results demonstrate that our method is capable of generating diverse photo-realistic multi-view fashion design images with fine-grained appearance details conditioned on the provided multiple inputs. The source code and trained models are available at https://github.com/gxl-groups/PAINT . Xiaoling Gu, Jie Huang 0033, Yongkang Wong, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Recurrent Appearance Flow for Occlusion-Free Virtual Try-OnabstractImage-based virtual try-on aims at transferring a target in-shop garment onto a reference person, and has garnered significant attention from the research communities recently. However, previous methods have faced severe challenges in handling occlusion problems. To address this limitation, we classify occlusion problems into three types based on the reference person’s arm postures: single-arm occlusion , two-arm non-crossed occlusion , and two-arm crossed occlusion . Specifically, we propose a novel Occlusion-Free Virtual Try-On Network (OF-VTON) that effectively overcomes these occlusion challenges. The OF-VTON framework consists of two core components: (i) a new Recurrent Appearance Flow based Deformation (RAFD) model that robustly aligns the in-shop garment to the reference person by adopting a multi-task learning strategy . This model jointly produces the dense appearance flow to warp the garment and predicts a human segmentation map to provide semantic guidance for the subsequent image synthesis model. (ii) a powerful Multi-mask Image SynthesiS (MISS) model that generates photo-realistic try-on results by introducing a new mask generation and selection mechanism . Experimental results demonstrate that our proposed OF-VTON significantly outperforms existing state-of-the-art methods by mitigating the impact of occlusion problems. Our code is available at https://github.com/gxl-groups/OF-VTON . Xiaoling Gu, Junkai Zhu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | NarSUM '23: The 2nd Workshop on User-Centric Narrative Summarization of Long VideosabstractWith video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization, etc. Mohan Kankanhalli, Ioannis Patras, Jianquan Liu, Yongkang Wong, Takahiro Komamizu, Satoshi Yamazaki, Karen Stephen, Kajal Kansal |
ACM Multimedia | 4 |
| 2023 | Fair Representation: Guaranteeing Approximate Multiple Group Fairness for Unknown TasksabstractMotivated by scenarios where data is used for diverse prediction tasks, we study whether fair representation can be used to guarantee fairness for unknown tasks and for multiple fairness notions. We consider seven group fairness notions that cover the concepts of independence, separation, and calibration. Against the backdrop of the fairness impossibility results, we explore approximate fairness. We prove that, although fair representation might not guarantee fairness for all prediction tasks, it does guarantee fairness for an important subset of tasks-the tasks for which the representation is discriminative. Specifically, all seven group fairness notions are linearly controlled by fairness and discriminativeness of the representation. When an incompatibility exists between different fairness notions, fair and discriminative representation hits the sweet spot that approximately satisfies all notions. Motivated by our theoretical findings, we propose to learn both fair and discriminative representations using pretext loss which self-supervises learning, and Maximum Mean Discrepancy as a fair regularizer. Experiments on tabular, image, and face datasets show that using the learned representation, downstream predictions that we are unaware of when learning the representation indeed become fairer. The fairness guarantees computed from our theoretical results are all valid. Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning to Minimize the Remainder in Supervised LearningabstractThe learning process of deep learning methods usually updates the model’s parameters in multiple iterations. Each iteration can be viewed as the first-order approximation of Taylor’s series expansion. The remainder, which consists of higher-order terms, is usually ignored in the learning process for simplicity. This learning scheme empowers various multimedia-based applications, such as image retrieval, recommendation system, and video search. Generally, multimedia data (e.g.images) are semantics-rich and high-dimensional, hence the remainders of approximations are possibly non-zero. In this work, we consider that the remainder is informative and study how it affects the learning process. To this end, we propose a new learning approach, namely gradient adjustment learning (GAL), to leverage the knowledge learned from the past training iterations to adjust vanilla gradients, such that the remainders are minimized and the approximations are improved. The proposed GAL is model- and optimizer-agnostic, and is easy to adapt to the standard learning framework. It is evaluated on three tasks,i.e.image classification, object detection, and regression, with state-of-the-art models and optimizers. The experiments show that the proposed GAL consistently enhances the evaluated models, whereas the ablation studies validate various aspects of the proposed GAL. The code is available athttps://github.com/luoyan407/gradient_adjustment.git. Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Semantic-Aware Triplet Loss for Image ClassificationabstractSuccessful image classification requires a discriminative representation learning model for images. To approach this idea, deep metric learning (DML), serving as building a basic feature space with a pre-defined metric, has demonstrated compelling performance over the years. DML is often implemented with a carefully crafted loss function, such as the representative triplet loss, which encourages a positive sample to be by a fixed margin closer to the anchor than the negative. Despite its efficacy, the negative samples are treated uniformly, rendering the feature space less informative since different negative samples can be largely different from the anchor. In this work, we, for the first time, propose to exploit the semantic information inherent in discrete class labels as an aid for the triplet loss. Specifically, we build a bi-level negative sampling strategy,i.e., strong negative and weak negative sampling, with the guidance of an external knowledge source, from which rich class semantics can be extracted. With several fine-grained and complementary triplet losses based on this strategy, our method is enhanced with semantic awareness for image classification. In addition, to coordinate with the complicated training dynamics, we devise an ad-hoc Semantic Relation Weighting module, which consistently inspects model states and dynamically adjusts the importance of each triplet loss. It is worth noting that our method is plug-and-play, and we thus test its validity over various backbones and knowledge sources. Both qualitative and quantitative experimental results on benchmark datasets demonstrate the effectiveness of employing semantics for image classification. Guangzhi Wang, Ziwei Xu 0001, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Multim. | 4 |
| 2022 | Chairs Can Be Stood On: Overcoming Object Bias in Human-Object Interaction Detection
Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli |
ECCV (24) | 3 |
| 2022 | A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQAabstractKnowledge-based Visual Question Answering (VQA) expects models to rely on external knowledge for robust answer prediction. Though significant it is, this paper discovers several leading factors impeding the advancement of current state-of-the-art methods. On the one hand, methods which exploit the explicit knowledge take the knowledge as a complement for the coarsely trained VQA model. Despite their effectiveness, these approaches often suffer from noise incorporation and error propagation. On the other hand, pertaining to the implicit knowledge, the multi-modal implicit knowledge for knowledge-based VQA still remains largely unexplored. This work presents a unified end-to-end retriever-reader framework towards knowledge-based VQA. In particular, we shed light on the multi-modal implicit knowledge from vision-language pre-training models to mine its potential in knowledge reasoning. As for the noise problem encountered by the retrieval operation on explicit knowledge, we design a novel scheme to create pseudo labels for effective knowledge supervision. This scheme is able to not only provide guidance for knowledge retrieval, but also drop these instances potentially error-prone towards question answering. To validate the effectiveness of the proposed method, we conduct extensive experiments on the benchmark dataset. The experimental results reveal that our method outperforms existing baselines by a noticeable margin. Beyond the reported numbers, this paper further spawns several insights on knowledge utilization for future research with some empirical findings. Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng 0001, Mohan Kankanhalli |
ACM Multimedia | 3 |
| 2022 | NarSUM '22: 1st Workshop on User-centric Narrative Summarization of Long VideosabstractWith video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization etc. Mohan Kankanhalli, Jianquan Liu, Yongkang Wong, Karen Stephen, Rishabh Sheoran, Anusha Bhamidipati |
ACM Multimedia | 3 |
| 2022 | Distance Matters in Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection has received considerable attention in the context of scene understanding. Despite the growing progress, we realize existing methods often perform unsatisfactorily on distant interactions, where the leading causes are two-fold: 1) Distant interactions are by nature more difficult to recognize than close ones. A natural scene often involves multiple humans and objects with intricate spatial relations, making the interaction recognition for distant human-object largely affected by complex visual context. 2) Insufficient number of distant interactions in datasets results in under-fitting on these instances. To address these problems, we propose a novel two-stage method for better handling distant interactions in HOI detection. One essential component in our method is a novel Far Near Distance Attention module. It enables information propagation between humans and objects, whereby the spatial distance is skillfully taken into consideration. Besides, we devise a novel Distance-Aware loss function which leads the model to focus more on distant yet rare interactions. We conduct extensive experiments on HICO-DET and V-COCO datasets. The results show that the proposed method surpass existing methods significantly, leading to new state-of-the-art results. Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli |
ACM Multimedia | 3 |
| 2022 | Compute to Tell the Tale: Goal-Driven Narrative GenerationabstractMan is by nature a social animal. One important facet of human evolution is through narrative imagination, be it fictional or factual, and to tell the tale to other individuals. The factual narrative, such as news, journalism, field report, etc., is based on real-world events and often requires extensive human efforts to create. In the era of big data where video capture devices are commonly available everywhere, a massive amount of raw videos (including life-logging, dashcam or surveillance footage) are generated daily. As a result, it is rather impossible for humans to digest and analyze these video data. This paper reviews the problem of computational narrative generation where a goal-driven narrative (in the form of text with or without video) is generated from a single or multiple long videos. Importantly, the narrative generation problem makes itself distinguished from the existing literature by its focus on a comprehensive understanding of user goal, narrative structure and open-domain input. We tentatively outline a general narrative generation framework and discuss the potential research problems and challenges in this direction. Informed by the real-world impact of narrative generation, we then illustrate several practical use cases in Video Logging as a Service platform which enables users to get more out of the data through a goal-driven intelligent storytelling AI agent. Yongkang Wong, Shaojing Fan, Ziwei Xu 0001, Karen Stephen, Rishabh Sheoran, Anusha Bhamidipati, Vivek Barsopia, Jianquan Liu, Mohan Kankanhalli |
ACM Multimedia | 1 |
| 2022 | Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationabstractWe propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output and the constraints. We propose a comprehensive set of constraints, which are implicit in data annotations, and incorporate them with deep networks via DTL. We evaluate the effectiveness of DTL on the temporal action segmentation task and observe improved performance and reduced logical errors in the output of different task models. Furthermore, we provide an extensive analysis to visualize the desirable effects of DTL. Ziwei Xu 0001, Yogesh S. Rawat, Yongkang Wong, Mohan Kankanhalli, Mubarak Shah |
NeurIPS | 3 |
| 2022 | Relation-Aware Compositional Zero-Shot Learning for Attribute-Object Pair RecognitionabstractThis paper proposes a novel model for recognizing images with composite attribute-object concepts, notably for composite concepts that are unseen during model training. We aim to explore the three key properties required by the task — relation-aware, consistent, and decoupled—to learn rich and robust features for primitive concepts that compose attribute-object pairs. To this end, we propose the Blocked Message Passing Network (BMP-Net). The model consists of two modules. The concept module generates semantically meaningful features for primitive concepts, whereas the visual module extracts visual features for attributes and objects from input images. A message passing mechanism is used in the concept module to capture the relations between primitive concepts. Furthermore, to prevent the model from being biased towards seen composite concepts and reduce the entanglement between attributes and objects, we propose a blocking mechanism that equalizes the information available to the model for both seen and unseen concepts. Extensive experiments and ablation studies on two benchmarks show the efficacy of the proposed model. Ziwei Xu 0001, Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Multim. | 3 |
| 2022 | Enhanced 3D Shape Reconstruction With Knowledge Graph of Category ConceptabstractReconstructing three-dimensional (3D) objects from images has attracted increasing attention due to its wide applications in computer vision and robotic tasks. Despite the promising progress of recent deep learning–based approaches, which directly reconstruct the full 3D shape without considering the conceptual knowledge of the object categories, existing models have limited usage and usually create unrealistic shapes. 3D objects have multiple forms of representation, such as 3D volume, conceptual knowledge, and so on. In this work, we show that the conceptual knowledge for a category of objects, which represents objects as prototype volumes and is structured by graph, can enhance the 3D reconstruction pipeline. We propose a novel multimodal framework that explicitly combines graph-based conceptual knowledge with deep neural networks for 3D shape reconstruction from a single RGB image. Our approach represents conceptual knowledge of a specific category as a structure-based knowledge graph. Specifically, conceptual knowledge acts as visual priors and spatial relationships to assist the 3D reconstruction framework to create realistic 3D shapes with enhanced details. Our 3D reconstruction framework takes an image as input. It first predicts the conceptual knowledge of the object in the image, then generates a 3D object based on the input image and the predicted conceptual knowledge. The generated 3D object satisfies the following requirements: (1) it is consistent with the predicted graph in concept, and (2) consistent with the input image in geometry. Extensive experiments on public datasets (i.e., ShapeNet, Pix3D, and Pascal3D+) with 13 object categories show that (1) our method outperforms the state-of-the-art methods, (2) our prototype volume-based conceptual knowledge representation is more effective, and (3) our pipeline-agnostic approach can enhance the reconstruction quality of various 3D shape reconstruction pipelines. Guofei Sun, Yongkang Wong, Mohan Kankanhalli, Weidong Geng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventionsabstract3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a step towards training robust models for cross-domain pose estimation task, which brings together ideas from causal representation learning and generative adversarial networks. Specifically, this paper introduces a novel framework for causal representation learning which explicitly exploits the causal structure of the task. We consider changing domain as interventions on images under the data-generation process and steer the generative model to produce counterfactual features. This help the model learn transferable and causal relations across different domains. Our framework is able to learn with various types of unlabeled datasets. We demonstrate the efficacy of our proposed method on both human and hand pose estimation task. The experiment results show the proposed approach achieves state-of-the-art performance on most datasets for both domain adaptation and domain generalization settings. Xiheng Zhang, Yongkang Wong, Juwei Lu, Mohan Kankanhalli, Weidong Geng |
ICCV | 2 |
| 2021 | Learning to Predict Trustworthiness with Steep Slope LossabstractUnderstanding the trustworthiness of a prediction yielded by a classifier is critical for the safe and effective use of AI models. Prior efforts have been proven to be reliable on small-scale datasets. In this work, we study the problem of predicting trustworthiness on real-world large-scale datasets, where the task is more challenging due to high-dimensional features, diverse visual concepts, and a large number of samples. In such a setting, we observe that the trustworthiness predictors trained with prior-art loss functions, i.e., the cross entropy loss, focal loss, and true class probability confidence loss, are prone to view both correct predictions and incorrect predictions to be trustworthy. The reasons are two-fold. Firstly, correct predictions are generally dominant over incorrect predictions. Secondly, due to the data complexity, it is challenging to differentiate the incorrect predictions from the correct ones on real-world large-scale datasets. To improve the generalizability of trustworthiness predictors, we propose a novel steep slope loss to separate the features w.r.t. correct predictions from the ones w.r.t. incorrect predictions by two slide-like curves that oppose each other. The proposed loss is evaluated with two representative deep learning models, i.e., Vision Transformer and ResNet, as trustworthiness predictors. We conduct comprehensive experiments and analyses on ImageNet, which show that the proposed loss effectively improves the generalizability of trustworthiness predictors. The code and pre-trained trustworthiness predictors for reproducibility are available at \url{https://github.com/luoyan407/predict_trustworthiness}. Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
NeurIPS | 2 |
| 2021 | Unsupervised Motion Representation Learning with Capsule AutoencodersabstractWe propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and semantic-agnostic snippets. In the higher level, the snippets are aggregated to form full-length semantic-aware segments. For both levels, we represent motion with a set of learned transformation invariant templates and the corresponding geometric transformations by using capsule autoencoders of a novel design. This leads to a robust and efficient encoding of viewpoint changes. MCAE is evaluated on a novel Trajectory20 motion dataset and various real-world skeleton-based human action datasets. Notably, it achieves better results than baselines on Trajectory20 with considerably fewer parameters and state-of-the-art performance on the unsupervised skeleton-based action recognition task. Ziwei Xu 0001, Yongkang Wong, Mohan Kankanhalli |
NeurIPS | 3 |
| 2021 | Direction Concentration Learning: Enhancing Congruency in Machine LearningabstractOne of the well-known challenges in computer vision tasks is the visual diversity of images, which could result in an agreement or disagreement between the learned knowledge and the visual content exhibited by the current observation. In this work, we first define such an agreement in a concepts learning process as congruency. Formally, given a particular task and sufficiently large dataset, the congruency issue occurs in the learning process whereby the task-specific semantics in the training data are highly varying. We propose a Direction Concentration Learning (DCL) method to improve congruency in the learning process, where enhancing congruency influences the convergence path to be less circuitous. The experimental results show that the proposed DCL method generalizes to state-of-the-art models and optimizers, as well as improves the performances of saliency prediction task, continual learning task, and classification task. Moreover, it helps mitigate the catastrophic forgetting problem in the continual learning task. The code is publicly available at https://github.com/luoyan407/congruency. Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Scene Graph Inference via Multi-Scale Context ModelingabstractThe scene graph generated for an image structurally represents its object interactions and it substantially aids image scene understanding. To the best of our knowledge, most current works on scene graph generation chiefly focus on pairwise object regions for object and relation inference while ignoring the global visual context outside of these regions. Guided by the intuition that object/relation inference can benefit from the visual context within an image, this paper proposes a multi-scale context modeling method, which can jointly discover and integrate the complementary object-centric and region-centric context for scene graph inference. While both the object-centric and region-centric contexts are separately modeled by their individual modules, a bi-directional message propagation strategy is designed to mutually reinforce the context modeling. A context-fused inference is then proposed to integrate the multi-scale context to guide scene graph inference. Extensive experiments establish that this method can achieve competitive performance compared to the state-of-the-art methods on three benchmarks. Additional ablation studies further validate its effectiveness. Code has been made available at: https://github.com/ningxu1990/MSCM. Ning Xu 0003, Anan Liu, Yongkang Wong, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Toward Multi-Modal Conditioned Fashion Image TranslationabstractHaving the capability to synthesize photo-realistic fashion product images conditioned on multiple attributes or modalities would bring many new exciting applications. In this work, we propose an end-to-end network architecture that built upon a new generative adversarial network for automatically synthesizing photo-realistic images of fashion products under multiple conditions. Given an input pose image that consists of a 2D skeleton pose and a sentence description of products, our model synthesizes a fashion image preserving the same pose and wearing the fashion products described as the text. Specifically, the generator$G$tries to generate realistic-looking fashion images based on a$\langle \mathsf {pose}, \mathsf {text} \rangle$pair condition to fool the discriminator. An attention network is added for enhancing the generator, which predicts a probability map indicating which part of the image needs to be attended for translation. In contrast, the discriminator$D$distinguishes real images from the translated ones based on the input pose image and text description. The discriminator is divided into two multi-scale sub-discriminators for improving image distinguishing task. Quantitative and qualitative analysis demonstrates that our method is capable of synthesizing realistic images that retain the poses of given images while matching the semantics of provided sentence descriptions. Xiaoling Gu, Jun Yu 0002, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Multim. | 3 |
| 2021 | DeepDance: Music-to-Dance Motion Choreography With Adversarial LearningabstractThe creation of improvised dancing choreographies is an important research field of cross-modal analysis. A key point of this task is how to effectively create and correlate music and dance with a probabilistic one-to-many mapping, which is essential to create realistic dances of various genres. To address this issue, we propose a GAN-based cross-modal association framework, DeepDance, which correlates two different modalities (dance motion and music) together, aiming at creating the desired dance sequence in terms of the input music. Its generator is to predictively produce the dance movements best-fit to current music piece by learning from examples. In another hand, its discriminator acts as an external evaluation from the audience and judges the whole performance. The generated dance movements and the corresponding input music are considered to be well-matched if the discriminator cannot distinguish the generated movements from the training samples according to the estimated probability. By adding motion consistency constraints in our loss function, the proposed framework is able to create long realistic dance sequences. To alleviate the problem of expensive and inefficient data collection, we propose an effective approach to create a large-scale dataset, YouTube-Dance3D, from open data source. Extensive experiments on currently available music-dance datasets and our YouTube-Dance3D dataset demonstrate that our approach effectively captures the correlation between music and dance and can be used to choreograph appropriate dance sequences. Guofei Sun, Yongkang Wong, Zhiyong Cheng 0001, Mohan Kankanhalli, Weidong Geng |
IEEE Trans. Multim. | 2 |
| 2020 | n-Reference Transfer Learning for Saliency Prediction
Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
ECCV (8) | 2 |
| 2020 | Weakly-Supervised Multi-Person Action Recognition in 360° VideosabstractThe recent development of commodity 360° cameras have enabled a single video to capture an entire scene, which endows promising potentials in surveillance scenarios. However, research in omnidirectional video analysis has lagged behind the hardware advances. In this work, we address the important problem of action recognition in topview 360° videos. Due to the wide filed-of-view, 360° videos usually capture multiple people performing actions at the same time. Furthermore, the appearance of people are deformed. The proposed framework first transforms top-view omnidirectional videos into panoramic videos using a calibrationfree method. Then spatial-temporal features are extracted using region-based 3D CNNs for action recognition. We propose a weakly-supervised method based on multiinstance multi-label learning, which trains the model to recognize and localize multiple actions in a video using only video-level action labels as supervision. We perform experiments to quantitatively validate the efficacy of the proposed method over state-of-the-art baselines and variants of our model, and qualitatively demonstrate action localization results. To enable research in this direction, we introduce the 360Action dataset. It is the first omnidirectional video dataset for multi-person action recognition with a diverse set of scenes, actors and actions. The dataset is available at https://github.com/ryukenzen/360action. Junnan Li 0001, Jianquan Liu, Yongkang Wong, Shoji Nishimura, Mohan Kankanhalli |
WACV | 3 |
| 2020 | GradMix: Multi-source Transfer across Domains and TasksabstractThe computer vision community is witnessing an unprecedented rate of new tasks being proposed and addressed, thanks to the deep convolutional networks' capability to find complex mappings from X to Y. The advent of each task often accompanies the release of a large-scale annotated dataset, for supervised training of deep network. However, it is expensive and time-consuming to manually label sufficient amount of training data. Therefore, it is important to develop algorithms that can leverage off-the-shelf labeled dataset to learn useful knowledge for the target task. While previous works mostly focus on transfer learning from a single source, we study multi-source transfer across domains and tasks (MS-DTT), in a semi-supervised setting. We propose GradMix, a model-agnostic method applicable to any model trained with gradient-based learning rule, to transfer knowledge via gradient descent by weighting and mixing the gradients from all sources during training. GradMix follows a meta-learning objective, which assigns layer-wise weights to the source gradients, such that the combined gradient follows the direction that minimize the loss for a small set of samples from the target dataset. In addition, we propose to adaptively adjust the learning rate for each mini-batch based on its importance to the target task, and a pseudo-labeling method to leverage the unlabeled samples in the target domain. We conduct MS-DTT experiments on two tasks: digit recognition and action recognition, and demonstrate the advantageous performance of the proposed method against multiple baselines. Junnan Li 0001, Ziwei Xu 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
WACV | 3 |
| 2020 | Visual Social Relationship Recognition
Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
Int. J. Comput. Vis. | 2 |
| 2020 | Unsupervised Online Video Object Segmentation With Motion Property UnderstandingabstractUnsupervised video object segmentation aims to automatically segment moving objects over an unconstrained video without any user annotation. So far, only few unsupervised online methods have been reported in the literature, and their performance is still far from satisfactory because the complementary information from future frames cannot be processed under online setting. To solve this challenging problem, in this paper, we propose a novel unsupervised online video object segmentation (UOVOS) framework by construing the motion property to mean moving in concurrence with a generic object for segmented regions. By incorporating the salient motion detection and the object proposal, a pixel-wise fusion strategy is developed to effectively remove detection noises, such as dynamic background and stationary objects. Furthermore, by leveraging the obtained segmentation from immediately preceding frames, a forward propagation algorithm is employed to deal with unreliable motion detection and object proposals. Experimental results on several benchmark datasets demonstrate the efficacy of the proposed method. Compared to state-of-the-art unsupervised online segmentation algorithms, the proposed method achieves an absolute gain of 6.2%. Moreover, our method achieves better performance than the best unsupervised offline algorithm on the DAVIS-2016 benchmark dataset. Our code is available on the project website: https://www.github.com/visiontao/uovos. Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Image Process. | 4 |
| 2020 | Video Storytelling: Textual Summaries for EventsabstractBridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph generation. In this paper, we introduce the problem of video storytelling, which aims at generating coherent and succinct stories for long videos. Video storytelling introduces new challenges, mainly due to the diversity of the story and the length and complexity of the video. We propose novel methods to address the challenges. First, we propose a context-aware framework for multimodal embedding learning, where we design a residual bidirectional recurrent neural network to leverage contextual information from past and future. The multimodal embedding is then used to retrieve sentences for video clips. Second, we propose a Narrator model to select clips that are representative of the underlying storyline. The Narrator is formulated as a reinforcement learning agent, which is trained by directly optimizing the textual metric of the generated story. We evaluate our method on the video story dataset, a new dataset that we have collected to enable the study. We compare our method with multiple state-of-the-art baselines and show that our method achieves better performance, in terms of quantitative measures and user study. Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
IEEE Trans. Multim. | 2 |
| 2020 | Interact as You Intend: Intention-Driven Human-Object Interaction DetectionabstractThe recent advances in instance-level detection tasks lay strong foundation for genuine comprehension of the visual scenes. However, the ability to fully comprehend a social scene is still in its preliminary stage. In this work, we focus on detecting human-object interactions (HOIs) in social scene images, which is demanding in terms of research and increasingly useful for practical applications. To undertake social tasks interacting with objects, humans direct their attention and move their body based on their intention. Based on this observation, we provide a unique computational perspective to explore human intention in HOI detection. Specifically, the proposed human intention-driven HOI detection (iHOI) framework models human pose with the relative distances from body joints to the object instances. It also utilizes human gaze to guide the attended contextual regions in a weakly-supervised setting. In addition, we propose a hard negative sampling strategy to address the problem of mis-grouping. We perform extensive experiments on two benchmark datasets, namely V-COCO and HICO-DET. The efficacy of each proposed component has also been validated. Bingjie Xu 0002, Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
IEEE Trans. Multim. | 3 |
| 2020 | G-Softmax: Improving Intraclass Compactness and Interclass Separability of FeaturesabstractIntraclass compactness and interclass separability are crucial indicators to measure the effectiveness of a model to produce discriminative features, where intraclass compactness indicates how close the features with the same label are to each other and interclass separability indicates how far away the features with different labels are. In this paper, we investigate intraclass compactness and interclass separability of features learned by convolutional networks and propose a Gaussian-based softmax ( G -softmax) function that can effectively improve intraclass compactness and interclass separability. The proposed function is simple to implement and can easily replace the softmax function. We evaluate the proposed G -softmax function on classification data sets (i.e., CIFAR-10, CIFAR-100, and Tiny ImageNet) and on multilabel classification data sets (i.e., MS COCO and NUS-WIDE). The experimental results show that the proposed G -softmax function improves the state-of-the-art models across all evaluated data sets. In addition, the analysis of the intraclass compactness and interclass separability demonstrates the advantages of the proposed function over the softmax function, which is consistent with the performance improvement. More importantly, we observe that high intraclass compactness and interclass separability are linearly correlated with average precision on MS COCO and NUS-WIDE. This implies that the improvement of intraclass compactness and interclass separability would lead to the improvement of average precision. Yan Luo 0002, Yongkang Wong, Mohan Kankanhalli, Qi Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Learning Controllable Face Generator from Disjoint Datasets
Jing Li 0050, Yongkang Wong, Terence Sim |
CAIP (1) | 2 |
| 2019 | Learning to Learn From Noisy Labeled DataabstractDespite the success of deep neural networks (DNNs) in image classification tasks, the human-level performance relies on massive training data with high-quality manual annotations, which are expensive and time-consuming to collect. There exist many inexpensive data sources on the web, but they tend to contain inaccurate labels. Training on noisy labeled datasets causes performance degradation because DNNs can easily overfit to the label noise. To overcome this problem, we propose a noise-tolerant training algorithm, where a meta-learning update is performed prior to conventional gradient update. The proposed meta-learning method simulates actual training by generating synthetic noisy labels, and train the model such that after one gradient update using each set of synthetic noisy labels, the model does not overfit to the specific noise. We conduct extensive experiments on the noisy CIFAR-10 dataset and the Clothing1M dataset. The results demonstrate the advantageous performance of the proposed method compared to several state-of-the-art baselines. Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
CVPR | 2 |
| 2019 | Learning to Detect Human-Object Interactions With KnowledgeabstractThe recent advances in instance-level detection tasks lay a strong foundation for automated visual scenes understanding. However, the ability to fully comprehend a social scene still eludes us. In this work, we focus on detecting human-object interactions (HOIs) in images, an essential step towards deeper scene understanding. HOI detection aims to localize human and objects, as well as to identify the complex interactions between them. Innate in practical problems with large label space, HOI categories exhibit a long-tail distribution, i.e., there exist some rare categories with very few training samples. Given the key observation that HOIs contain intrinsic semantic regularities despite they are visually diverse, we tackle the challenge of long-tail HOI categories by modeling the underlying regularities among verbs and objects in HOIs as well as general relationships. In particular, we construct a knowledge graph based on the ground-truth annotations of training dataset and external source. In contrast to direct knowledge incorporation, we address the necessity of dynamic image-specific knowledge retrieval by multi-modal learning, which leads to an enhanced semantic embedding space for HOI comprehension. The proposed method shows improved performance on V-COCO and HICO-DET benchmarks, especially when predicting the rare HOI categories. Bingjie Xu 0002, Yongkang Wong, Junnan Li 0001, Qi Zhao 0001, Mohan Kankanhalli |
CVPR | 2 |
| 2019 | Self-supervised Representation Learning Using 360° DataabstractThe amount of 360-degree panoramas shared online has been rapidly increasing due to the availability of affordable and compact omnidirectional cameras, which offers huge amount of new information unavailable before. In this paper, we present the first work to exploit unlabeled 360-degree data for image representation learning. We propose middle-out, a new self-supervised learning task, which leverages the spatial configuration of normal field-of-view images sampled from a 360-degree image as supervisory signal. We train a Siamese ConvNet model to identify the middle image among three shuffled images sampled from a panorama by perspective projection. Compared to previous self-supervised methods that train models using image patches or video frames with limited field-of-view, our method leverages the rich semantic information contained in 360-degree images and enforces the model to not only learn about objects, but also develop a higher-level understanding about object relationships and scene structures. We quantitatively demonstrate that the feature representation learned using the proposed task is useful for a wide range of vision tasks including object classification, object detection, scene classification, semantic segmentation, and geometry estimation. We also qualitatively show that the proposed method can enforce the ConvNet to extract high-level semantic concepts, an ability which previous self-supervised learning methods have not acquired. Junnan Li 0001, Jianquan Liu, Yongkang Wong, Shoji Nishimura, Mohan Kankanhalli |
ACM Multimedia | 3 |
| 2019 | Human-imperceptible Privacy Protection Against MachinesabstractPrivacy concerns with social media have recently been under the spotlight, due to a few incidents on user data leakage on social networking platforms. With the current advances in machine learning and big data, computer algorithms often act as a first-step filter for privacy breaches, by automatically selecting content with sensitive information, such as photos that contain faces or vehicle license plate. In this paper we propose a novel algorithm to protect the sensitive attributes against machines, meanwhile keeping the changes imperceptible to humans. In particular, we first conducted a series of human studies to investigate multiple factors that influence human sensitivity to the visual changes. We discover that human sensitivity is influenced by multiple factors, from low-level features such as illumination, texture, to high-level attributes like object sentiment and semantics. Based on our human data, we propose for the first time the concept of human sensitivity map. With the sensitivity map, we design a human-sensitivity-aware image perturbation model, which is able to modify the computational classification results of sensitive attributes while preserving the remaining attributes. Experiments on real world data demonstrate the superior performance of the proposed model on human-imperceptible privacy protection. Zhiqi Shen 0002, Shaojing Fan, Yongkang Wong, Tian-Tsong Ng, Mohan Kankanhalli |
ACM Multimedia | 3 |
| 2019 | Unsupervised Domain Adaptation for 3D Human Pose EstimationabstractTraining an accurate 3D human pose estimator often requires a large amount of 3D ground-truth data which is inefficient and costly to collect. Previous methods have either resorted to weakly supervised methods to reduce the demand of ground-truth data for training, or using synthetically-generated but photo-realistic samples to enlarge the training data pool. Nevertheless, the former methods mainly require either additional supervision, such as unpaired 3D ground-truth data, or the camera parameters in multiview settings. On the other hand, the latter methods require accurately textured models, illumination configurations and background which need careful engineering. To address these problems, we propose a domain adaptation framework with unsupervised knowledge transfer, which aims at leveraging the knowledge in multi-modality data of the easy-to-get synthetic depth datasets to better train a pose estimator on the real-world datasets. Specifically, the framework first trains two pose estimators on synthetically-generated depth images and human body segmentation masks with full supervision, while jointly learning a human body segmentation module from the predicted 2D poses. Subsequently, the learned pose estimator and the segmentation module are applied to the real-world dataset to unsupervisedly learn a new RGB image based 2D/3D human pose estimator. Here, the knowledge encoded in the supervised learning modules are used to regularize a pose estimator without ground-truth annotations. Comprehensive experiments demonstrate significant improvements over weakly supervised methods when no ground-truth annotations are available. Further experiments with ground-truth annotations show that the proposed framework can outperform state-of-the-art fully supervised methods. In addition, we conducted ablation studies to examine the impact of each loss term, as well as with different amount of supervisions signal. Xiheng Zhang, Yongkang Wong, Mohan Kankanhalli, Weidong Geng |
ACM Multimedia | 2 |
| 2019 | Explainable Video Action Reasoning via Prior Knowledge and State TransitionsabstractHuman action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel action reasoning framework that uses prior knowledge to explain semantic-level observations of video state changes. Our method takes advantage of both classical reasoning and modern deep learning approaches. Specifically, prior knowledge is defined as the information of a target video domain, including a set of objects, attributes and relationships in the target video domain, as well as relevant actions defined by the temporal attribute and relationship changes (i.e. state transitions). Given a video sequence, we first generate a scene graph on each frame to represent concerned objects, attributes and relationships. Then those scene graphs are associated by tracking objects across frames to form a spatio-temporal graph (also called video graph), which represents semantic-level video states. Finally, by sequentially examining each state transition in the video graph, our method can detect and explain how those actions are executed with prior knowledge, just like the logical manner of thinking by humans. Compared to previous works, the action reasoning results of our method can be explained by both logical rules and semantic-level observations of video content changes. Besides, the proposed method can be used to detect multiple concurrent actions with detailed information, such as who (particular objects), when (time), where (object locations) and how (what kind of changes). Experiments on a re-annotated dataset CAD-120 show the effectiveness of our method. Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli |
ACM Multimedia | 4 |
| 2019 | LSTM-based multi-label video event detection
Anan Liu, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli |
Multim. Tools Appl. | 3 |
| 2019 | A multi-stream convolutional neural network for sEMG-based gesture recognition in muscle-computer interface
Yongkang Wong, Yu Du 0016, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng |
Pattern Recognit. Lett. | 2 |
| 2019 | Dual-Stream Recurrent Neural Network for Video CaptioningabstractRecent progress in using recurrent neural networks (RNNs) for video description has attracted an increasing interest, due to its capability to encode a sequence of frames for caption generation. While existing methods have studied various features (e.g., CNN, 3D CNN, and semantic attributes) for visual encoding, the representation and fusion of heterogeneous information from multi-modal spaces have not fully explored. Consider that different modalities are often asynchronous, frame-level multi-modal fusion (e.g., concatenation and linear fusion) will negatively influence each modality. In this paper, we propose a dual-stream RNN (DS-RNN) framework to jointly discover and integrate the hidden states of both visual and semantic streams for video caption generation. First, an encoding RNN is used for each stream to flexibly exploit the hidden states of respective modality. Specifically, we proposed an attentive multi-grained encoder module to enhance the local feature learning with global semantics feature. Then, a dual-stream decoder is deployed to integrate the asynchronous yet complementary sequential hidden states from both streams for caption generation. Extensive experiments on three benchmark datasets, namely, MSVD, MSR-VTT, and MPII-MD, show that DS-RNN achieves competitive performance against the state-of-the-art. Additional ablation studies were conducted on various variants of the proposed DS-RNN. Ning Xu 0003, Anan Liu, Yongkang Wong, Yongdong Zhang 0001, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Multi-Modal and Multi-Domain Embedding Learning for Fashion Retrieval and AnalysisabstractBig data analytics has been revolutionizing the fashion industry in recent years. This is evidenced by the fact that popular fashion brands and designers have relied on big data analytics to trace fashion trends and predict market patterns. In this paper, we propose learning a common latent feature representation from heterogeneous fashion data. Specifically, we design a multi-modal and multi-domain embedding learning framework for fashion analysis and data retrieval. Unlike most of the existing multi-view embedding methods, which only consider the heterogeneous similarity constraint, our proposed framework jointly considers both the homogeneous and heterogeneous similarity constraints to capture cross-view similarity and preserve the similarity of the same view. The proposed framework is comprised of two projection steps. In the first projection, a quintuplet-based ranking loss is proposed for multi-domain fashion data to preserve the homogeneous similarity. In the second projection, a cross-view similarity ranking loss is designed for multi-modal fashion data to capture heterogeneous similarity. By utilizing the learned common latent feature representation, the distance between any vector pairs from same or different modalities can reflect its semantic similarity. Quantitative evaluation on a new large-scale dataset and a fashion analysis case study demonstrate the effectiveness of our proposed method. Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli |
IEEE Trans. Multim. | 2 |
| 2019 | A Multi-sensor Framework for Personal Presentation AnalyticsabstractPresentation has been an effective method for delivering information to an audience for many years. Over the past few decades, technological advancements have revolutionized the way humans deliver presentation. Conventionally, the quality of a presentation is usually evaluated through painstaking manual analysis with experts. Although the expert feedback is effective in assisting users to improve their presentation skills, manual evaluation suffers from high cost and is often not available to most individuals. In this work, we propose a novel multi-sensor self-quantification system for presentations, which is designed based on a new proposed assessment rubric. We present our analytics model with conventional ambient sensors (i.e., static cameras and Kinect sensor) and the emerging wearable egocentric sensors (i.e., Google Glass). In addition, we performed a cross-correlation analysis of speaker’s vocal behavior and body language. The proposed framework is evaluated on a new presentation dataset, namely, NUS Multi-Sensor Presentation dataset, which consists of 51 presentations covering a diverse range of topics. To validate the efficacy of the proposed system, we have conducted a series of user studies with the speakers and an interview with an English communication expert, which reveals positive and promising feedback. Tian Gan 0002, Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Unsupervised Learning of View-invariant Action RepresentationsabstractThe recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, label collection is an expensive and time-consuming process. In this work, we propose an unsupervised learning framework, which exploits unlabeled data to learn video representations. Different from previous works in video representation learning, our unsupervised learning task is to predict 3D motion in multiple target views using video representation from a source view. By learning to extrapolate cross-view motions, the representation can capture view-invariant motion dynamics which is discriminative for the action. In addition, we propose a view-adversarial training method to enhance learning of view-invariant features. We demonstrate the effectiveness of the learned representations for action recognition on multiple datasets. Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
NeurIPS | 2 |
| 2017 | Dual-Glance Model for Deciphering Social RelationshipsabstractSince the beginning of early civilizations, social relationships derived from each individual fundamentally form the basis of social structure in our daily life. In the computer vision literature, much progress has been made in scene understanding, such as object detection and scene parsing. Recent research focuses on the relationship between objects based on its functionality and geometrical relations. In this work, we aim to study the problem of social relationship recognition, in still images. We have proposed a dual-glance model for social relationship recognition, where the first glance fixates at the individual pair of interest and the second glance deploys attention mechanism to explore contextual cues. We have also collected a new large scale People in Social Context (PISC) dataset, which comprises of 22,670 images and 76,568 annotated samples from 9 types of social relationship. We provide benchmark results on the PISC dataset, and qualitatively demonstrate the efficacy of the proposed model. Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
ICCV | 2 |
| 2017 | Semi-Supervised Learning for Surface EMG-based Gesture RecognitionabstractConventionally, gesture recognition based on non-intrusive muscle-computer interfaces required a strongly-supervised learning algorithm and a large amount of labeled training signals of surface electromyography (sEMG). In this work, we show that temporal relationship of sEMG signals and data glove provides implicit supervisory signal for learning the gesture recognition model. To demonstrate this, we present a semi-supervised learning framework with a novel Siamese architecture for sEMG-based gesture recognition. Specifically, we employ auxiliary tasks to learn visual representation; predicting the temporal order of two consecutive sEMG frames; and, optionally, predicting the statistics of 3D hand pose with a sEMG frame. Experiments on the NinaPro, CapgMyo and csl-hdemg datasets validate the efficacy of our proposed approach, especially when the labeled samples are very scarce. Yu Du 0016, Yongkang Wong, Wenguang Jin, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng |
IJCAI | 2 |
| 2017 | Understanding Fashion Trends from Street Photos via Neighbor-Constrained Embedding LearningabstractDriven by the increasing popular image-dominated social networks, such as Instagram, Pinterest and Chictopica, sharing of daily-life street photos now plays an influential role in fashion adoption between fashion trend-setters and followers. In this work, we propose a deep learning based fine-grained embedding learning approach for street fashion analysis by leveraging user-generated street fashion data. Specifically, we present QuadNet, an effective CNN based image embedding network driven by both multi-task classification loss and neighbor-constrained similarity loss. The latter loss function is computed with a novel quadruplet loss function, which considers both hard and soft positive neighbors as well as a negative neighbor for each anchor image. The embedded feature learned from co-optimization is effective for both fine-grained classification task and image retrieval task. Quantitative evaluation on a newly collected large-scale multi-task street photo dataset shows that our QuadNet outperforms the state-of-the-art triplet network by a significant margin. In order to further evaluate the effectiveness of the learned embedding, we analyze and trace the fashion trends of New York City from 2011 to 2016. In our analysis, we are able to identify some short-term and long-term fashion styles. Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2017 | Attention Transfer from Web Images for Video RecognitionabstractTraining deep learning based video classifiers for action recognition requires a large amount of labeled videos. The labeling process is labor-intensive and time-consuming. On the other hand, large amount of weakly-labeled images are uploaded to the Internet by users everyday. To harness the rich and highly diverse set of Web images, a scalable approach is to crawl these images to train deep learning based classifier, such as Convolutional Neural Networks (CNN). However, due to the domain shift problem, the performance of Web images trained deep classifiers tend to degrade when directly deployed to videos. One way to address this problem is to fine-tune the trained models on videos, but sufficient amount of annotated videos are still required. In this work, we propose a novel approach to transfer knowledge from image domain to video domain. The proposed method can adapt to the target domain (i.e. video data) with limited amount of training data. Our method maps the video frames into a low-dimensional feature space using the class-discriminative spatial attention map for CNNs. We design a novel Siamese EnergyNet structure to learn energy functions on the attention maps by jointly optimizing two loss functions, such that the attention map corresponding to a ground truth concept would have higher energy. We conduct extensive experiments on two challenging video recognition datasets (i.e. TVHI and UCF101), and demonstrate the efficacy of our proposed method. Junnan Li 0001, Yongkang Wong, Qi Zhao 0001, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2017 | Multi-Camera Action Dataset for Cross-Camera Action Recognition BenchmarkingabstractAction recognition has received increasing attention from the computer vision and machine learning communities in the last decade. To enable the study of this problem, there exist a vast number of action datasets, which are recorded under controlled laboratory settings, real-world surveillance environments, or crawled from the Internet. Apart from the "in-the-wild" datasets, the training and test split of conventional datasets often possess similar environments conditions, which leads to close to perfect performance on constrained datasets. In this paper, we introduce a new dataset, namely Multi-Camera Action Dataset (MCAD), which is designed to evaluate the open view classification problem under the surveillance environment. In total, MCAD contains 14,298 action samples from 18 action categories, which are performed by 20 subjects and independently recorded with 5 cameras. Inspired by the well received evaluation approach on the LFW dataset, we designed a standard evaluation protocol and benchmarked MCAD under several scenarios. The benchmark shows that while an average of 85% accuracy is achieved under the closed-view scenario, the performance suffers from a significant drop under the cross-view scenario. In the worst case scenario, the performance of 10-fold cross validation drops from 87.0% to 47.4%. Wenhui Li 0001, Yongkang Wong, Anan Liu, Yang Li 0108, Yuting Su 0001, Mohan Kankanhalli |
WACV | 2 |
| 2017 | Hierarchical & multimodal video captioning: Discovering and transferring multimodal knowledge for vision to language
Anan Liu, Ning Xu 0003, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli |
Comput. Vis. Image Underst. | 3 |
| 2017 | Benchmarking a Multimodal and Multiview and Interactive Dataset for Human Action RecognitionabstractHuman action recognition is an active research area in both computer vision and machine learning communities. In the past decades, the machine learning problem has evolved from conventional single-view learning problem, to cross-view learning, cross-domain learning and multitask learning, where a large number of algorithms have been proposed in the literature. Despite having large number of action recognition datasets, most of them are designed for a subset of the four learning problems, where the comparisons between algorithms can further limited by variances within datasets, experimental configurations, and other factors. To the best of our knowledge, there exists no dataset that allows concurrent analysis on the four learning problems. In this paper, we introduce a novel multimodal and multiview and interactive (M2I) dataset, which is designed for the evaluation of human action recognition methods under all four scenarios. This dataset consists of 1760 action samples from 22 action categories, including nine person-person interactive actions and 13 person-object interactive actions. We systematically benchmark state-of-the-art approaches on M2I dataset on all four learning problems. Overall, we evaluated 13 approaches with nine popular feature and descriptor combinations. Our comprehensive analysis demonstrates that M2I dataset is challenging due to significant intraclass and view variations, and multiple similar action categories, as well as provides solid foundation for the evaluation of existing state-of-the-art algorithms. Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Cybern. | 5 |
| 2016 | Marker-Less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps
Yu Du 0016, Yongkang Wong, Feilin Han, Yilin Gui, Zhen Wang 0003, Mohan Kankanhalli, Weidong Geng |
ECCV (4) | 2 |
| 2016 | Towards protecting biometric templates without sacrificing performanceabstractThe ideal biometric template protection scheme possesses the properties of irreversibility, revocability, unlinkability, and good performance. These properties protect the security of the biometrics system as well as users' privacy. Practical systems, however, fall short of this ideal. In this paper, we present a novel protection scheme that achieves this ideal under the circumstance that a subject's token and his biometric template are not concurrently exposed. Moreover, our scheme can add template protection to any face verifier. We do this by rendering virtual faces, rather than by devising new biometric features, which is the more common approach. Experimental evaluations using two public face recognition systems show that accuracy is not adversely affected with our scheme. Jing Li 0050, Yongkang Wong, Terence Sim |
ICPR | 2 |
| 2016 | Multi-stream Deep Learning Framework for Automated Presentation AssessmentabstractPresentation is one of the most effective methods to disseminate information. Traditional methods to evaluate the quality of a presentation generally involves a human instructor, which is infeasible in many scenarios. Recent studies have focused on the automated assessment of presentations. A variety of systems have been developed that focus on analyzing various aspects of presentations. However, those systems are mainly limited by their performance, as they mostly adopt hand-crafted features and ad-hoc algorithms. In this work, we propose a multi-stream deep learning framework customized for presentation assessment. The framework uses Bidirectional Long Short-Term Memory with attention mechanism for temporal modeling, and fuses information from multiple modalities for the final decision. We also design a novel assessment rubric based on input from a domain expert. Experimental results on the NUS Multi-Sensor Presentation (NUSMAP) dataset show that the proposed framework is computationally efficient and achieves significant improvement in classification accuracy. Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli |
ISM | 2 |
| 2016 | Demo Paper: PreSense - An Assistive Presentation Self-Quantification SystemabstractThis technical demo presents PreSense, an Assistive Presentation Self-Quantification System. Oral presentation is traditionally evaluated by a human instructor, which is cost-ineffective and time-consuming. PreSense allows individuals to self-evaluate their presentation skills. The system is designed to receive multimodal inputs from multiple sources, including webcam, Kinect sensor and Google Glass. The multimodal data is processed by a deep assessment framework, which outputs the evaluation results based on a carefully designed assessment rubric. In order to systematically visualize the individual assessment results, we develop an interactive graphical interface and demonstrate its efficacy using presentation data from the NUS Multi-Sensor Presentation (NUSMAP) dataset. Junnan Li 0001, Yongkang Wong, Mohan Kankanhalli |
ISM | 2 |
| 2015 | Label Consistent Quadratic Surrogate model for visual saliency predictionabstractRecently, an increasing number of works have proposed to learn visual saliency by leveraging human fixations. However, the collection of human fixations is time consuming and the existing eye tracking datasets are generally small when compared with other domains. Thus, it contains a certain degree of dataset bias due to the large image variations (e.g., outdoor scenes vs. emotion-evoking images). In the learning based saliency prediction literature, most models are trained and evaluated within the same dataset and cross dataset validation is not yet a common practice. Instead of directly applying model learned from another dataset in cross dataset fashion, it is better to transfer the prior knowledge obtained from one dataset to improve the training and prediction on another. In addition, since new datasets are built and shared in the community from time to time, it would be good not to retrain the entire model when new data are added. To address these problems, we proposed a new learning based saliency model, namely Label Consistent Quadratic Surrogate algorithm, which employs an iterative online algorithm to learn a sparse dictionary with label consistent constraint. The advantages of the proposed model are three-folds: (1) the quadratic surrogate function guarantees convergence at each iteration, (2) the label consistent constraint enforces the predicted sparse code to be discriminative, and (3) the online properties enable the proposed algorithm to adapt existing model with new data without retraining. As shown in this work, the proposed saliency model achieves better performance than the state-of-the-art saliency models. Yan Luo 0002, Yongkang Wong, Qi Zhao 0001 |
CVPR | 2 |
| 2015 | Multi-sensor Self-Quantification of PresentationsabstractPresentations have been an effective means of delivering information to groups for ages. Over the past few decades, technological advancements have revolutionized the way humans deliver presentations. Despite that, the quality of presentations can be varied and affected by a variety of reasons. Conventional presentation evaluation usually requires painstaking manual analysis by experts. Although the expert feedback can definitely assist users in improving their presentation skills, manual evaluation suffers from high cost and is often not accessible to most people. In this work, we propose a novel multi-sensor self-quantification framework for presentations. Utilizing conventional ambient sensors (i.e., static cameras, Kinect sensor) and the emerging wearable egocentric sensors (i.e., Google Glass), we first analyze the efficacy of each type of sensor with various nonverbal assessment rubrics, which is followed by our proposed multi-sensor presentation analytics framework. The proposed framework is evaluated on a new presentation dataset, namely NUS Multi-Sensor Presentation (NUSMSP) dataset, which consists of 51 presentations covering a diverse set of topics. The dataset was recorded with ambient static cameras, Kinect sensor, and Google Glass. In addition to multi-sensor analytics, we have conducted a user study with the speakers to verify the effectiveness of our system generated analytics, which has received positive and promising feedback. Tian Gan 0002, Yongkang Wong, Bappaditya Mandal, Vijay Chandrasekhar 0001, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2015 | Multi-modal & Multi-view & Interactive Benchmark Dataset for Human Action RecognitionabstractHuman action recognition is one of the most active research areas in both computer vision and machine learning communities. Several methods for human action recognition have been proposed in the literature and promising results have been achieved on the popular datasets. However, the comparison of existing methods is often limited given the different datasets, experimental settings, feature representations, and so on. In particularly, there are no human action dataset that allow concurrent analysis on three popular scenarios, namely single view, cross view, and cross domain. In this paper, we introduce a Multi-modal & Multi-view & Interactive (M2I) dataset, which is designed for the evaluation of the performances of human action recognition under multi-view scenario. This dataset consists of 1760 action samples, including 9 person-person interaction actions and 13 person-object interaction actions. Moreover, we respectively evaluate three representative methods for the single-view, cross-view, and cross domain human action recognition on this dataset with the proposed evaluation protocol. It is experimentally demonstrated that this dataset is extremely challenging due to large intraclass variation, multiple similar actions, significant view difference. This benchmark can provide solid basis for the evaluation of this task and will benefit advancing related computer vision and machine learning research topics. Ning Xu 0003, Anan Liu, Weizhi Nie, Yongkang Wong, Fuwu Li, Yuting Su 0001 |
ACM Multimedia | 4 |
| 2015 | Multi-Camera SaliencyabstractA significant body of literature on saliency modeling predicts where humans look in a single image or video. Besides the scientific goal of understanding how information is fused from multiple visual sources to identify regions of interest in a holistic manner, there are tremendous engineering applications of multi-camera saliency due to the widespread of cameras. This paper proposes a principled framework to smoothly integrate visual information from multiple views to a global scene map, and to employ a saliency algorithm incorporating high-level features to identify the most important regions by fusing visual information. The proposed method has the following key distinguishing features compared with its counterparts: (1) the proposed saliency detection is global (salient regions from one local view may not be important in a global context), (2) it does not require special ways for camera deployment or overlapping field of view, and (3) the key saliency algorithm is effective in highlighting interesting object regions though not a single detector is used. Experiments on several data sets confirm the effectiveness of the proposed principled framework. Yan Luo 0002, Ming Jiang 0019, Yongkang Wong, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | View-invariant feature discovering for multi-camera human action recognitionabstractIntelligent video surveillance system is built to automatically detect events of interest, especially on object tracking and behavior understanding. In this paper, we focus on the task of human action recognition under surveillance environment, specifically in a multi-camera monitoring scene. Despite many approaches have achieved success in recognizing human action from video sequences, they are designed for single view and generally not robust against viewpoint invariant. Human action recognition across different views remains challenging due to the large variations from one view to another. We present a framework to solve the problem of transferring action models learned in one view (source view) to another view (target view). First, local space-time interest point feature and global shape-flow feature are extracted as low-level feature, followed by building the hybrid Bag-of-Words model for each action sequence. The data distribution of relevant actions from source view and target view are linked via a cross-view discriminative dictionary learning method. Through the view-adaptive dictionary pair learned by the method, the data from source and target view can be respectively mapped into a common space which is view-invariant. Furthermore, We extend our framework to transfer action models from multiple views to one view when there are multiple source views available. Experiments on the IXMAS human action dataset, which contains videos captured with five viewpoints, show the efficacy of our framework. Lekha Chaisorn, Yongkang Wong, Anan Liu, Yuting Su 0001, Mohan Kankanhalli |
MMSP | 3 |
| 2014 | Multi-view action recognition by cross-domain learningabstractThis paper proposes a novel multi-view human action recognition method by discovering and sharing common knowledge among different video sets captured in multiple viewpoints. To our knowledge, we are the first to treat a specific view as target domain and the others as source domains and consequently formulate the multi-view action recognition into the cross-domain learning framework. First, the classic bag-of-visual word framework is implemented for visual feature extraction in individual viewpoints. Then, we propose a cross-domain learning method with block-wise weighted kernel function matrix to highlight the saliency components and consequently augment the discriminative ability of the model. Extensive experiments are implemented on IXMAS, the popular multi-view action dataset. The experimental results demonstrate that the proposed method can consistently outperform the state of the arts. Weizhi Nie, Anan Liu, Yuting Su 0001, Lekha Chaisorn, Yongkang Wong, Mohan Kankanhalli |
MMSP | 6 |
| 2014 | Automatic classification of Human Epithelial type 2 cell Indirect Immunofluorescence images using Cell Pyramid Matching
Arnold Wiliem, Conrad Sanderson, Yongkang Wong, Peter Hobson, Rodney F. Minchin, Brian C. Lovell |
Pattern Recognit. | 3 |
| 2013 | Temporal encoded F-formation system for social interaction detectionabstractIn the context of a social gathering, such as a cocktail party, the memorable moments are generally captured by professional photographers or by the participants. The latter case is often undesirable because many participants would rather enjoy the event instead of being occupied by the photo-taking task. Motivated by this scenario, we propose the use of a set of cameras to automatically take photos. Instead of performing dense analysis on all cameras for photo capturing, we first detect the occurrence and location of social interactions via F-formation detection. In the sociology literature, F-formation is a concept used to define social interactions, where each detection only requires the spatial location and orientation of each participant. This information can be robustly obtained with additional Kinect depth sensors. In this paper, we propose an extended F-formation system for robust detection of interactions and interactants. The extended F-formation system employs a heat-map based feature representation for each individual, namely Interaction Space (IS), to model their location, orientation, and temporal information. Using the temporally encoded IS for each detected interactant, we propose a best-view camera selection framework to detect the corresponding best view camera for each detected social interaction. The extended F-formation system is evaluated with synthetic data on multiple scenarios. To demonstrate the effectiveness of the proposed system, we conducted a user study to compare our best view camera ranking with human's ranking using real-world data. Tian Gan 0002, Yongkang Wong, Daqing Zhang 0001, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2013 | Classification of Human Epithelial type 2 cell indirect immunofluoresence images via codebook based descriptorsabstractThe Anti-Nuclear Antibody (ANA) clinical pathology test is commonly used to identify the existence of various diseases. A hallmark method for identifying the presence of ANAs is the Indirect Immunofluorescence method on Human Epithelial (HEp-2) cells, due to its high sensitivity and the large range of antigens that can be detected. However, the method suffers from numerous shortcomings, such as being subjective as well as time and labour intensive. Computer Aided Diagnostic (CAD) systems have been developed to address these problems, which automatically classify a HEp-2 cell image into one of its known patterns (eg., speckled, homogeneous). Most of the existing CAD systems use handpicked features to represent a HEp-2 cell image, which may only work in limited scenarios. In this paper, we propose a cell classification system comprised of a dual-region codebook-based descriptor, combined with the Nearest Convex Hull Classifier. We evaluate the performance of several variants of the descriptor on two publicly available datasets: ICPR HEp-2 cell classification contest dataset and the new SNPHEp-2 dataset. To our knowledge, this is the first time codebook-based descriptors are applied and studied in this domain. Experiments show that the proposed system has consistent high performance and is more robust than two recent CAD systems. Arnold Wiliem, Yongkang Wong, Conrad Sanderson, Peter Hobson, Shaokang Chen, Brian C. Lovell |
WACV | 2 |
| 2012 | Combined Learning of Salient Local Descriptors and Distance Metrics for Image Set Face VerificationabstractIn contrast to comparing faces via single exemplars, matching sets of face images increases robustness and discrimination performance. Recent image set matching approaches typically measure similarities between subspaces or manifolds, while representing faces in a rigid and holistic manner. Such representations are easily affected by variations in terms of alignment, illumination, pose and expression. While local feature based representations are considerably more robust to such variations, they have received little attention within the image set matching area. We propose a novel image set matching technique, comprised of three aspects: (i) robust descriptors of face regions based on local features, partly inspired by the hierarchy in the human visual system, (ii) use of several subspace and exemplar metrics to compare corresponding face regions, (iii) jointly learning which regions are the most discriminative while finding the optimal mixing weights for combining metrics. Experiments on LFW, PIE and MOBIO face datasets show that the proposed algorithm obtains considerably better performance than several recent state of-the-art techniques, such as Local Principal Angle and the Kernel Affine Hull Method. Conrad Sanderson, Mehrtash Harandi, Yongkang Wong, Brian C. Lovell |
AVSS | 3 |
| 2012 | On robust biometric identity verification via sparse encoding of faces: Holistic vs local approachesabstractIn the field of face recognition, Sparse Representation (SR) has received considerable attention during the past few years. Most of the related literature focuses on holistic descriptors in closed-set identification applications. The underlying assumption in identification is that the gallery always has sufficient samples per subject to linearly reconstruct a query image. Unfortunately, such assumption is easily violated in the more challenging and realistic face verification scenario. A verification algorithm is required to determine if two faces (where one or both have not been seen before) belong to the same person, while explicitly taking into account the possibility of impostor attacks. In this paper, we first discuss why most of the SR literature is not applicable to verification problems. Motivated by the success of bag-of-words methods in the field of object recognition, which describe an image as a set of local patches or interest points, we then propose to tackle the verification problem by encoding each local face patch through SR. The locally encoded sparse vectors are pooled to form regional descriptors, where each descriptor covers a relatively large portion of the face. Experiments in various challenging conditions show that the proposed method achieves high and robust verification performance. Yongkang Wong, Mehrtash Harandi, Conrad Sanderson, Brian C. Lovell |
IJCNN | 1 |
| 2010 | Dynamic Amelioration of Resolution Mismatches for Local Feature Based Identity InferenceabstractWhile existing face recognition systems based on local features are robust to issues such as misalignment, they can exhibit accuracy degradation when comparing images of differing resolutions. This is common in surveillance environments where a gallery of high resolution mugshots is compared to low resolution CCTV probe images, or where the size of a given image is not a reliable indicator of the underlying resolution (e.g. poor optics). To alleviate this degradation, we propose a compensation framework which dynamically chooses the most appropriate face recognition system for a given pair of image resolutions. This framework applies a novel resolution detection method which does not rely on the size of the input images, but instead exploits the sensitivity of local features to resolution using a probabilistic multi-region histogram approach. Experiments on a resolution-modified version of the "Labeled Faces in the Wild" dataset show that the proposed resolution detector frontend obtains a 99% average accuracy in selecting the most appropriate face recognition system, resulting in higher overall face discrimination accuracy (across several resolutions) compared to the individual baseline face recognition systems. Yongkang Wong, Conrad Sanderson, Sandra Mau, Brian C. Lovell |
ICPR | 1 |
| 2009 | Regression Based Non-frontal Face Synthesis for Improved Identity Verification
Yongkang Wong, Conrad Sanderson, Brian C. Lovell |
CAIP | 1 |