John See

dblp:36/4809 · DBLP profile ↗
← Back
94ranked-venue papers
6as first author
50since 2021 · last 2026
0000-0003-3005-4109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 3 first-author · 33 since 2021Artificial intelligence and machine learning · 41 · 5 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unleashing Semantic and Geometric Priors for 3D Scene Completion
abstract
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric priors, which forces the model to make a trade-off between conflicting demands and limits its overall performance. To tackle these challenges, we propose FoundationSSC, a novel framework that performs dual decoupling at both the source and pathway levels. At the source level, we introduce a foundation encoder that provides rich semantic feature priors for the semantic branch and high-fidelity stereo cost volumes for the geometric branch. At the pathway level, these priors are refined through specialised, decoupled pathways, yielding superior semantic context and depth distributions. Our dual-decoupling design produces disentangled and refined inputs, which are then utilised by a hybrid view transformation to generate complementary 3D features. Additionally, we introduce a novel Axis-Aware Fusion (AAF) module that addresses the often-overlooked challenge of fusing these features by anisotropically merging them into a unified representation. Extensive experiments demonstrate the advantages of FoundationSSC, achieving simultaneous improvements in both semantic and geometric metrics, surpassing prior bests by +0.23 mIoU and +2.03 IoU on SemanticKITTI. Additionally, we achieve state-of-the-art performance on SSCBench-KITTI-360, with 21.78 mIoU and 48.61 IoU.
Shiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers, John See
AAAI5
2026 MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering
Xinqi Fan, Jingting Li 0001, John See, Moi Hoon Yap, Adrian K. Davison
FG3
2026 MoiréNet: Leveraging Directional Priors for Compact Dual-Domain Image Demoiréing
abstract
Digital images serve as the fundamental carrier for information exchange within multimedia ecosystems. However, the ubiquitous practice of screen recapturing often introduces complex moiré patterns due to spectral aliasing, which severely degrades the visual quality and impedes downstream multimedia analysis tasks. In this paper, we propose MoiréNet, a compact convolutional neural framework designed for effective moiré removal, thereby synergistically restoring high-fidelity content. To address the anisotropic and multi-scale nature of these artifacts, we introduce two novel modules: the Directional Frequency-Spatial Encoder (DFSE), which explicitly discerns moiré orientation via directional difference convolutions, and the Frequency-Spatial Adaptive Selector (FSAS), which enables feature-adaptive artifact suppression across dual domains. Extensive experiments demonstrate that MoiréNet achieves state-of-the-art performance on public and widely used datasets while being highly parameter-efficient. With only 5.513M parameters, representing a 48% reduction compared to ESDNet-L, MoiréNet combines superior restoration quality with parameter efficiency for storage-constrained multimedia applications.
Shuwei Guo, Simin Luan, John See, Zeyd Boukhers, Miho Ohsaki, Kimiaki Shirahama
ICMR3
2026 TrinitySeg: A Benchmark and Baseline for Monochrome NIR Smoke Segmentation
abstract
Monochrome Near-Infrared (NIR) surveillance is widely used for low-light monitoring, yet pixel-level smoke segmentation remains underexplored due to scarce annotated data and strong appearance ambiguity. We introduce NIR-Smoke, a monochrome NIR smoke segmentation benchmark containing real frame-wise annotations and synthetic composites, and provide TrinitySeg as a strong baseline. Specifically, TrinitySeg integrates three complementary cues by capturing motion via lightweight frame differences, structure via multi-scale deep supervision with an auxiliary boundary head, and frequency via a compact Fourier-domain filtering module. We further study synthetic-to-real mixing for training. On NIR-Smoke, TrinitySeg achieves 85.64 ± 0.10 mIoU and 72.57 ± 0.21 Smoke IoU, outperforming representative baselines while retaining near-baseline inference cost.
Haowen Hua, Zeyi Shao, Shuwei Guo, John See, Zeyd Boukhers, Miho Ohsaki, Kimiaki Shirahama
ICMR4
2026 Sparse4D: Sparse-Based End-to-End Multi-Sensor Temporal Perception
abstract
In the field of autonomous driving, the structural design of perception models is of paramount importance. Unlike mainstream BEV-based algorithms, we propose a novel end-to-end sparse perception framework, Sparse4D, to achieve better performance and higher efficiency. Starting from the sparsity of perception results, we define an instance that decouples implicit features and explicit anchors, using the instance as the core for feature fusion to accomplish perception tasks. For spatial modelling, we develop a novel operator called deformable aggregation, which enables the transfer of information from the dense image feature to the sparse instance. For temporal modelling, we design a recurrent instance feature propagation structure, which not only realizes long-term feature fusion but also ensures the computational efficiency of the temporal module. Lastly, we explored the performance of Sparse4D in multi-object tracking tasks and proposed a minimalist joint detection and tracking model. We conduct extensive experimental validation of Sparse4D on the nuScenes benchmark. Sparse4D achieved state-of-the-art performance in multi-camera 3D detection and tracking tasks, and it also outperformed other algorithms in terms of training and inference efficiency. Furthermore, we extended Sparse4D to a multi-modal model, achieving excellent performance and better generalization.
Xuewu Lin, Zixiang Pei, Lichao Huang, Chenyao Yu, John See, Zhizhong Su
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Few-Shot Action Recognition via Intra- and Inter-Video Information Maximization
abstract
Current few-shot action recognition involves two primary sources of information for classification: (1) intra-video information, determined by frame content within a single video clip, and (2) inter-video information, measured by relationships (e.g., feature similarity) among videos. However, existing methods inadequately exploit these two information sources. In terms of intra-video information, current sampling operations for input videos may omit critical action information, reducing the utilization efficiency of video data. For the inter-video information, the action misalignment among videos makes it challenging to calculate precise relationships. Moreover, how to jointly consider both inter- and intra-video information remains under-explored for few-shot action recognition. To this end, we propose a novel framework, Video Information Maximization (VIM), for few-shot video action recognition. VIM is equipped with an adaptive spatial-temporal video sampler and a spatial-temporal action alignment model to maximize intra- and inter-video information, respectively. The video sampler adaptively selects important frames and amplifies critical spatial regions for each input video based on the task at hand. This preserves and emphasizes informative parts of video clips while eliminating interference at the data level. The alignment model performs temporal and spatial action alignment sequentially at the feature level, leading to more precise measurements of inter-video similarity. Finally, based on the mutual information measurement, we introduce a new training objective into few-shot learning, which provides explicit guidance in jointly maximizing intra- and inter-video information in our VIM. Extensive experimental results on public datasets for few-shot action recognition demonstrate the effectiveness of our framework.
Huabin Liu 0001, Tieyuan Chen, Yuxi Li 0009, Shuyuan Li, John See, Weiyao Lin
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
abstract
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In recent years, substantial advancements have been made in the areas of ME recognition, spotting, and generation. However, conventional approaches that treat spotting and recognition as separate tasks are suboptimal, particularly for analyzing long-duration videos in realistic settings. Concurrently, the emergence of multimodal large language models (MLLMs) and large vision-language models (LVLMs) offers promising new avenues for enhancing ME analysis through their powerful multimodal reasoning capabilities. The ME grand challenge (MEGC) 2025 introduces two tasks that reflect these evolving research directions: (1) ME spot-then-recognize (ME-STR), which integrates ME spotting and subsequent recognition in a unified sequential pipeline; and (2) ME visual question answering (ME-VQA), which explores ME understanding through visual question answering, leveraging MLLMs or LVLMs to address diverse question types related to MEs. All participating algorithms are required to run on this test set and submit their results on a leaderboard. More details are available at https://megc2025.github.io.
Xinqi Fan, Jingting Li 0001, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaopeng Hong, Adrian K. Davison
ACM Multimedia3
2025 Encoding Music Score Data for Emotional Expression Prediction in Western Classical Music Pieces
abstract
Music’s ability to evoke emotions has long been recognised, with specific musical features shaping listeners’ perceptions. This study investigates the relationship between symbolic musical features and perceived emotional responses (valence and arousal) in Western classical music. Focusing on 72 preludes by J.S. Bach (Baroque period), F. Chopin (Romantic period), and D. Shostakovich (Modernist period), we develop an approach to predict emotional ratings directly from symbolic music that reduces interpretation bias, while examining how compositional structures shape emotional expression. It is observed that a GRU model integrated with Dyna-Octuple feature set achieves the best trade-off between accuracy and efficiency. Segment-level analysis reveals that emotional predictions remain stable across beginning, middle, and end sections, while period-based comparisons uncover distinctive stylistic patterns: Bach’s balanced control, Chopin’s expressive emotional intensity, and Shostakovich’s contrasting emotional states. This preliminary study demonstrates how symbolic features in Western classical music can relate well to emotional responses, as well as providing insights into stylistic differences across different periods in history.
Wei En Sui, Zhi Lin Chong, John See
MMAsia3
2025 Multimodal Engagement Prediction in Human-Robot Interaction Using Transformer Neural Networks
Jia Yap Lim, John See, Christian Dondrup
MMM (5)2
2025 Low-rank Winograd transformation for 3D convolutional neural networks
Ziran Qin, Mingbao Lin, Huabin Liu 0001, John See, Gui Zou, Weiyao Lin
Sci. China Inf. Sci.4
2025 Review of state-of-the-art surface defect detection on wind turbine blades through aerial imagery: Challenges and recommendations
Imad Gohar, Weng Kean Yew, Abderrahim Halimi, John See
Eng. Appl. Artif. Intell.4
2025 BladeView: Toward Automatic Wind Turbine Inspection With Unmanned Aerial Vehicle
abstract
This paper presents a fully automatic method, BladeView, for drone-based wind turbine blade inspection using an Unmanned Aerial Vehicle (UAV). With the need for highly efficient blade inspection coupled with the rapid increase of wind turbines, existing methods provide limited automation on wind turbine parameter estimation, full blade coverage, and safety control. We introduce an Automatic Parameter Calculation (APC) algorithm and an Automatic Flight System (AFS) in BladeView to compute wind turbine parameters and inspection paths, respectively. Leveraging triangulation and linear fitting integration techniques, the APC automatically calculates the wind turbine parameters and estimates the relative angle and position between a drone and the turbine. Furthermore, with dynamic path finding and B-spline optimization, the AFS plans a path covering 3 blades within specified flight corridors, in compliance with the turbine parameters obtained from APC. Thus, the proposed BladeView can properly ensure an inspection’s automation, coverage, safety, and smoothness. The efficiency and usability of BladeView are validated through 100,000 flight simulations in the Gazebo simulation environment and 9,239 field runs at various wind farms, including offshore, near-shore, deserts, mountainous areas, farmlands, and suburbs.Note to Practitioners—The proposed BladeView is distinguished in three aspects: (1) It automatically adapts wind turbines with varying geometric properties and physical locations relative to the take-off point. (2) It dramatically improves the quality of collected data with optimal UAV speed and flight corridors. (3) It thoroughly covers all three blades of a Horizontal-Axis Wind Turbine (HAWT), including regions where defects frequently occur. Thus, BladeView is more efficient and robust than existing UAV-based methods for blade inspection, with only around 25 minutes per HAWT. Moreover, it does not require experienced pilots to fly the UAV and manual interventions are rarely needed. Extensive simulation and real-world experiments demonstrate the efficiency and usability of BladeView in various on- and offshore wind farms.
Huan Zhou 0002, Yan Ke, Marcin Grzegorzek, Zeyd Boukhers, John See
IEEE Trans Autom. Sci. Eng.9
2025 CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
abstract
Continual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced sequentially. For video data, the task becomes more complex than image data because it requires learning and preserving both spatial appearance and temporal action involvement. To address this challenge, we propose a novel exemplar-free framework that equips separate spatiotemporal adapters to learn new class patterns, accommodating the incremental information representation requirements unique to each class. While separate adapters are proven to mitigate forgetting and fit unique requirements, naively applying them hinders the intrinsic connection between spatial and temporal information increments, affecting the efficiency of representing newly learned class information. Motivated by this, we introduce two key innovations from a causal perspective. First, a causal distillation module is devised to maintain the relation between spatial-temporal knowledge for a more efficient representation. Second, a causal compensation mechanism is proposed to reduce the conflicts during increment and memorization between different types of information. Extensive experiments conducted on benchmark datasets demonstrate that our framework can achieve new state-of-the-art results, surpassing current example-based methods by 4.2% in accuracy on average. The codes are accessible in https://github.com/tychen-SJTU/CSTA.
Tieyuan Chen, Huabin Liu 0001, Chern Hong Lim, John See, Xing Gao 0005, Junhui Hou, Weiyao Lin
IEEE Trans. Circuits Syst. Video Technol.4
2024 MEGC2024: ACM Multimedia 2024 Facial Micro-Expression Grand Challenge
abstract
Facial micro-expressions (MEs) are involuntary spontaneous movements of the face that typically appear in high-stakes situations where a person attempts to conceal a certain emotion from being known. A decade after the inception of the widely used CASME II and SMIC datasets, research in computational analysis of MEs has now advanced toward new pathways, exploring problems crucial to model generalization and real-world practicality. It is often challenging to design robust algorithms or models for spotting micro-expressions due to the high variability across diverse cultural backgrounds. Also, treating spotting and recognition as separate tasks is undesirable when handling long-spanning videos under realistic settings. This Grand Challenge comprises two distinct tracks: the Cross-Cultural Spotting (CCS) track, and the Spot-Then-Recognize (STR) track. All participating solutions submitted their results to a leaderboard, and several submissions performed well surpassing their respective baseline results. More details are available at: https://megc2024.github.io.
John See, Jingting Li 0001, Adrian K. Davison, Gen-Bing Liong, Moi Hoon Yap, Wen-Huang Cheng, Xiaopeng Hong
ACM Multimedia1
2024 Skeleton Ground Truth Extraction: Methodology, Annotation Tool and Benchmarks
abstract
Abstract Skeleton Ground Truth (GT) is critical to the success of supervised skeleton extraction methods, especially with the popularity of deep learning techniques. Furthermore, we see skeleton GTs used not only for training skeleton detectors with Convolutional Neural Networks (CNN), but also for evaluating skeleton-related pruning and matching algorithms. However, most existing shape and image datasets suffer from the lack of skeleton GT and inconsistency of GT standards. As a result, it is difficult to evaluate and reproduce CNN-based skeleton detectors and algorithms on a fair basis. In this paper, we present a heuristic strategy for object skeleton GT extraction in binary shapes and natural images. Our strategy is built on an extended theory of diagnosticity hypothesis, which enables encoding human-in-the-loop GT extraction based on clues from the target’s context, simplicity, and completeness. Using this strategy, we developed a tool, SkeView, to generate skeleton GT of 17 existing shape and image datasets. The GTs are then structurally evaluated with representative methods to build viable baselines for fair comparisons. Experiments demonstrate that GTs generated by our strategy yield promising quality with respect to standard consistency, and also provide a balance between simplicity and completeness.
Bipin Indurkhya, John See, Yan Ke, Zeyd Boukhers, Marcin Grzegorzek
Int. J. Comput. Vis.3
2024 SFAMNet: A scene flow attention-based micro-expression network
Gen-Bing Liong, Sze-Teng Liong, Chee Seng Chan, John See
Neurocomputing4
2024 Scene Graph Lossless Compression with Adaptive Prediction for Objects and Relations
abstract
The scene graph is a novel data structure describing objects and their pairwise relationship within image scenes. As the size of scene graphs in vision and multimedia applications increases, the need for lossless storage and transmission of such data becomes more critical. However, the compression of scene graphs is less studied because of the complicated data structures involved and complex distributions. Existing solutions usually involve general-purpose compressors or graph structure compression methods, which are weak at reducing the redundancy in scene graph data. This article introduces a novel lossless compression framework with adaptive predictors for the joint compression of objects and relations in scene graph data. The proposed framework comprises a unified prior extractor and specialized element predictors to adapt to different data elements. Furthermore, to exploit the context information within and between graph elements, Graph Context Convolution is proposed to support different graph context modeling schemes for different graph elements. Finally, an overarching framework incorporates the learned distribution model to predict numerical data under complicated conditional constraints. Experiments conducted on labeled or generated scene graphs demonstrate the effectiveness of the proposed framework for scene graph lossless compression.
Weiyao Lin, Wenrui Dai, Huabin Liu 0001, John See, Hongkai Xiong
ACM Trans. Multim. Comput. Commun. Appl.5
2024 A Unified Framework for Jointly Compressing Visual and Semantic Data
abstract
The rapid advancement of multimedia and imaging technologies has resulted in increasingly diverse visual and semantic data. A large range of applications such as remote-assisted driving requires the amalgamated storage and transmission of various visual and semantic data. However, existing works suffer from the limitation of insufficiently exploiting the redundancy between different types of data. In this article, we propose a unified framework to jointly compress a diverse spectrum of visual and semantic data, including images, point clouds, segmentation maps, object attributes, and relations. We develop a unifying process that embeds the representations of these data into a joint embedding graph according to their categories, which enables flexible handling of joint compression tasks for various visual and semantic data. To fully leverage the redundancy between different data types, we further introduce an embedding-based adaptive joint encoding process and a Semantic Adaptation Module to efficiently encode diverse data based on the learned embeddings in the joint embedding graph. Experiments on the Cityscapes, MSCOCO, and KITTI datasets demonstrate the superiority of our framework, highlighting promising steps toward scalable multimedia processing.
Shizhan Liu, Weiyao Lin, Yihang Chen 0002, Wenrui Dai, John See, Hongkai Xiong
ACM Trans. Multim. Comput. Commun. Appl.6
2023 What Modality Matters? Exploiting Highly Relevant Features for Video Advertisement Insertion
abstract
Video advertising is a thriving industry that has recently turned its attention to the use of intelligent algorithms for automating tasks. In advertisement insertion, the integration of contextual relevance is essential in influencing the viewer’s experience. Despite the wide spectrum of audio-visual semantic modalities available, there is a lack of research that analyzes their individual and complementary strengths in a systematic manner. In this paper, we propose an ad-insertion framework that maximizes the contextual relevance between advertisement and content video by employing high-level multi-modal semantic features. Prediction vectors are derived via clip-level and image-level extractors, which are then matched accordingly to yield relevance scores. We also established a new user study methodology that produces gold standard annotations based on multiple expert selections. By comprehensive human-centered approaches and analysis, we demonstrate that automatic ad-insertion can be improved by exploiting effective combinations of semantic modalities.
Onn Keat Chong, Hui-Ngo Goh, John See
ICIP3
2023 Context-Aware Multi-Stream Networks for Dimensional Emotion Prediction in Images
abstract
Teaching machines to comprehend the nuances of emotion from photographs is a particularly challenging task. Emotion perception— naturally a subjective problem, is often simplified for computational purposes into categorical states or valence-arousal dimensional space, the latter being a lesser-explored problem in the literature. This paper proposes a multi-stream context-aware neural network model for dimensional emotion prediction in images. Models were trained using a set of object and scene data along with deep features for valence, arousal, and dominance estimation. Experimental evaluation on a large-scale image emotion dataset demonstrates the viability of our proposed approach. Our analysis postulates that the understanding of the depicted object in an image is vital for successful predictions whilst relying on scene information can lead to somewhat confounding effects.
Sidharrth Nagappan, Jia Qi Tan, Lai-Kuan Wong, John See
ICIP4
2023 MEGC2023: ACM Multimedia 2023 ME Grand Challenge
abstract
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. Unfortunately, the small sample problem severely limits the automation of ME analysis. Furthermore, due to the weak and transient nature of MEs, it is difficult for models to distinguish it from other types of facial actions. Therefore, ME in long videos is a challenging task, and the current performance cannot meet the practical application requirements. Addressing these issues, this challenge focuses on ME and the macro-expression (MaE) spotting task. This year, in order to evaluate algorithms' performance more fairly, based on CAS(ME)2, SAMM Long Videos, SMIC-E-long, CAS(ME)3 and 4DME, we build an unseen cross-cultural long-video test set. All participating algorithms are required to run on this test set and submit their results on a leaderboard with a baseline result.
Adrian K. Davison, Jingting Li 0001, Moi Hoon Yap, John See, Wen-Huang Cheng, Xiaopeng Hong
ACM Multimedia4
2023 FME '23: 3rd Facial Micro-Expression Workshop
abstract
Micro-expressions are facial movements that are extremely short and not easily detected, which often reflect the genuine emotions of individuals. Micro-expressions are important cues for understanding real human emotions and can be used for non-contact, non-perceptual deception detection, or abnormal emotion recognition. It has broad application prospects in national security, judicial practice, health prevention, and clinical practice. However, micro-expression feature extraction and learning are highly challenging because they are typically short in duration, low intensity, and have local facial asymmetry. In addition, the intelligent micro-expression analysis combined with deep learning technology is also plagued by the problem of relatively small data samples. Not only is micro-expression elicitation very difficult, micro-expression annotation is also very time-consuming and laborious. More importantly, the micro-expression generation mechanism is not yet clear, which shackles the application of micro-expressions in real scenarios. FME'23 is the inaugural workshop in this area of research, with the aim of promoting interactions between researchers and scholars from within this niche area of research. This year we hope to discuss the growing ethical conversations when using face data, and how we can come to a consensus on micro-expression standards within affective computing.
Adrian K. Davison, Jingting Li 0001, Moi Hoon Yap, John See, Wen-Huang Cheng, Xiaopeng Hong
ACM Multimedia4
2023 Editorial for pattern recognition letters special issue on face-based emotion understanding
Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong
Pattern Recognit. Lett.4
2023 Spot-then-Recognize: A Micro-Expression Analysis Network for Seamless Evaluation of Long Videos
Gen-Bing Liong, John See, Chee Seng Chan
Signal Process. Image Commun.2
2023 ERNet: An Efficient and Reliable Human-Object Interaction Detection Network
abstract
Human-Object Interaction (HOI) detection recognizes how persons interact with objects, which is advantageous in autonomous systems such as self-driving vehicles and collaborative robots. However, current HOI detectors are often plagued by model inefficiency and unreliability when making a prediction, which consequently limits its potential for real-world scenarios. In this paper, we address these challenges by proposing ERNet, an end-to-end trainable convolutional-transformer network for HOI detection. The proposed model employs an efficient multi-scale deformable attention to effectively capture vital HOI features. We also put forward a novel detection attention module to adaptively generate semantically rich instance and interaction tokens. These tokens undergo pre-emptive detections to produce initial region and vector proposals that also serve as queries which enhances the feature refinement process in the transformer decoders. Several impactful enhancements are also applied to improve the HOI representation learning. Additionally, we utilize a predictive uncertainty estimation framework in the instance and interaction classification heads to quantify the uncertainty behind each prediction. By doing so, we can accurately and reliably predict HOIs even under challenging scenarios. Experiment results on the HICO-Det, V-COCO, and HOI-A datasets demonstrate that the proposed model achieves state-of-the-art performance in detection accuracy and training efficiency. Codes are publicly available at https://github.com/Monash-CyPhi-AI-Research-Lab/ernet.
JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Koksheik Wong, John See, Massimo Tistarelli
IEEE Trans. Image Process.5
2023 Doing More With Moiré Pattern Detection in Digital Photos
abstract
Detecting moiré patterns in digital photographs is meaningful as it provides priors towards image quality evaluation and demoiréing tasks. In this paper, we present a simple yet efficient framework to extract moiré edge maps from images with moiré patterns. The framework includes a strategy for training triplet (natural image, moiré layer, and their synthetic mixture) generation, and a Moiré Pattern Detection Neural Network (MoireDet) for moiré edge map estimation. This strategy ensures consistent pixel-level alignments during training, accommodating characteristics of a diverse set of camera-captured screen images and real-world moiré patterns from natural images. The design of three encoders in MoireDet exploits both high-level contextual and low-level structural features of various moiré patterns. Through comprehensive experiments, we demonstrate the advantages of MoireDet: better identification precision of moiré images on two datasets, and a marked improvement over state-of-the-art demoiréing methods.
Yan Ke, Marcin Grzegorzek, John See
IEEE Trans. Image Process.6
2023 FatigueView: A Multi-Camera Video Dataset for Vision-Based Drowsiness Detection
abstract
Although vision-based drowsiness detection approaches have achieved great success on empirically organized datasets, it remains far from being satisfactory for deployment in practice. One crucial issue lies in the scarcity and lack of datasets that represent the actual challenges in real-world applications, e.g. tremendous variation and aggregation of visual signs, challenges brought on by different camera positions and camera types. To promote research in this field, we introduce a new large-scale dataset, FatigueView, that is collected by both RGB and infrared (IR) cameras from five different positions. It contains real sleepy driving videos and various visual signs of drowsiness from subtle to obvious, e.g. with 17,403 different yawning sets totaling more than 124 million frames, far more than recent actively used datasets. We also provide hierarchical annotations for each video, ranging from spatial face landmarks and visual signs to temporal drowsiness locations and levels to meet different research requirements. We structurally evaluate representative methods to build viable baselines. With FatigueView, we would like to encourage the community to adapt computer vision models to address practical real-world concerns, particularly the challenges posed by this dataset.
John See
IEEE Trans. Intell. Transp. Syst.4
2023 Spatio-Temporal Point Process for Multiple Object Tracking
abstract
Multiple object tracking (MOT) focuses on modeling the relationship of detected objects among consecutive frames and merge them into different trajectories. MOT remains a challenging task as noisy and confusing detection results often hinder the final performance. Furthermore, most existing research are focusing on improving detection algorithms and association strategies. As such, we propose a novel framework that can effectively predict and mask-out the noisy and confusing detection results before associating the objects into trajectories. In particular, we formulate such "bad" detection results as a sequence of events and adopt the spatio-temporal point process to model such events. Traditionally, the occurrence rate in a point process is characterized by an explicitly defined intensity function, which depends on the prior knowledge of some specific tasks. Thus, designing a proper model is expensive and time-consuming, with also limited ability to generalize well. To tackle this problem, we adopt the convolutional recurrent neural network (conv-RNN) to instantiate the point process, where its intensity function is automatically modeled by the training data. Furthermore, we show that our method captures both temporal and spatial evolution, which is essential in modeling events for MOT. Experimental results demonstrate notable improvements in addressing noisy and confusing detection results in MOT data sets. An improved state-of-the-art performance is achieved by incorporating our baseline MOT algorithm with the spatio-temporal point process model.
Tao Wang 0002, Kean Chen, Weiyao Lin, John See, Zenghui Zhang, Xia Jia
IEEE Trans. Neural Networks Learn. Syst.4
2022 TA2N: Two-Stage Action Alignment Network for Few-Shot Action Recognition
abstract
Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this similarity is not ideal since different action instances may show distinctive temporal distribution, resulting in severe misalignment issues across query and support videos. In this paper, we arrest this problem from two distinct aspects -- action duration misalignment and action evolution misalignment. We address them sequentially through a Two-stage Action Alignment Network (TA2N). The first stage locates the action by learning a temporal affine transform, which warps each video feature to its action duration while dismissing the action-irrelevant feature (e.g. background). Next, the second stage coordinates query feature to match the spatial-temporal action evolution of support by performing temporally rearrange and spatially offset prediction. Extensive experiments on benchmark datasets show the potential of the proposed method in achieving state-of-the-art performance for few-shot action recognition.
Shuyuan Li, Huabin Liu 0001, Rui Qian 0001, Yuxi Li 0009, John See, Mengjuan Fei, Xiaoyuan Yu, Weiyao Lin
AAAI5
2022 Speed up Object Detection on Gigapixel-level Images with Patch Arrangement
abstract
With the appearance of super high-resolution (e.g., gigapixel-level) images, performing efficient object detection on such images becomes an important issue. Most ex-isting works for efficient object detection on high-resolution images focus on generating local patches where objects may exist, and then every patch is detected independently. How-ever, when the image resolution reaches gigapixel-level, they will suffer from a huge time cost for detecting numerous patches. Different from them, we devise a novel patch ar-rangement frameworkfor fast object detection on gigapixel-level images. Under this framework, a Patch Arrangement Network (PAN) is proposed to accelerate the detection by determining which patches could be packed together into a compact canvas. Specifically, PAN consists of (1) a Patch Filter Module (PFM) (2) a Patch Packing Module (PPM). PFM filters patch candidates by learning to select patches between two granularities. Subsequently, from the remaining patches, PPM determines how to pack these patches to-gether into a smaller number of canvases. Meanwhile, it generates an ideal layout of patches on canvas. These can-vases are fed to the detector to get final results. Experiments show that our method could improve the inference speed on gigapixel-level images by 5 x while maintaining great performance.
Huabin Liu 0001, John See, Aixin Zhang, Weiyao Lin
CVPR4
2022 Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition
abstract
A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algorithms at the feature level while little attention is paid to processing input video data. Moreover, existing frame sampling strategies may omit critical action information in temporal and spatial dimensions, which further impacts video utilization efficiency. In this paper, we propose a novel video frame sampler for few-shot action recognition to address this issue, where task-specific spatial-temporal frame sampling is achieved via a temporal selector (TS) and a spatial amplifier (SA). Specifically, our sampler first scans the whole video at a small computational cost to obtain a global perception of video frames. The TS plays its role in selecting top-T frames that contribute most significantly and subsequently. The SA emphasizes the discriminative information of each frame by amplifying critical regions with the guidance of saliency maps. We further adopt task-adaptive learning to dynamically adjust the sampling strategy according to the episode task at hand. Both the implementations of TS and SA are differentiable for end-to-end optimization, facilitating seamless integration of our proposed sampler with most few-shot action recognition methods. Extensive experiments show a significant boost in the performances on various benchmarks including long-term videos.
Huabin Liu 0001, Weixian Lv, John See, Weiyao Lin
ACM Multimedia3
2022 FME '22: 2nd Workshop on Facial Micro-Expression: Advanced Techniques for Multi-Modal Facial Expression Analysis
abstract
Micro-expressions are facial movements that are extremely short and not easily detected, which often reflect the genuine emotions of individuals. Micro-expressions are important cues for understanding real human emotions and can be used for non-contact non-perceptual deception detection, or abnormal emotion recognition. It has broad application prospects in national security, judicial practice, health prevention, clinical practice, etc. However, micro-expression feature extraction and learning are highly challenging because micro-expressions have the characteristics of short duration, low intensity, and local asymmetry. In addition, the intelligent micro-expression analysis combined with deep learning technology is also plagued by the problem of small samples. Not only is micro-expression elicitation very difficult, micro-expression annotation is also very time-consuming and laborious. More importantly, the micro-expression generation mechanism is not yet clear, which shackles the application of micro-expressions in real scenarios. FME'22 is the inaugural workshop in this area of research, with the aim of promoting interactions between researchers and scholars from within this niche area of research and also including those from broader, general areas of expression and psychology research. The complete FME'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552465.
Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong
ACM Multimedia4
2022 MEGC2022: ACM Multimedia 2022 Micro-Expression Grand Challenge
abstract
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. Unfortunately, the small sample problem severely limits the automation of ME analysis. Furthermore, due to the brief and subtle nature of ME, ME spotting is a challenging task, and the performance is still not satisfactory yet. This challenge focuses on two tasks, i.e., the micro- and macro-expression spotting task, and the ME Generation task.
Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong, Adrian K. Davison, Yante Li, Zizhao Dong
ACM Multimedia4
2022 BlumNet: Graph Component Detection for Object Skeleton Extraction
abstract
In this paper, we present a simple yet efficient framework, BlumNet, for extracting object skeletons in natural images and binary shapes. With the need for highly reliable skeletons in various multimedia applications, the proposed BlumNet is distinguished in three aspects: (1) The inception of graph decomposition and reconstruction strategies further simplifies the skeleton extraction task into a graph component detection problem, which significantly improves the accuracy and robustness of extracted skeletons. (2) The intuitive representation of each skeleton branch with multiple structured and overlapping line segments can effectively prevent the skeleton branch vanishing problem. (3) In comparison to traditional skeleton heatmaps, our approach directly outputs skeleton graphs, which is more feasible for real-world applications. Through comprehensive experiments, we demonstrate the advantages of BlumNet: significantly higher accuracy than the state-of-the-art AdaLSN (0.826 vs. 0.786) on the SK1491 dataset, a marked improvement in robustness on mixed object deformations, and also a state-of-the-art performance on binary shape datasets (e.g. 0.893 on the MPEG7 dataset).
Liang Sang, Marcin Grzegorzek, John See
ACM Multimedia4
2022 Exploring the Semi-Supervised Video Object Segmentation Problem from a Cyclic Perspective
Yuxi Li 0009, Ning Xu 0007, John See, Weiyao Lin
Int. J. Comput. Vis.4
2022 Needle in a Haystack: Spotting and recognising micro-expressions "in the wild"
Yee Siang Gan, John See, Huai-Qian Khor, Kunhong Liu 0001, Sze-Teng Liong
Neurocomputing2
2022 Learning Scale-Consistent Attention Part Network for Fine-Grained Image Recognition
abstract
Discriminative region localization and feature learning are crucial for fine-grained visual recognition. Existing approaches solve this issue by attention mechanism or part based methods while neglecting consistency between attention and local parts, as well as the rich relation information among parts. This paper proposes a Scale-consistent Attention Part Network (SCAPNet) to address that issue, which seamlessly integrates three novel modules: grid gate attention unit (gGAU), scale-consistent attention part selection (SCAPS), and part relation modeling (PRM). The gGAU module represents the grid region at a certain fine-scale with middle layer CNN features and produces hard attention maps with the lightweight Gumbel-Max based gate. The SCAPS module utilizes attention to guide part selection across multi-scales and keep the selection scale-consistent. The PRM module utilizes the self-attention mechanism to build the relationship among parts based on their appearance and relative geo-positions. SCAPNet can be learned in an end-to-end way and demonstrates state-of-the-art accuracy on several publicly available fine-grained recognition datasets (CUB-200-2011, FGVC-Aircraft, Veg200, and Fru92).
Huabin Liu 0001, John See, Weiyao Lin
IEEE Trans. Multim.4
2021 Variational Pedestrian Detection
abstract
Pedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the current IoU-based ground truth assignment procedure in classical object detection methods. In this paper, we develop a unique perspective of pedestrian detection as a variational inference problem. We formulate a novel and efficient algorithm for pedestrian detection by modeling the dense proposals as a latent variable while proposing a customized Auto-Encoding Variational Bayes (AEVB) algorithm. Through the optimization of our proposed algorithm, a classical detector can be fashioned into a variational pedestrian detector. Experiments conducted on CrowdHuman and CityPersons datasets show that the proposed algorithm serves as an efficient solution to handle the dense pedestrian detection problem for the case of single-stage detectors. Our method can also be flexibly applied to two-stage detectors, achieving notable performance enhancement.
Huanyu He, Yuxi Li 0009, John See, Weiyao Lin
CVPR5
2021 MLife: A Lite Framework for Machine Learning Lifecycle Initialization
abstract
Machine learning (ML) lifecycle is a cyclic process to build an efficient ML system. Though a lot of commercial and community (non-commercial) frameworks have been proposed to streamline the major stages in the ML lifecycle, they are normally overqualified and insufficient for an ML system in its nascent phase. Driven by real-world experience in building and maintaining ML systems, we find that it is more efficient to initialize the major stages of ML lifecycle first for trial and error, followed by the extension of specific stages to acclimatize towards more complex scenarios. For this, we introduce a simple yet flexible framework, MLife, for fast ML lifecycle initialization. This is built on the fact that data flow in MLife is in a closed loop driven by badcases, especially those which impact ML model performance the most but also provide the most value for further ML model development - a key factor towards enabling enterprises to fast track their ML capabilities.
Yunhui Zhang, Lina Shen, John See
DSAA7
2021 Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization
abstract
The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics and neglected lower-level representations and their temporal relationship which are crucial for general video understanding. To address these challenges, this paper proposes a multi-level feature optimization framework to improve the generalization and temporal modeling ability of learned video representations. Concretely, high-level features obtained from naive and prototypical contrastive learning are utilized to build distribution graphs, guiding the process of low-level and mid-level feature learning. We also devise a simple temporal modeling module from multi-level features to enhance motion pattern learning. Experiments demonstrate that multi-level feature optimization with the graph constraint and temporal modeling can greatly improve the representation ability in video understanding. Code is available$here$.
Rui Qian 0001, Yuxi Li 0009, Huabin Liu 0001, John See, Shuangrui Ding, Weiyao Lin
ICCV4
2021 Shallow Optical Flow Three-Stream CNN For Macro- And Micro-Expression Spotting From Long Videos
abstract
Facial expressions vary from the visible to the subtle. In recent years, the analysis of micro-expressions— a natural occurrence resulting from the suppression of one’s true emotions, has drawn the attention of researchers with a broad range of potential applications. However, spotting micro-expressions in long videos becomes increasingly challenging when intertwined with normal or macro-expressions. In this paper, we propose a shallow optical flow three-stream CNN (SOFTNet) model to predict a score that captures the likelihood of a frame being in an expression interval. By fashioning the spotting task as a regression problem, we introduce pseudo-labeling to facilitate the learning process. We demonstrate the efficacy and efficiency of the proposed approach on the recent MEGC 2020 benchmark, where state-of-the-art performance is achieved on CAS(ME)2with equally promising results on SAMM Long Videos.
Gen-Bing Liong, John See, Lai-Kuan Wong
ICIP2
2021 Generating Aesthetic Based Critique For Photographs
abstract
The recent surge in deep learning methods across multiple modalities has resulted in an increased interest in image captioning. Most advances in image captioning are still focused on the generation of factual-centric captions, which mainly describe the contents of an image. However, generating captions to provide a meaningful and opinionated critique of photographs is less studied. This paper presents a framework for leveraging aesthetic features encoded from an image aesthetic scorer, to synthesize human-like textual critique via a sequence decoder. Experiments on a large-scale dataset show that the proposed method is capable of producing promising results on relevant metrics relating to semantic diversity and synonymity, with qualitative observations demonstrating likewise. We also suggest the use of Word Mover’s Distance as a semantically intuitive and informative metric for this task.
Yong-Yaw Yeo, John See, Lai-Kuan Wong, Hui-Ngo Goh
ICIP2
2021 FME'21: 1st Workshop on Facial Micro-Expression: Advanced Techniques for Facial Expressions Generation and Spotting
abstract
Facial micro-expressions (FMEs) are involuntary facial movements that occur spontaneously when a person experiences an emotion but tries to suppress or repress the facial expression and usually occur in high-risk situations. Thus, FMEs are very short in duration, an important feature that distinguishes them from ordinary facial expressions. And MEs are considered to be one of the most valuable cues for complex human emotion understanding and lie detection. Since 2014, the computational analysis and automation of MEs have been an emerging area of face research. The workshop will explore various dimensions of the human mind through emotion understanding and FME analysis, as well as extended research based on multi modal approaches.
Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong
ACM Multimedia4
2021 Deep multi-level feature pyramids: Application for non-canonical firearm detection in video surveillance
JunYi Lim, Md Istiaque Al Jobayer, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, John See, Koksheik Wong
Eng. Appl. Artif. Intell.5
2021 MLife: a lite framework for machine learning lifecycle initialization
Yunhui Zhang, Lina Shen, John See
Mach. Learn.7
2021 AP-Loss for Accurate One-Stage Object Detection
abstract
One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework to replace the classification task in one-stage detectors with a ranking task, and adopting the average-precision loss (AP-loss) for the ranking problem. Due to its non-differentiability and non-convexity, the AP-loss cannot be optimized directly. For this purpose, we develop a novel optimization algorithm, which seamlessly combines the error-driven update scheme in perceptron learning and backpropagation algorithm in deep networks. We provide in-depth analyses on the good convergence property and computational complexity of the proposed algorithm, both theoretically and empirically. Experimental results demonstrate notable improvement in addressing the imbalance issue in object detection over existing AP-based optimization algorithms. An improved state-of-the-art performance is achieved in one-stage detectors based on AP-loss over detectors using classification-losses on various standard benchmarks. The proposed framework is also highly versatile in accommodating different network architectures. Code is available at https://github.com/cccorn/AP-loss.
Kean Chen, Weiyao Lin, John See, Junni Zou
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Pic2PolyArt: Transforming a photograph into polygon-based geometric art
Pau-Ek Low, Lai-Kuan Wong, John See, Ruisheng Ng
Signal Process. Image Commun.3
2021 Group Reidentification with Multigrained Matching and Integration
abstract
The task of reidentifying groups of people under different camera views is an important yet less-studied problem. Group reidentification (Re-ID) is a very challenging task since it is not only adversely affected by common issues in traditional single-object Re-ID problems, such as viewpoint and human pose variations, but also suffers from changes in group layout and group membership. In this paper, we propose a novel concept of group granularity by characterizing a group image by multigrained objects: individual people and subgroups of two and three people within a group. To achieve robust group Re-ID, we first introduce multigrained representations which can be extracted via the development of two separate schemes, that is, one with handcrafted descriptors and another with deep neural networks. The proposed representation seeks to characterize both appearance and spatial relations of multigrained objects, and is further equipped with importance weights which capture variations in intragroup dynamics. Optimal group-wise matching is facilitated by a multiorder matching process which, in turn, dynamically updates the importance weights in iterative fashion. We evaluated three multicamera group datasets containing complex scenarios and large dynamics, with experimental results demonstrating the effectiveness of our approach.
Weiyao Lin, Yuxi Li 0009, John See, Junni Zou, Hongkai Xiong, Jingdong Wang 0001, Tao Mei 0001
IEEE Trans. Cybern.4
2021 Dress With Style: Learning Style From Joint Deep Embedding of Clothing Styles and Body Shapes
abstract
Body shape is about proportion, and fashion style is all about dressing those proportions to look their very best. Figuring out the styles to suit a body shape can be a daunting task for many people. It is, therefore, essential to develop a framework for learning the compatibility of body shapes and clothing styles. Though fashion designers and fashion stylists have analyzed the correlation between human body shapes and fashion styles for a long time, this issue did not receive much attention in multimedia science. In this paper, we present a novel style recommender, on the basis of the user's body attributes. The rich amount of fashion styling knowledge from social big data is exploited for this purpose. We first construct a joint embedding of clothing styles and human body measurements with deep multimodal representation learning on a reference dataset that has been sorted to meet the fashion rules. We then discover the relevant semantic features by propagation and selection in clothing style and body shape graphs. Experiments demonstrate the effectiveness of the proposed framework when compared with several baseline methods.
Shintami Chusnul Hidayati, Ting Wei Goh, Ji-Sheng Gary Chan, Cheng-Chun Hsu, John See, Lai-Kuan Wong, Kai-Lung Hua, Yu Tsao 0001, Wen-Huang Cheng
IEEE Trans. Multim.5
2021 Towards Automatic Skeleton Extraction With Skeleton Grafting
abstract
This article introduces a novel approach to generate visually promising skeletons automatically without any manual tuning. In practice, it is challenging to extract promising skeletons directly using existing approaches. This is because they either cannot fully preserve shape features, or require manual intervention, such as boundary smoothing and skeleton pruning, to justify the eye-level view assumption. We propose an approach here that generates backbone and dense skeletons by shape input, and then extends the backbone branches via skeleton grafting from the dense skeleton to ensure a well-integrated output. Based on our evaluation, the generated skeletons best depict the shapes at levels that are similar to human perception. To evaluate and fully express the properties of the extracted skeletons, we introduce two potential functions within the high-order matching protocol to improve the accuracy of skeleton-based matching. These two functions fuse the similarities between skeleton graphs and geometrical relations characterized by multiple skeleton endpoints. Experiments on three high-order matching protocols show that the proposed potential functions can effectively reduce the number of incorrect matches.
Bipin Indurkhya, John See, Marcin Grzegorzek
IEEE Trans. Vis. Comput. Graph.3
2020 Finding Action Tubes with a Sparse-to-Dense Framework
abstract
The task of spatial-temporal action detection has attracted increasing researchers. Existing dominant methods solve this problem by relying on short-term information and dense serial-wise detection on each individual frames or clips. Despite their effectiveness, these methods showed inadequate use of long-term information and are prone to inefficiency. In this paper, we propose for the first time, an efficient framework that generates action tube proposals from video streams with a single forward pass in a sparse-to-dense manner. There are two key characteristics in this framework: (1) Both long-term and short-term sampled information are explicitly utilized in our spatio-temporal network, (2) A new dynamic feature sampling module (DTS) is designed to effectively approximate the tube output while keeping the system tractable. We evaluate the efficacy of our model on the UCF101-24, JHMDB-21 and UCFSports benchmark datasets, achieving promising results that are competitive to state-of-the-art methods. The proposed sparse-to-dense strategy rendered our framework about 7.6 times more efficient than the nearest competitor.
Yuxi Li 0009, Weiyao Lin, Tao Wang 0002, John See, Rui Qian 0001, Ning Xu 0007, Limin Wang 0002, Shugong Xu
AAAI4
2020 PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
Kean Chen, Weiyao Lin, John See, Yan Ke
ECCV (5)4
2020 CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li 0009, Weiyao Lin, John See, Ning Xu 0007, Shugong Xu, Yan Ke
ECCV (16)3
2020 MEGC2020 - The Third Facial Micro-Expression Grand Challenge
abstract
The recent emergence of automatic facial micro-expression analysis has attracted a lot of attention in the last five years. Compared to the advances made in micro-expression recognition, the task of micro-expression spotting from long videos is tremendously in need of more effective methods. This paper summarises the 3rd Facial Micro-Expression Grand Challenge (MEGC 2020) held in conjunction with the 15th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2020. In this workshop, we propose a new challenge of spotting both macro- and micro-expressions from long videos, to spur the community to develop new techniques for micro-expression spotting and also to extend facial micro-expression analysis to more complex real-world scenarios where micro-expressions are likely to be intertwined among normal expressions. In this paper, we outline the evaluation protocols for the challenge task, and describe the datasets involved. Then, we summarize the methods from the accepted challenge papers, present the comparison and analysis of results, as well as future directions.
Jingting Li 0001, Moi Hoon Yap, John See, Xiaopeng Hong
FG4
2020 Learning Image Aesthetics by Learning Inpainting
abstract
Due to the high capability of learning robust features, convolutional neural networks (CNN) are becoming a mainstay solution for many computer vision problems, including aesthetic quality assessment (AQA). However, there remains the issue that learning with CNN requires time-consuming and expensive data annotations especially for a task like AQA. In this paper, we present a novel approach to AQA that incorporates self-supervised learning (SSL) by learning how to inpaint images according to photographic rules such as rules-of-thirds and visual saliency. We conduct extensive quantitative experiments on a variety of pretext tasks and also different ways of masking patches for inpainting, reporting fairer distribution-based metrics. We also show the suitability and practicality of the inpainting task which yielded comparably good benchmark results with much lighter model complexity.
June Hao Ching, John See, Lai-Kuan Wong
ICIP2
2020 Image Dehazing With Contextualized Attentive U-NET
abstract
Haze, which occurs due to the accumulation of fine dust or smoke particles in the atmosphere, degrades outdoor imaging, resulting in reduced attractiveness of outdoor photography and the effectiveness of vision-based systems. In this paper, we present an end-to-end convolutional neural network for image dehazing. Our proposed U-Net based architecture employs Squeeze-and-Excitation (SE) blocks at the skip connections to enforce channel-wise attention and parallelized dilated convolution blocks at the bottleneck to capture both local and global context, resulting in a richer representation of the image features. Experimental results demonstrate the effectiveness of the proposed method in achieving state-of-the-art performance on the benchmark SOTS dataset.
Yean-Wei Lee, Lai-Kuan Wong, John See
ICIP3
2020 Where Is The Emotion? Dissecting A Multi-Gap Network For Image Emotion Classification
abstract
Image emotion recognition has become an increasingly popular research domain in the area of image processing and affective computing. Despite fast-improving classification performance in this task, the understanding and interpretability of its performance are still lacking as there are limited studies on which part of an image would invoke a particular emotion. In this work, we propose a Multi-GAP deep neural network for image emotion classification, which is extensible to accommodate multiple streams of information. We also incorporate feature dependency into our network blocks by adding a bidirectional GRU network to learn transitional features. We report extensive results on the variants of our proposed network and provide valuable perspectives into the class-activated regions via Grad-CAM, and network depth contributions by truncation strategy.
Lucinda Lim, Huai-Qian Khor, Phatcharawat Chaemchoy, John See, Lai-Kuan Wong
ICIP4
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia13
2020 Spectrogram-Based Classification Of Spoken Foul Language Using Deep CNN
abstract
Excessive content of profanity in audio and video files has proven to shape one's character and behavior. Currently, conventional methods of manual detection and censorship are being used. Manual censorship method is time consuming and prone to misdetection of foul language. This paper proposed an intelligent model for foul language censorship through automated and robust detection by deep convolutional neural networks (CNNs). A dataset of foul language was collected and processed for the computation of audio spectrogram images that serve as an input to evaluate the classification of foul language. The proposed model was first tested for 2-class (Foul vs Normal) classification problem, the foul class is then further decomposed into a 10-class classification problem for exact detection of profanity. Experimental results show the viability of proposed system by demonstrating high performance of curse words classification with 1.24-2.71 Error Rate (ER) for 2-class and 5.49-8.30 F1- score. Proposed Resnet50 architecture outperforms other models in terms of accuracy, sensitivity, specificity, F1-score.
Abdulaziz Saleh Ba Wazir, H. Abdul Karim, Mohd Haris Lye, Sarina Mansor, Nouar AlDahoul, Mohammad Faizal Ahmad Fauzi, John See
MMSP7
2020 Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation
abstract
In this paper, we take attempt to incorporate the cyclic mechanism with the vision task of semi-supervised video object segmentation. By resorting to the accurate reference mask of the first frame, we try to mitigate the error propagation problem in most of current video object segmentation pipelines. Firstly, we propose a cyclic scheme for offline training of segmentation networks. Then, we extend the offline pipeline to an online method by introducing a simple gradient correction module while keeping high efficiency as other offline methods. Finally we develop cycle effective receptive field (cycle-ERF) from gradient correction to provide a new perspective for analyzing object-specific regions of interests. We conduct comprehensive experiments on benchmarks of DAVIS17 and Youtube-VOS, demonstrating that our introduced cyclic mechanism is helpful to boost the segmentation quality.
Yuxi Li 0009, Ning Xu 0007, Jinlong Peng, John See, Weiyao Lin
NeurIPS4
2020 TPM: Multiple object tracking with tracklet-plane matching
Jinlong Peng, Tao Wang 0002, Weiyao Lin, Jian Wang 0066, John See, Shilei Wen, Errui Ding
Pattern Recognit.5
2020 Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC
abstract
This article addresses neural network based post-processing for the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC). We first propose a partition-aware convolution neural network (CNN) that utilizes the partition information produced by the encoder to assist in the post-processing. In contrast to existing CNN-based approaches, which only take the decoded frame as input, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the artifacts introduced by HEVC are efficiently reduced. We further introduce an adaptive-switching neural network (ASN) that consists of multiple independent CNNs to adaptively handle the variations in content and distortion within compressed-video frames, providing further reduction in visual artifacts. Additionally, an iterative training procedure is proposed to train these independent CNNs attentively on different local patch-wise classes. Experiments on benchmark sequences demonstrate the effectiveness of our partition-aware and adaptive-switching neural networks.
Weiyao Lin, Xiaoyi He, Xintong Han, Dong Liu 0002, John See, Junni Zou, Hongkai Xiong, Feng Wu 0001
IEEE Trans. Multim.5
2019 Towards Accurate One-Stage Object Detection With AP-Loss
abstract
One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework to replace the classification task in one-stage detectors with a ranking task, and adopting the Average-Precision loss (AP-loss) for the ranking problem. Due to its non-differentiability and non-convexity, the AP-loss cannot be optimized directly. For this purpose, we develop a novel optimization algorithm, which seamlessly combines the error-driven update scheme in perceptron learning and backpropagation algorithm in deep networks. We verify good convergence property of the proposed algorithm theoretically and empirically. Experimental results demonstrate notable performance improvement in state-of-the-art one-stage detectors based on AP-loss over different kinds of classification-losses on various benchmarks, without changing the network architectures.
Kean Chen, Weiyao Lin, John See, Ling-Yu Duan, Zhibo Chen 0006, Changwei He, Junni Zou
CVPR4
2019 Shallow Triple Stream Three-dimensional CNN (STSTNet) for Micro-expression Recognition
abstract
In the recent year, state-of-the-art for facial micro-expression recognition have been significantly advanced by deep neural networks. The robustness of deep learning has yielded promising performance beyond that of traditional handcrafted approaches. Most works in literature emphasized on increasing the depth of networks and employing highly complex objective functions to learn more features. In this paper, we design a Shallow Triple Stream Three-dimensional CNN (STSTNet) that is computationally light whilst capable of extracting discriminative high level features and details of micro-expressions. The network learns from three optical flow features (i.e., optical strain, horizontal and vertical optical flow fields) computed based on the onset and apex frames of each video. Our experimental results demonstrate the effectiveness of the proposed STSTNet, which obtained an unweighted average recall rate of 0.7605 and unweighted F1-score of 0.7353 on the composite database consisting of 442 samples from the SMIC, CASME II and SAMM databases.
Sze-Teng Liong, Yee Siang Gan, John See, Huai-Qian Khor, Yen-Chang Huang
FG3
2019 MEGC 2019 - The Second Facial Micro-Expressions Grand Challenge
abstract
Automatic facial micro-expression (ME) analysis is a growing field of research that has gained much attention in the last five years. With many recent works testing on limited data, there is a need to spur better approaches that are both robust and effective. This paper summarises the 2nd Facial Micro-Expression Grand Challenge (MEGC 2019) held in conjunction with the 14th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2019. In this workshop, we proposed challenges for two micro-expression (ME) tasks- spotting and recognition, with the aim of encouraging rigorous evaluation and development of new robust techniques that can accommodate data captured across a variety of settings. In this paper, we outline the evaluation protocols for the two challenge tasks, the datasets involved, and an analysis of the best performing works from the participating teams, together with a summary of results. Finally, we highlight some possible future directions.
John See, Moi Hoon Yap, Jingting Li 0001, Xiaopeng Hong
FG1
2019 Dual-stream Shallow Networks for Facial Micro-expression Recognition
abstract
Micro-expressions are spontaneous, brief and subtle facial muscle movements that exposes underlying emotions. Motivated by recent exploits into deep learning for micro-expression analysis, we propose a lightweight dual-stream shallow network in the form of a pair of truncated CNNs with heterogeneous input features. The merging of the convolutional features allows for discriminative learning of micro-expression classes stemming from both streams. Using activation heatmaps, we further demonstrate that salient facial areas are well emphasized, and correspond closely to relevant action units belonging to emotion classes. We empirically validate the proposed network on three benchmark databases, obtaining state-of-the-art performance on the CASME II and SAMM while remaining competitive on the SMIC. Further observations point towards the sufficiency of utilizing shallower deep networks for micro-expression recognition.
Huai-Qian Khor, John See, Sze-Teng Liong, Raphael C.-W. Phan, Weiyao Lin
ICIP2
2019 Localization Guided Fight Action Detection in Surveillance Videos
abstract
Automatic detection of fight behaviors in surveillance videos is an important task for surveillance systems. In this work, we propose a novel localization guided framework for detecting fight actions in surveillance videos. Specifically, we exploit optical flow maps to extract motion activation information, which indicates the location of active regions. Then, a detection guided alignment module is designed to adjust the localized active regions. This approach employs a two-stream based 3D convolution network as the backbone network with a novel motion acceleration representation on the temporal stream. While most existing methods are still evaluated on three benchmark datasets which were not originally collected from surveillance scenarios, we present a novel Fight Action Detection in Surveillance-videos (FADS) dataset for this purpose. With a total of 1,520 video clips, the FADS is the largest known dataset in terms of number of surveillance videos with fight scenes. Experimental results on both the benchmark datasets and the FADS show that our proposed localization guided method outperforms state-of-the-art techniques.
Qichao Xu, John See, Weiyao Lin
ICME2
2019 LiteEmo: Lightweight Deep Neural Networks for Image Emotion Recognition
abstract
Psychology studies have shown that an image can invoke various emotions, depending on the visual features as well as semantic content of the image. Ability to identify image emotion can be very useful for many applications, including image retrieval and aesthetics prediction. Notably, most of the existing deep learning-based emotion recognition models do not capitalize on additional semantics or contextual information and are computational expensive. Inspired to overcome these limitations, we proposed a lightweight multi-stream deep network that concatenates several MobileNet networks for performing image emotion analysis. Each stream in the multi-stream deep network represents the core emotion recognition, object recognition and image category recognition models respectively. Experimental results demonstrate the effectiveness of the additional contextual information in producing comparable performance as the state-of-the-art emotion models, but with lesser parameters, thus improving its practicality.
Yan-Han Chew, Lai-Kuan Wong, John See, Huai-Qian Khor, Balasubramanian Abivishaq
MMSP3
2019 Micro-expression recognition based on 3D flow convolutional neural network
Yandan Wang, John See
Pattern Anal. Appl.3
2019 Beauty Is in the Eye of the Beholder: Demographically Oriented Analysis of Aesthetics in Photographs
abstract
Aesthetics is a subjective concept that is likely to be perceived differently among people of different ages, genders, and cultural backgrounds. While techniques that directly compute this concept in images has seen increasing attention by the multimedia and machine-learning community, there are very few attempts at encoding the influences from the photographer’s viewpoint. This work demonstrates how the aesthetic quality of photos can be better learned by accounting for the demographic background of a photographer. A new AVA-PD (Photographer Demographic) dataset is created to supplement the AVA dataset by providing photographers’ age, gender and location attributes. Two deep convolutional neural network (CNN) architectures are proposed to utilize demographic information for aesthetic prediction of photos; both are shown to yield better prediction capabilities compared to most existing approaches. By leveraging on AVA-PD meta-data, we also present some additional machine-learnable tasks such as identifying the photographer and predicting photography styles from a person’s gallery of photos.
Magzhan Kairanbay, John See, Lai-Kuan Wong
ACM Trans. Multim. Comput. Commun. Appl.2
2018 Enriched Long-Term Recurrent Convolutional Network for Facial Micro-Expression Recognition
abstract
Facial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved superior performance in micro-expression recognition but at the cost of domain specificity and cumbersome parametric tunings. In this paper, we propose an Enriched Long-term Recurrent Convolutional Network (ELRCN) that first encodes each micro-expression frame into a feature vector through CNN module(s), then predicts the micro-expression by passing the feature vector through a Long Short-term Memory (LSTM) module. The framework contains 2 different network variants: (1) Channel-wise stacking of input data for spatial enrichment, (2) Feature-wise stacking of features for temporal enrichment. We demonstrate that the proposed approach is able to achieve reasonably good performance, without data augmentation. In addition, we also present ablation studies conducted on the framework and visualizations of what CNN "sees" when predicting the micro-expression classes.
Huai-Qian Khor, John See, Raphael C.-W. Phan, Weiyao Lin
FG2
2018 Micro-Expression Motion Magnification: Global Lagrangian vs. Local Eulerian Approaches
abstract
Micro-expressions are difficult to spot but are utterly important for engaging in a conversation or negotiation. Through motion magnification, these expressions become much more distinguishable and easily recognized. This work proposes Global Lagrangian Motion Magnification (GLMM) for consistent exaggeration of facial expressions and dynamics across a whole video. As the proposal takes an opposite approach to a previous pivotal work, i.e. local Amplitude-based Eulerian Motion Magnification (AEMM). GLMM and AEMM are theoretically analyzed for potential advantages and disadvantages, especially with respect to how magnified noise and distortions are dealt with. Then, both GLMM and AEMM are empirically evaluated and compared using the CASME II micro-expression corpus.
Anh Cat Le Ngo, Alan Johnston, Raphael C.-W. Phan, John See
FG4
2018 Facial Micro-Expressions Grand Challenge 2018 Summary
abstract
This paper summarises the Facial Micro-Expression Grand Challenge (MEGC 2018) held in conjunction with the 13th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2018. In this workshop, we aim to stimulate new ideas and techniques for facial micro-expression analysis by proposing a new cross-database challenge. Two state-of-the-art datasets, CASME II and SAMM, are used to validate the performance of existing and new algorithms. Also, the challenge advocates the recognition of micro-expressions based on AU-centric objective classes rather than emotional classes. We present a summary and analysis of the baseline results using LBP-TOP, HOOF and 3DHOG, together with results from the challenge submissions.
Moi Hoon Yap, John See, Xiaopeng Hong
FG2
2018 Vehicle Semantics Extraction and Retrieval for Long-Term Carpark Video Surveillance
Clarence Weihan Cheong, Ryan Woei-Sheng Lim, John See, Lai-Kuan Wong, Ian K. T. Tan, Azrin Aris
MMM (2)3
2018 Towards Demographic-Based Photographic Aesthetics Prediction for Portraitures
Magzhan Kairanbay, John See, Lai-Kuan Wong
MMM (1)2
2018 Multi-scale Spatiotemporal Information Fusion Network for Video Action Recognition
abstract
Two-stream convolutional networks have shown excellent performance in video action recognition in recent years. However, it remains unclear how to model the correlation between the temporal and spatial streams more effectively. First, the spatial stream and temporal stream pay attention to different aspects, which can lead to different recognition results. Second, the variety in the length of optical flow fields tends to have a great impact on the classification results. In this paper, we propose a novel multi-scale spatiotemporal information fusion network to fuse the spatial and temporal features. Specifically, our network takes advantage of multi-scale temporal information to better utilize the motion cues. Considering the complementary relationship between the spatial and temporal features, we take the hierarchical fusion strategies and asynchronous fusion method to fuse the two-stream features. Experimental results on two benchmark datasets (UCF101 and HMDB51) show that the proposed network achieves competitive performance.
Yutong Cai, Weiyao Lin, John See, Ming-Ming Cheng, Guangcan Liu, Hongkai Xiong
VCIP3
2018 Tracklet Siamese Network with Constrained Clustering for Multiple Object Tracking
abstract
Multiple object tracking (MOT) is an important yet challenging task in video understanding and analysis. Basically, MOT aims to associate detected objects into trajectories based on their temporal relationships. The occlusion among moving objects poses a major challenge towards robust modeling of these relationships. In this paper, we propose a novel Tracklet Siamese Network (TSN) for learning similarities between track-lets characterized by appearance information, achieving superior performance on two MOTChallenge benchmark datasets. Our framework constructs short tracklets from highly-related object detections by excluding inaccurate object detections. We also adopt a constrained clustering technique to piece tracklets together into long trajectories, thus recovering many missing detections caused by original detector or the detection removing in the previous step. Comparisons against state-of-the-art methods were reported while ablation studies further substantiate the viability of components in our approach.
Jinlong Peng, Fan Qiu, John See, Shaoshuai Huang, Ling-Yu Duan, Weiyao Lin
VCIP3
2018 Semantic facial scores and compact deep transferred descriptors for scalable face image retrieval
Rasoul Banaeeyan, Mohd Haris Lye, Mohammad Faizal Ahmad Fauzi, H. Abdul Karim, John See
Neurocomputing5
2018 Less is more: Micro-expression recognition from video using apex frame
Sze-Teng Liong, John See, Koksheik Wong, Raphael C.-W. Phan
Signal Process. Image Commun.2
2017 Multigap: Multi-pooled inception network with text augmentation for aesthetic prediction of photographs
abstract
With the advent of deep learning, convolutional neural networks have solved many imaging problems to a large extent. However, it remains to be seen if the image “bottleneck” can be unplugged by harnessing complementary sources of data. In this paper, we present a new approach to image aesthetic evaluation that learns both visual and textual features simultaneously. Our network extracts visual features by appending global average pooling blocks on multiple inception modules (MultiGAP), while textual features from associated user comments are learned from a recurrent neural network. Experimental results show that the proposed method is capable of achieving state-of-the-art performance on the AVA / AVA-Comments datasets. We also demonstrate the capability of our approach in visualizing aesthetic activations.
Yong-Lian Hii, John See, Magzhan Kairanbay, Lai-Kuan Wong
ICIP2
2017 Filling the gaps: Reducing the complexity of networks for multi-attribute image aesthetic prediction
abstract
Computational aesthetics have seen much progress in recent years with the increasing popularity of deep learning methods. In this paper, we present two approaches that leverage on the benefits of using Global Average Pooling (GAP) to reduce the complexity of deep convolutional neural networks. The first model fine-tunes a standard CNN with a newly introduced GAP layer. The second approach extracts global and local CNN codes by reducing the dimensionality of convolution layers with individual GAP operations. We also extend these approaches to a multi-attribute network which uses a style network to regularize the aesthetic network. Experiments demonstrate the capability of attaining comparable accuracy results while reducing training complexity substantially.
Magzhan Kairanbay, John See, Lai-Kuan Wong, Yong-Lian Hii
ICIP2
2017 Effective recognition of facial micro-expressions with video motion magnification
Yandan Wang, John See, Yee-Hui Oh, Raphael C.-W. Phan, Yo Rahul, Huo-Chong Ling, Su-Wei Tan, Xujie Li 0002
Multim. Tools Appl.2
2017 Sparsity in Dynamics of Spontaneous Subtle Emotions: Analysis and Application
abstract
Subtle emotions are present in diverse real-life situations: in hostile environments, enemies and/or spies maliciouslyconceal their emotions as part of their deception; in life-threatening situations, victims under duress have no choice but to withhold theirreal feelings; in the medical scene, patients with psychological conditions such as depression could either be intentionally orsubconsciously suppressing their anguish from loved ones. Under such circumstances, it is often crucial that these subtle emotions arerecognized before it is too late. These spontaneous subtle emotions are typically expressed through micro-expressions, which are tiny,sudden and short-lived dynamics of facial muscles; thus, such micro-expressions pose a great challenge for visual recognition. Theabrupt but significant dynamics for the recognition task are temporally sparse while the rest, i.e. irrelevant dynamics, are temporallyredundant. In this work, we analyze and enforce sparsity constraints to learn significant temporal and spectral structures whileeliminating irrelevant facial dynamics of micro-expressions, which would ease the challenge in the visual recognition of spontaneoussubtle emotions. The hypothesis is confirmed through experimental results of automatic spontaneous subtle emotion recognition withseveral sparsity levels on CASME II and SMIC, the two well-established and publicly available spontaneous subtle emotion databases.The overall performances of the automatic subtle emotion recognition are boosted when only significant dynamics of the originalsequences are preserved.
Anh Cat Le Ngo, John See, Raphael C.-W. Phan
IEEE Trans. Affect. Comput.2
2016 DUKE: Enhancing Virtual Reality based FPS Game with Full-body Interactions
abstract
Project DUKE explores the design and development of a virtual reality based first-person shooter (FPS) game using gesture based technologies. By integrating Oculus Rift, Microsoft Kinect and Leap Motion technologies, we attempt to foster natural in-game hand presence while maintaining foot- and torso-based navigation. We demonstrate in this exploratory work, a fully immersive prototype that enables a full-body FPS gaming experience involving interactions for complete navigation, combat and viewing control.
Mohd Hezri Amir, Albert Quek, Nur Rasyid Bin Sulaiman, John See
ACE4
2016 Eulerian emotion magnification for subtle expression recognition
abstract
Subtle emotions are expressed through tiny and brief movements of facial muscles, called micro-expressions; thus, recognition of these hidden expressions is as challenging as inspection of microscopic worlds without microscopes. In this paper, we show that through motion magnification, subtle expressions can be realistically exaggerated and become more easily recognisable. We magnify motions of facial expressions in the Eulerian perspective by manipulating their amplitudes or phases. To evaluate effects of exaggerating facial expressions, we use a common framework (LBP-TOP features and SVM classifiers) to perform 5-class subtle emotion recognition on the CASME II corpus, a spontaneous subtle emotion database. According to experimental results, significant improvements in recognition rates of magnified micro-expressions over normal ones are confirmed and measured. Furthermore, we estimate upper bounds of effective magnification factors and empirically corroborate these theoretical calculations with experimental data.
Anh Cat Le Ngo, Yee-Hui Oh, Raphael C.-W. Phan, John See
ICASSP4
2016 Intrinsic two-dimensional local structures for micro-expression recognition
abstract
An elapsed facial emotion involves changes of facial contour due to the motions (such as contraction or stretch) of facial muscles located at the eyes, nose, lips and etc. Thus, the important information such as corners of facial contours that are located in various regions of the face are crucial to the recognition of facial expressions, and even more apparent for micro-expressions. In this paper, we propose the first known notion of employing intrinsic two-dimensional (i2D) local structures to represent these features for micro-expression recognition. To retrieve i2D local structures such as phase and orientation, higher order Riesz transforms are employed by means of monogenic curvature tensors. Experiments performed on micro-expression datasets show the effectiveness of i2D local structures in recognizing micro-expressions.
Yee-Hui Oh, Anh Cat Le Ngo, Raphael C.-W. Phan, John See, Huo-Chong Ling
ICASSP4
2016 Spatio-temporal mid-level feature bank for action recognition in low quality video
abstract
It is a great challenge to perform high level recognition tasks on videos that are poor in quality. In this paper, we propose a new spatio-temporal mid-level (STEM) feature bank for recognizing human actions in low quality videos. The feature bank comprises of a trio of local spatio-temporal features, i.e. shape, motion and textures, which respectively encode structural, dynamic and statistical information in video. These features are encoded into mid-level representations and aggregated to construct STEM. Based on the recent binarized statistical image feature (BSIF), we also design a new spatiotemporal textural feature that extracts discriminately from 3D salient patches. Extensive experiments on the poor quality versions/subsets of the KTH and HMDB51 datasets demonstrate the effectiveness of the proposed approach.
Saimunur Rahman, John See
ICASSP2
2016 Spontaneous subtle expression detection and recognition based on facial strain
Sze-Teng Liong, John See, Raphael C.-W. Phan, Yee-Hui Oh, Anh Cat Le Ngo, Koksheik Wong, Su-Wei Tan
Signal Process. Image Commun.2
2014 Spontaneous Subtle Expression Recognition: Imbalanced Databases and Solutions
Anh Cat Le Ngo, Raphael C.-W. Phan, John See
ACCV (4)3
2014 LBP with Six Intersection Points: Reducing Redundant Information in LBP-TOP for Micro-expression Recognition
Yandan Wang, John See, Raphael C.-W. Phan, Yee-Hui Oh
ACCV (1)2
2013 Fusing cluster-centric feature similarities for face recognition in video sequences
John See, Mohammad Faizal Ahmad Fauzi, Chikkannan Eswaran
Pattern Recognit. Lett.1
2012 Dual-Feature Bayesian MAP Classification: Exploiting Temporal Information for Video-Based Face Recognition
John See, Chikkannan Eswaran, Mohammad Faizal Ahmad Fauzi
ICONIP (5)1
2011 Exemplar extraction using spatio-temporal hierarchical agglomerative clustering for face recognition in video
abstract
Many recent works have attempted to improve object recognition by exploiting temporal dynamics, an intrinsic property of video sequences. In this paper, a new spatio-temporal hierarchical agglomerative clustering (STHAC) method is proposed for automatic extraction of face exemplars for face recognition in video sequences. Two variants of STHAC are presented - a global variety that unifies spatial and temporal distances between points, and a local variety that introduces perturbation of distances based on a local spatio-temporal neighborhood criterion. Faces that are nearest to the cluster means are chosen as exemplars for the testing stage, where subjects in the test video sequences are recognized using a probabilistic-based classifier. Extensive evaluation on a face video database demonstrates the effectiveness of our proposed method, and the significance of incorporating temporal information for exemplar extraction.
John See, Chikkannan Eswaran
ICCV1
2011 Probabilistic Bayesian network classifier for face recognition in video sequences
abstract
The inherent properties of video sequences allow for representation of data in both spatial and temporal dimensions. Using conventional image-based methods for face recognition in video is often an ineffective approach as the essential spatio-temporal properties are not fully harnessed. This paper proposes a probabilistic Bayesian network classifier to accomplish effective recognition of faces in video sequences. In our model, we introduce a joint probability function that encodes the causal dependencies between video frames, selected exemplars or representative images of a video, and subject classes. This enables both the temporal continuity between video frames and also the spatial relationships between exemplars and their respective exemplar-set classes to be captured. To simplify the tedious estimation of densities, the proposed method also utilizes probabilistic similarity scores that are computationally inexpensive. Good recognition rates were achieved by our proposed method in comprehensive experiments conducted on two standard face video datasets.
John See
ISDA1