Lai-Kuan Wong

dblp:04/8109 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-4517-0391ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond illusions of competence: Revisiting zero-shot learning in emotion recognition with FEAr dataset
Zhong Ken Hew, Lai-Kuan Wong, Chee Seng Chan
Signal Process. Image Commun.2
2026 A Lightweight Deep and Wide Network for Image-Based Detection of Industrial Waste Gas
abstract
Due to inadequate monitoring, key pollutants (e.g., PM2.5, VOCs, etc) very possibly leak into atmosphere, thus to endanger the long-term and short-term life safety of people that work and live in the environment. Therefore, it is imperative to effectively and efficiently detect the leakage of industrial waste gas, for the purpose of timely lowering the risk of pollution and explosions. To solve such a problem, we in this paper propose a new lightweight deep and wide network (LdwNet) for detecting the leakage of industrial waste gas from an image, which brings about the two main merits: 1) Compensating for the deficiencies of sensor-based detection methods, which can accurately detect the leakage of waste gas and even measure its concentrations but require to seek leakage sources beforehand; 2) Overcoming the shortcomings of image-based detection methods, which leverage DNN-based recognition technologies and usually suffer from low efficacy, low efficiency and high energy consumption during the model training and inference. To specify, the proposed LdwNet is developed by simulating human perception, motivated by the method which detects the leakage of industrial waste gas from surveillance images with the human observation and judgement. First, based on the inspiration that the human eyes are highly sensitive to horizontal and vertical stimuli, we construct a novel lightweight parallel-series-stripe (PS2) module to validly extract features with very few parameters. Second, to fully exploit deep and shallow features for fusing the global and local information, we extend the PS2 module as a backbone along both the deep and wide directions to build the multi-channel network. Third, to achieve effective, efficient and low-carbon detection in model running, we constraint the extended PS2 modules with parameter sharing to prodigiously reduce the model parameters and thus to make the proposed model ultra-lightweight. Experiments on the datasets of carbon particulate matters and ethylene leakage prove that our LdwNet with ten thousand parameters outperforms the state-of-the-art models with millions of parameters in detection accuracy and implementation cost, and this renders our proposed LdwNet more suitable for real industrial applications.
Ke Gu 0001, Hongyan Liu 0004, Jingchao Cao, Lai-Kuan Wong, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001, Weisi Lin, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2026 Infrared Image Quality Estimation With Node-to-Graph Regression
abstract
By comparison with the commonly seen visible light images that can be effectively characterized within a Euclidean space, infrared images have non-Euclidean characteristics since their pixels contain rich thermal radiation information, such as heat distribution, surface temperature and thermal radiation. Considering the advantages of Graph Convolutional Networks (GCNs) in processing non-Euclidean data, this study proposes to introduce the GCNs to estimate the quality of infrared images by developing the Node-to-Graph Regression (NGR) model. To specify, the proposed NGR model is composed of two main steps, namely network establishment and network training. In the first step, following the classical researches of image quality estimation that include local distortion measurement followed by pooling for inferring the image quality score, this study captures the local distortion of the input infrared images by stacking up a set of Vision Graph (VSG) blocks to generate one node map, and then conducts the weighted pooling method on the node map to yield the graph output as the estimated quality score. In the second step, for enhancing the model's performance and generalization ability in the network training process, this study implements the node regression with the big data pre-training method to raise the local distortion extraction ability in a broad range of image scenarios and distortion intensities, and then performs the graph regression by using the knowledge distillation method to reduce the over-fitting risk. Using the largest-size infrared image quality evaluation database (I2QED), this study compared the proposed NGR model with three dozen mainstream and state-of-the-art competitors, and results showed that our proposed NGR model achieved the optimal performance.
Ke Gu 0001, Hongyan Liu 0004, Yubin Gao, Chen Wang 0019, Lai-Kuan Wong, Weisi Lin, Guangtao Zhai, Wenjun Zhang 0001, Daniel Thalmann
IEEE Trans. Multim.5
2025 Mix-YOLONet: Deep Image Dehazing for Improving Object Detection
Xin Lim, Lai-Kuan Wong, Yuen Peng Loh, Ke Gu 0001, Weisi Lin
MMM (2)2
2024 KBY-Net: A Dual Learning Framework for Improving Object Detection in Rainy Weather Conditions
abstract
Rainy weather conditions significantly degrade image quality, posing a major challenge for object detection tasks. Conventional methods often address this issue through domain adaptation, or the "derain then detect" approach that utilizes image deraining as the preprocessing technique. This paper presents KBY-Net, a novel end-to-end Y-Net architecture that is built upon the YOLOv8 architecture and leverages multi-task learning for concurrent image restoration and object detection. First, KBY-Net incorporates a novel KBY-decoder designed for image deraining. This decoder leverages Cross Stage Partial (CSP) layer and kernel basis attention (KBA) module to improve feature representation. Second, KBY-Net adopted two innovative modules; a multi-Dconv head transposed attention (MDTA) module at the bottleneck and a multi-axis feature fusion (MFF) block at the neck of the Y-Net. The multi-DConv module empowers the model to capture long-range dependencies and complex representations, and the MFF block refines the extracted features – both contribute significantly to accurate object detection in challenging rainy scenes. Empirical evaluations on benchmark rainy datasets demonstrate that KBY-Net outperforms the state-ofthe-art object detection approaches by a significant margin both quantitatively and qualitatively
Zheng-Xian Keh, Lai-Kuan Wong, Yuen Peng Loh, Ke Gu 0001, Weisi Lin
MMAsia2
2023 Context-Aware Multi-Stream Networks for Dimensional Emotion Prediction in Images
abstract
Teaching machines to comprehend the nuances of emotion from photographs is a particularly challenging task. Emotion perception— naturally a subjective problem, is often simplified for computational purposes into categorical states or valence-arousal dimensional space, the latter being a lesser-explored problem in the literature. This paper proposes a multi-stream context-aware neural network model for dimensional emotion prediction in images. Models were trained using a set of object and scene data along with deep features for valence, arousal, and dominance estimation. Experimental evaluation on a large-scale image emotion dataset demonstrates the viability of our proposed approach. Our analysis postulates that the understanding of the depicted object in an image is vital for successful predictions whilst relying on scene information can lead to somewhat confounding effects.
Sidharrth Nagappan, Jia Qi Tan, Lai-Kuan Wong, John See
ICIP3
2023 HER2-Sish Histopathology Image Classification Using Deep Neural Networks
abstract
The status of the human epidermal growth factor receptor 2 (HER2) gene amplification is an important marker for assessing the efficacy of clinical treatments for breast cancer. This article discusses the application of deep learning to classify HER2-SISH (silver-enhanced in situ hybridization) pathological images and identifies their HER2/Chr17 amplification status. We used four pre-trained models for classifying the cases into either amplified or non-amplified: two models from the convolutional neural networks, CNNs (DenseNet, and MobileNet), and two transformer models (Vision Transformer, and Data-Efficient Image Transformers). Apart from these single models, we also built two ensemble models by concatenating the transformer and CNN architectures to observe their performances. A private dataset obtained from our collaborating hospital is used in this project, with several preprocessing techniques applied to the raw images prior to feeding the models. Promising results are reported with ViT emerged as the best performing model with a high accuracy of 87.48%, with 92.93% recall in detecting amplified HER2-SISH samples.
Choo Hui Tan, Wei Jie Lim, Wan Siti Halimatul Munirah Wan Ahmad, Lai-Kuan Wong, Zaka Ur Rehman, Lai-Meng Looi, Phaik-Leng Cheah, Toh Yen Fa, Mohammad Faizal Ahmad Fauzi
ICIP4
2023 Efficient anisotropic scaling and translation invariants of Tchebichef moments using image normalization
Chih-Yang Pee, Seng-Huat Ong, Raveendran Paramesran, Lai-Kuan Wong
Pattern Recognit. Lett.4
2023 LAU-Net: A low light image enhancer with attention and resizing mechanisms
Choon Chen Lim, Yuen Peng Loh, Lai-Kuan Wong
Signal Process. Image Commun.3
2022 Social-SSL: Self-supervised Cross-Sequence Representation Learning Based on Transformers for Multi-agent Trajectory Prediction
Li-Wu Tsao, Yan-Kai Wang, Hao-Siang Lin, Hong-Han Shuai, Lai-Kuan Wong, Wen-Huang Cheng
ECCV (22)5
2021 Shallow Optical Flow Three-Stream CNN For Macro- And Micro-Expression Spotting From Long Videos
abstract
Facial expressions vary from the visible to the subtle. In recent years, the analysis of micro-expressions— a natural occurrence resulting from the suppression of one’s true emotions, has drawn the attention of researchers with a broad range of potential applications. However, spotting micro-expressions in long videos becomes increasingly challenging when intertwined with normal or macro-expressions. In this paper, we propose a shallow optical flow three-stream CNN (SOFTNet) model to predict a score that captures the likelihood of a frame being in an expression interval. By fashioning the spotting task as a regression problem, we introduce pseudo-labeling to facilitate the learning process. We demonstrate the efficacy and efficiency of the proposed approach on the recent MEGC 2020 benchmark, where state-of-the-art performance is achieved on CAS(ME)2with equally promising results on SAMM Long Videos.
Gen-Bing Liong, John See, Lai-Kuan Wong
ICIP3
2021 Generating Aesthetic Based Critique For Photographs
abstract
The recent surge in deep learning methods across multiple modalities has resulted in an increased interest in image captioning. Most advances in image captioning are still focused on the generation of factual-centric captions, which mainly describe the contents of an image. However, generating captions to provide a meaningful and opinionated critique of photographs is less studied. This paper presents a framework for leveraging aesthetic features encoded from an image aesthetic scorer, to synthesize human-like textual critique via a sequence decoder. Experiments on a large-scale dataset show that the proposed method is capable of producing promising results on relevant metrics relating to semantic diversity and synonymity, with qualitative observations demonstrating likewise. We also suggest the use of Word Mover’s Distance as a semantically intuitive and informative metric for this task.
Yong-Yaw Yeo, John See, Lai-Kuan Wong, Hui-Ngo Goh
ICIP3
2021 Pic2PolyArt: Transforming a photograph into polygon-based geometric art
Pau-Ek Low, Lai-Kuan Wong, John See, Ruisheng Ng
Signal Process. Image Commun.2
2021 Dress With Style: Learning Style From Joint Deep Embedding of Clothing Styles and Body Shapes
abstract
Body shape is about proportion, and fashion style is all about dressing those proportions to look their very best. Figuring out the styles to suit a body shape can be a daunting task for many people. It is, therefore, essential to develop a framework for learning the compatibility of body shapes and clothing styles. Though fashion designers and fashion stylists have analyzed the correlation between human body shapes and fashion styles for a long time, this issue did not receive much attention in multimedia science. In this paper, we present a novel style recommender, on the basis of the user's body attributes. The rich amount of fashion styling knowledge from social big data is exploited for this purpose. We first construct a joint embedding of clothing styles and human body measurements with deep multimodal representation learning on a reference dataset that has been sorted to meet the fashion rules. We then discover the relevant semantic features by propagation and selection in clothing style and body shape graphs. Experiments demonstrate the effectiveness of the proposed framework when compared with several baseline methods.
Shintami Chusnul Hidayati, Ting Wei Goh, Ji-Sheng Gary Chan, Cheng-Chun Hsu, John See, Lai-Kuan Wong, Kai-Lung Hua, Yu Tsao 0001, Wen-Huang Cheng
IEEE Trans. Multim.6
2020 Learning Image Aesthetics by Learning Inpainting
abstract
Due to the high capability of learning robust features, convolutional neural networks (CNN) are becoming a mainstay solution for many computer vision problems, including aesthetic quality assessment (AQA). However, there remains the issue that learning with CNN requires time-consuming and expensive data annotations especially for a task like AQA. In this paper, we present a novel approach to AQA that incorporates self-supervised learning (SSL) by learning how to inpaint images according to photographic rules such as rules-of-thirds and visual saliency. We conduct extensive quantitative experiments on a variety of pretext tasks and also different ways of masking patches for inpainting, reporting fairer distribution-based metrics. We also show the suitability and practicality of the inpainting task which yielded comparably good benchmark results with much lighter model complexity.
June Hao Ching, John See, Lai-Kuan Wong
ICIP3
2020 Image Dehazing With Contextualized Attentive U-NET
abstract
Haze, which occurs due to the accumulation of fine dust or smoke particles in the atmosphere, degrades outdoor imaging, resulting in reduced attractiveness of outdoor photography and the effectiveness of vision-based systems. In this paper, we present an end-to-end convolutional neural network for image dehazing. Our proposed U-Net based architecture employs Squeeze-and-Excitation (SE) blocks at the skip connections to enforce channel-wise attention and parallelized dilated convolution blocks at the bottleneck to capture both local and global context, resulting in a richer representation of the image features. Experimental results demonstrate the effectiveness of the proposed method in achieving state-of-the-art performance on the benchmark SOTS dataset.
Yean-Wei Lee, Lai-Kuan Wong, John See
ICIP2
2020 Where Is The Emotion? Dissecting A Multi-Gap Network For Image Emotion Classification
abstract
Image emotion recognition has become an increasingly popular research domain in the area of image processing and affective computing. Despite fast-improving classification performance in this task, the understanding and interpretability of its performance are still lacking as there are limited studies on which part of an image would invoke a particular emotion. In this work, we propose a Multi-GAP deep neural network for image emotion classification, which is extensible to accommodate multiple streams of information. We also incorporate feature dependency into our network blocks by adding a bidirectional GRU network to learn transitional features. We report extensive results on the variants of our proposed network and provide valuable perspectives into the class-activated regions via Grad-CAM, and network depth contributions by truncation strategy.
Lucinda Lim, Huai-Qian Khor, Phatcharawat Chaemchoy, John See, Lai-Kuan Wong
ICIP5
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia14
2019 LiteEmo: Lightweight Deep Neural Networks for Image Emotion Recognition
abstract
Psychology studies have shown that an image can invoke various emotions, depending on the visual features as well as semantic content of the image. Ability to identify image emotion can be very useful for many applications, including image retrieval and aesthetics prediction. Notably, most of the existing deep learning-based emotion recognition models do not capitalize on additional semantics or contextual information and are computational expensive. Inspired to overcome these limitations, we proposed a lightweight multi-stream deep network that concatenates several MobileNet networks for performing image emotion analysis. Each stream in the multi-stream deep network represents the core emotion recognition, object recognition and image category recognition models respectively. Experimental results demonstrate the effectiveness of the additional contextual information in producing comparable performance as the state-of-the-art emotion models, but with lesser parameters, thus improving its practicality.
Yan-Han Chew, Lai-Kuan Wong, John See, Huai-Qian Khor, Balasubramanian Abivishaq
MMSP2
2019 Warping-Based Stereoscopic 3D Video Retargeting With Depth Remapping
abstract
Due to the recent availability of different stereoscopic display devices and online 3D media resources (e.g. 3D movies), there is a growing demand for stereoscopic video retargeting that can automatically resize a given stereoscopic video to fit the target display device. In this paper, we propose a warping-based approach that can simultaneously resize and remap the depth of a stereoscopic video to produce a better 3D viewing experience. Firstly, our method computes the significance map for each stereo video frame. It then performs volume warping using non-homogeneous scaling optimization to resize the stereoscopic video. A depth remapping constraint is used to remap the depth and a constraint is applied to preserve the significant content during warping process. Experimental results demonstrate the effectiveness of our method in preserving the significant content, ensuring motion consistency, and enhancing the depth perception of the retargeted video sequences within the comfort depth range.
Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low, Chee-Onn Wong
WACV2
2019 Beauty Is in the Eye of the Beholder: Demographically Oriented Analysis of Aesthetics in Photographs
abstract
Aesthetics is a subjective concept that is likely to be perceived differently among people of different ages, genders, and cultural backgrounds. While techniques that directly compute this concept in images has seen increasing attention by the multimedia and machine-learning community, there are very few attempts at encoding the influences from the photographer’s viewpoint. This work demonstrates how the aesthetic quality of photos can be better learned by accounting for the demographic background of a photographer. A new AVA-PD (Photographer Demographic) dataset is created to supplement the AVA dataset by providing photographers’ age, gender and location attributes. Two deep convolutional neural network (CNN) architectures are proposed to utilize demographic information for aesthetic prediction of photos; both are shown to yield better prediction capabilities compared to most existing approaches. By leveraging on AVA-PD meta-data, we also present some additional machine-learnable tasks such as identifying the photographer and predicting photography styles from a person’s gallery of photos.
Magzhan Kairanbay, John See, Lai-Kuan Wong
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Vehicle Semantics Extraction and Retrieval for Long-Term Carpark Video Surveillance
Clarence Weihan Cheong, Ryan Woei-Sheng Lim, John See, Lai-Kuan Wong, Ian K. T. Tan, Azrin Aris
MMM (2)4
2018 Towards Demographic-Based Photographic Aesthetics Prediction for Portraitures
Magzhan Kairanbay, John See, Lai-Kuan Wong
MMM (1)3
2018 The design and empirical evaluations of 3D positioning techniques for pressure-based touch control on mobile devices
Lu Wang 0007, Lai-Kuan Wong, Yajie Xu, Siyuan Qiu, Xiangxu Meng, Chenglei Yang
Pers. Ubiquitous Comput.2
2018 Aesthetics-Driven Stereoscopic 3-D Image Recomposition With Depth Adaptation
abstract
Due to the availability and affordability of the stereoscopic equipment (e.g., stereo camera, lens, and display devices), stereoscopic image manipulation has been receiving considerable research attention in recent years. In this paper, we present a semiautomatic, aesthetic-driven stereoscopic image recomposition approach, which capacitates the change of the spatial position of the foreground object(s) in a given stereoscopic image to enhance human visual aesthetics. Our algorithm recomposes both the left and right stereo images simultaneously using a global optimization algorithm. To maximize image aesthetics, our algorithm minimizes a set of aesthetic quality errors, which is derived from selected photographic composition rules. In addition, depth adaptation is applied to the resized objects and change in vertical disparity of the resulting stereo image pair is minimized to ensure a pleasant three-dimensional (3-D) viewing experience. Our method can be used to perform stereoscopic image retargeting and recomposition simultaneously by providing the target image scale as the input. Empirical evaluations demonstrate the effectiveness of our approach in enhancing the aesthetics of stereoscopic 3-D images. Notably, depth adaptation is shown to play an important role in aesthetics enhancement.
Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low, Chee-Onn Wong
IEEE Trans. Multim.2
2017 Multigap: Multi-pooled inception network with text augmentation for aesthetic prediction of photographs
abstract
With the advent of deep learning, convolutional neural networks have solved many imaging problems to a large extent. However, it remains to be seen if the image “bottleneck” can be unplugged by harnessing complementary sources of data. In this paper, we present a new approach to image aesthetic evaluation that learns both visual and textual features simultaneously. Our network extracts visual features by appending global average pooling blocks on multiple inception modules (MultiGAP), while textual features from associated user comments are learned from a recurrent neural network. Experimental results show that the proposed method is capable of achieving state-of-the-art performance on the AVA / AVA-Comments datasets. We also demonstrate the capability of our approach in visualizing aesthetic activations.
Yong-Lian Hii, John See, Magzhan Kairanbay, Lai-Kuan Wong
ICIP4
2017 Filling the gaps: Reducing the complexity of networks for multi-attribute image aesthetic prediction
abstract
Computational aesthetics have seen much progress in recent years with the increasing popularity of deep learning methods. In this paper, we present two approaches that leverage on the benefits of using Global Average Pooling (GAP) to reduce the complexity of deep convolutional neural networks. The first model fine-tunes a standard CNN with a newly introduced GAP layer. The second approach extracts global and local CNN codes by reducing the dimensionality of convolution layers with individual GAP operations. We also extend these approaches to a multi-attribute network which uses a style network to regularize the aesthetic network. Experiments demonstrate the capability of attaining comparable accuracy results while reducing training complexity substantially.
Magzhan Kairanbay, John See, Lai-Kuan Wong, Yong-Lian Hii
ICIP3
2017 A survey of aesthetics-driven image recomposition
Md Baharul Islam, Lai-Kuan Wong, Chee-Onn Wong
Multim. Tools Appl.2
2016 Pressure-based touch positioning techniques for 3D objects
abstract
Most of the previous 3 DOF (Degree Of Freedom) 3D touch positioning techniques require more than one finger (usually two hands) to be performed, which limits their using space on small mobile devices such as phones and tablets that need one hand to be held in most occasions. Given that the pressure sensitive touch screen would become ubiquitous in near future, we present the pressure-based 3DOF 3D objects positioning and manipulating techniques that only use one hand in operating.
Siyuan Qiu, Lu Wang 0007, Lai-Kuan Wong
I3D3
2015 Semantics-Preserving Warping for Stereoscopic Image Retargeting
Chun-Hau Tan, Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low
PSIVT3
2012 Enhancing visual dominance by semantics-preserving image recomposition
abstract
We present a semi-automatic photographic recomposition approach that employs a semantics-preserving warp of the input image to enhance the visual dominance of the main subjects. Our method uses the tearable image warping method to shift the subjects against the background (and vice versa), so that their visual dominance is improved, and yet preserve the desired spatial semantics between the subjects and the background. The recomposition is guided by a measure of the resulting visual dominance of the main subjects. Our user experiment shows the effectiveness of the approach.
Lai-Kuan Wong, Kok-Lim Low
ACM Multimedia1
2011 Saliency retargeting: An approach to enhance image aesthetics
abstract
A photograph that has visually dominant subjects in general induces stronger aesthetic interest. Inspired by this, we have developed a new approach to enhance image aesthetics through saliency retargeting. Our method alters low-level image features of the objects in the photograph such that their computed saliency measurements in the modified image become consistent with the intended order of their visual importance. The goal of our approach is to produce an image that can redirect the viewers' attention to the most important objects in the image, and thus making these objects the main subjects. Since many modified images can satisfy the same specified order of visual importance, we trained an aesthetics score prediction model to pick the one with the best aesthetics. Results from our user experiments support the effectiveness of our approach.
Lai-Kuan Wong, Kok-Lim Low
WACV1
2009 Saliency-enhanced image aesthetics class prediction
abstract
We present a saliency-enhanced method for the classification of professional photos and snapshots. First, we extract the salient regions from an image by utilizing a visual saliency model. We assume that the salient regions contain the photo subject. Then, in addition to a set of discriminative global image features, we extract a set of salient features that characterize the subject and depict the subject-background relationship. Our high-level perceptual approach produces a promising 5-fold cross-validation (5-CV) classification accuracy of 78.8%, significantly higher than existing approaches that concentrate mainly on global features.
Lai-Kuan Wong, Kok-Lim Low
ICIP1