VLDB 2026 Research / reviewers in the wild / expert
Kim-Hui Yap
dblp:49/1306
· DBLP profile ↗
127ranked-venue papers
12as first author
42since 2021 · last 2026
0000-0003-1933-4986ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 86 · 9 first-author · 27 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 17 since 2021Systems, architecture and hardware · 16 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Demystifying Data Organization for Enhanced LLM TrainingabstractYalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Xin Zhang, Wenshan Wu, Qihao Zhao, Hao Li, Yuanyuan Gao, Kim-Hui Yap, Scarlett Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Wenshan Wu, Qihao Zhao, Hao Li 0069, Kim-Hui Yap, Scarlett Li |
ACL (1) | 10 |
| 2026 | PromptSR: Cascade Prompting for Lightweight Image Super-ResolutionabstractAlthough the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-based self-attention modeling. The quadratic computational complexity relative to window size restricts its ability to use a large window size for expanding the receptive field while maintaining low computational costs. To address this challenge, we propose PromptSR, a novel prompt-empowered lightweight image SR method. The core component is the proposed cascade prompting block (CPB), which enhances global information access and local refinement via three cascaded prompting layers: a global anchor prompting layer (GAPL) and two local prompting layers (LPLs). The GAPL leverages downscaled features as anchors to construct low-dimensional anchor prompts (APs) through cross-scale attention, significantly reducing computational costs. These APs, with enhanced global perception, are then used to provide global prompts, efficiently facilitating long-range token connections. The two LPLs subsequently combine category-based self-attention and window-based self-attention to refine the representation in a coarse-to-fine manner. They leverage attention maps from the GAPL as additional global prompts, enabling them to perceive features globally at different granularities for adaptive local refinement. In this way, the proposed CPB effectively combines global priors and local details, significantly enlarging the receptive field while maintaining the low computational costs of our PromptSR. The experimental results demonstrate the superiority of our method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations. Our code will be released at https://github.com/wenyang001/PromptSR. Wenyang Liu, Jianjun Gao 0005, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 6 |
| 2025 | Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationabstractOnline stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches for online stereo rectification involve estimating relative camera poses in real time to derive rectification homographies. However, they do not directly optimize for rectification constraints. Additionally, the general-purpose correspondence matchers used in these methods are not trained for rectification, while training of these matchers typically requires ground-truth correspondences which are not available in stereo rectification datasets. To address these limitations, we propose a matching-based stereo rectification framework that is directly optimized for rectification and does not require ground-truth correspondence annotations for training. We assume intrinsics are known as they are generally available on modern devices and are relatively stable. Our framework incorporates a rectification-constrained estimator and applies multi-level, rectification-specific supervision that trains the matcher network for rectification without relying on ground-truth correspondences. Additionally, we create a new rectification dataset with ground-truth optical flow annotations, eliminating bias from evaluation metrics used in prior work that relied on pretrained keypoint matching or optical flow models. Extensive experiments show that our approach outperforms both state-of-the-art matching-based and matching-free methods in vertical flow metric by 10.7% on the Carla-Flowguided dataset and 21.3% on the Semi-Truck Highway dataset, offering superior rectification accuracy. Kim-Hui Yap, Weide Liu, Xulei Yang, Jun Cheng 0003 |
CVPR | 2 |
| 2025 | A Structure-Aware and Motion-Adaptive Framework for 3D Human Pose Estimation with Mamba
Jianjun Gao 0005, Kim-Hui Yap |
ICCV | 6 |
| 2025 | Towards Blind Bitstream-corrupted Video Recovery: A Visual Foundation Model-driven FrameworkabstractVideo signals are vulnerable in multimedia communication and storage systems, as even slight bitstream-domain corruption can lead to significant pixel-domain degradation. To recover faithful spatio-temporal content from corrupted inputs, bitstream-corrupted video recovery has recently emerged as a challenging and understudied task. However, existing methods require time-consuming and labor-intensive annotation of corrupted regions for each corrupted video frame, resulting in a large workload in practice. In addition, high-quality recovery remains difficult as part of the local residual information in corrupted frames may mislead feature completion and successive content recovery. In this paper, we propose the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with recovery model, which is adapted to different types of corruption and bitstream-level prompts. Within the framework, the proposed Detect Any Corruption (DAC) model leverages the rich priors of the visual foundation model while incorporating bitstream and corruption knowledge to enhance corruption localization and blind recovery. Additionally, we introduce a novel Corruption-aware Feature Completion (CFC) module, which adaptively processes residual contributions based on high-level corruption understanding. With VFM-guided hierarchical feature augmentation and high-level coordination in a mixture-of-residual-experts (MoRE) structure, our method suppresses artifacts and enhances informative residuals. Comprehensive evaluations show that the proposed method achieves outstanding performance in bitstream-corrupted video recovery without requiring a manually labeled mask sequence. The demonstrated effectiveness will help to realize improved user experience, wider application scenarios, and more reliable multimedia communication and storage systems. Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
ACM Multimedia | 5 |
| 2025 | SSH-Net: A self-supervised and hybrid network for noisy image watermark removal
Wenyang Liu, Jianjun Gao 0005, Kim-Hui Yap |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | CL-HOI: Cross-level human-object interaction distillation from multimodal large language models
Jianjun Gao 0005, Wenyang Liu, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
Knowl. Based Syst. | 5 |
| 2025 | Open World Object Detection: A SurveyabstractExploring new knowledge is a fundamental human ability that can be mirrored in the development of deep neural networks, especially in the field of object detection. Open world object detection (OWOD) is an emerging area of research that adapts this principle to explore new knowledge. It focuses on recognizing and learning from objects absent from initial training sets, thereby incrementally expanding its knowledge base when new class labels are introduced. This survey paper offers a thorough review of the OWOD domain, covering essential aspects, including problem definitions, benchmark datasets, source codes, evaluation metrics, and a comparative study of existing methods. Additionally, we investigate related areas like open set recognition (OSR) and incremental learning (IL), underlining their relevance to OWOD. Finally, the paper concludes by addressing the limitations and challenges faced by current OWOD algorithms and proposes directions for future research. To our knowledge, this is the first comprehensive survey of the emerging OWOD field with over one hundred references, marking a significant step forward for object detection technology. A comprehensive source code and benchmarks are archived and concluded athttps://github.com/ArminLee/OWOD_Review. Yi Wang 0068, Dan Lin 0008, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian TrackingabstractMultiple pedestrian tracking is crucial for enhancing safety and efficiency in intelligent transport and autonomous driving systems by predicting movements and enabling adaptive decision-making in dynamic environments. It optimizes traffic flow, facilitates human interaction, and ensures compliance with regulations. However, it faces the challenge of tracking pedestrians in the presence of occlusion. Existing methods overlook effects caused by abnormal detections during partial occlusion. Subsequently, these abnormal detections can lead to inaccurate motion estimation, unreliable appearance features, and unfair association. To address these issues, we propose an adaptive occlusion-aware multiple pedestrian tracker, OccluTrack, to mitigate the effects caused by partial occlusion. Specifically, we first introduce a plug-and-play abnormal motion suppression mechanism into the Kalman Filter to adaptively detect and suppress outlier motions caused by partial occlusion. Second, we develop a pose-guided re-identification (Re-ID) module to extract discriminative part features for partially occluded pedestrians. Last, we develop a new occlusion-aware association method towards fair Intersection over Union (IoU) and appearance embedding distance measurement for occluded pedestrians. Extensive evaluation results demonstrate that our method outperforms state-of-the-art methods on MOTChallenge and DanceTrack datasets. Particularly, the performance improvements on IDF1 and ID Switches, as well as visualized results, demonstrate the effectiveness of our method in multiple pedestrian tracking. Jianjun Gao 0005, Yi Wang 0068, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | ByteNet: Rethinking Multimedia File Fragment Classification Through Visual PerspectivesabstractMultimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multimedia storage and communication. Existing MFFC methods typically treat fragments as 1D byte sequences and emphasize the relations between separate bytes (interbytes) for classification. However, the more informative relations inside bytes (intrabytes) are overlooked and seldom investigated. By looking inside bytes, the bit-level details of file fragments can be accessed, enabling a more accurate classification. Motivated by this, we first proposeByte2Image, a novel visual representation model that incorporates previously overlooked intrabyte information into file fragments and reinterprets these fragments as 2D grayscale images. This model involves a sliding byte window to reveal the intrabyte information and a rowwise stacking of intrabyte n-grams for embedding fragments into a 2D space. Thus, complex interbyte and intrabyte correlations can be mined simultaneously using powerful vision networks. Additionally, we propose an end-to-end dual-branch networkByteNetto enhance robust correlation mining and feature representation. ByteNet makes full use of the raw 1D byte sequence and the converted 2D image through a shallow byte branch feature extraction (BBFE) and a deep image branch feature extraction (IBFE) network. In particular, the BBFE, composed of a single fully-connected layer, adaptively recognizes the co-occurrence of several some specific bytes within the raw byte sequence, while the IBFE, built on a vision Transformer, effectively mines the complex interbyte and intrabyte correlations from the converted image. Experiments on the two representative benchmarks, including 14 cases, validate that our proposed method outperforms state-of-the-art approaches on different cases by up to 12.2%. Wenyang Liu, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 5 |
| 2024 | CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringabstractDiagram Question Answering (DQA) is a challenging task, requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring, technical support, and more practical applications. DQA poses significant challenges, such as the demand for domain-specific knowledge and the scarcity of annotated data, which restrict the applicability of large-scale deep models. Previous approaches have explored external knowledge integration through pretraining, but these methods are costly and can be limited by domain disparities. While Large Language Models (LLMs) show promise in question-answering, there is still a gap in how to cooperate and interact with the diagram parsing process. In this paper, we introduce the Chain-of-Guiding Learning Model for Diagram Question Answering (CoG-DQA), a novel framework that effectively addresses DQA challenges. CoG-DQA leverages LLMs to guide diagram parsing tools (DPTs) through the guiding chains, enhancing the precision of diagram parsing while introducing rich background knowledge. Our experimental findings reveal that CoG-DQA surpasses all comparison models in various DQA scenarios, achieving an average accuracy enhancement exceeding 5% and peaking at 11% across four datasets. These results underscore CoG-DQA's capacity to advance the field of visual question answering and promote the integration of LLMs into specialized domains. Lingling Zhang 0005, Longji Zhu, Tao Qin 0002, Kim-Hui Yap, Xinyu Zhang 0021, Jun Liu 0002 |
CVPR | 5 |
| 2024 | Empowering Large Language Model for Continual Video Question Answering with Collaborative PromptingabstractIn recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content.In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting.To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting.These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research.Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14% accuracy on NExT-QA and 71.24% accuracy on DramaQA, highlighting its practical relevance and effectiveness. Jianjun Gao 0005, Wenyang Liu, Runzhong Zhang, Kim-Hui Yap |
EMNLP | 7 |
| 2024 | Contextual Human Object Interaction Understanding from Pre-Trained Large Language ModelabstractExisting human object interaction (HOI) detection methods have introduced zero-shot learning techniques to recognize unseen interactions, but they still have limitations in understanding context information and comprehensive reasoning. To overcome these limitations, we propose a novel HOI learning framework, ContextHOI, which serves as an effective contextual HOI detector to enhance contextual understanding and zero-shot reasoning ability. The main contributions of the proposed ContextHOI are a novel context-mining decoder and a powerful interaction reasoning large language model (LLM). The context-mining decoder aims to extract linguistic contextual information from a pre-trained vision-language model. Based on the extracted context information, the proposed interaction reasoning LLM further enhances the zero-shot reasoning ability by leveraging rich linguistic knowledge. Extensive evaluation demonstrates that our proposed framework outperforms existing zero-shot methods on the HICO-DET and SWIG-HOI datasets, as high as 19.34% mAP on unseen interaction can be achieved. Jianjun Gao 0005, Kim-Hui Yap, Kejun Wu, Duc Tri Phan, Kratika Garg, Boon Siew Han |
ICASSP | 2 |
| 2024 | Multi-Modality Action Recognition Based on Dual Feature Shift in Vehicle Cabin MonitoringabstractDriver Action Recognition (DAR) is crucial in vehicle cabin monitoring systems. In real-world applications, it is common for vehicle cabins to be equipped with cameras featuring different modalities. However, multi-modality fusion strategies for the DAR task within car cabins have rarely been studied. In this paper, we propose a novel yet efficient multi-modality driver action recognition method based on dual feature shift, named DFS. DFS first integrates complementary features across modalities by performing modality feature interaction. Meanwhile, DFS achieves the neighbour feature propagation within single modalities, by feature shifting among temporal frames. To learn common patterns and improve model efficiency, DFS shares feature extracting stages among multiple modalities. Extensive experiments have been carried out to verify the effectiveness of the proposed DFS model on the Drive&Act dataset. The results demonstrate that DFS achieves good performance and improves the efficiency of multi-modality driver action recognition. Dan Lin 0008, Philip Hann Yung Lee, Kim-Hui Yap, You Shing Ngim |
ICASSP | 5 |
| 2024 | Hdplifter: Hierarchical Dynamics Perception For 2D-to-3D Human Pose LiftingabstractRecent 2D-to-3D pose lifting networks have achieved remarkable success in monocular 3D human pose estimation through learning joint dependencies. We observed that the extracted 2D pose sequences encountered spatial pose topology ambiguity and temporal movement patterns information loss. Existing methods overlook these intrinsic limitations, resulting in inferior inferences of the corresponding 3D poses. To address these, we introduce the Hierarchical Dynamics Pose Lifter (HDPLifter), which captures spatial human joint connections and subtle temporal movement patterns while maintaining global modeling through hierarchical perception. Specifically, we propose a Structure-aware Spatial Transformer using adaptive topology learning to efficiently integrate spatial joint connections. Moreover, a novel Hierarchical Temporal Transformer is utilized to comprehensively capture subtle joint movement patterns along with global movement patterns with a scaleable receptive field. In both modules, we utilize a 2D depth-wise convolution as a feedforward network to further gather local joint correlations in the spatial and temporal domains simultaneously. Our model, HDPLifter, surpasses the state-of-the-art approach (Motion-BERT) on Human3.6M and MPI-INF-3DHP datasets with P1 errors of 38.0mm and 14.4mm, respectively, while utilizing only 1/5 of the parameters compared to it. Jianjun Gao 0005, Duc Tri Phan, Kim-Hui Yap |
ICIP | 6 |
| 2024 | CM2-Net: Continual Cross-Modal Mapping Network For Driver Action RecognitionabstractDriver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and depth. Nevertheless, compared to RGB modality only, it is always laborious and costly to collect extensive data for all types of non-RGB modalities in car cabin environments. Therefore, previous works have suggested independently learning each non-RGB modality by fine-tuning a model pretrained on RGB videos, but these methods are less effective in extracting informative features when faced with newly-incoming modalities due to large domain gaps. In contrast, we propose a Continual Cross-Modal Mapping Network (CM2Net) to continually learn each newly-incoming modality with instructive prompts from the previously-learned modalities. Specifically, we have developed Accumulative Cross-modal Mapping Prompting (ACMP), to map the discriminative and informative features learned from previous modalities into the feature space of newly-incoming modalities. Then, when faced with newly-incoming modalities, these mapped features are able to provide effective prompts for which features should be extracted and prioritized. These prompts are accumulating throughout the continual learning process, thereby boosting further recognition performances. Extensive experiments conducted on the Drive&Act dataset demonstrate the performance superiority of $\mathrm{CM}^{2}-\mathrm{Net}$ on both uni- and multi-modal driver action recognition. Jianjun Gao 0005, Dan Lin 0008, Wenyang Liu, Kim-Hui Yap |
ICIP | 7 |
| 2024 | Temporal Sentence Grounding with Temporally Global Textual KnowledgeabstractTemporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multimodal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method. Runzhong Zhang, Jianjun Gao 0005, Kejun Wu, Kim-Hui Yap, Yi Wang 0068 |
ICME | 5 |
| 2024 | Multi-scale Attentive Fusion Network for Remote Sensing Image Change CaptioningabstractRemote-sensing Image Change Captioning (RSICC) aims to automatically generate sentences describing the difference of content in remote-sensing bitemporal images. Most of the methods often address shortcomings in model architecture to enhance previous work, overlooking the distinctive characteristics that set remote sensing images apart from natural images, such as recognizing the change of objects with various scales (e.g., small/large-scale objects). By considering the difference, we proposed a Multi-scale Attentive Fusion Network (MAF-Net) to adaptively capture and describe the object change with a wide range of scales. The MAF-Net first extracts multi-scale visual features of bitemporal images from different stages of the CNN backbone, then captures the changes in each pair of the features with the proposed Multi-scale Change Aware Encoders (MCAE). Specifically, the MCAE captures the change-aware discriminative information over the paired multi-scale bitemporal features by Transformer-based different and content cross-attention encoding. Furthermore, a Gated Attentive Fusion (GAF) module is introduced to adaptively aggregate the relevant change-aware features to enhance the change caption performance. We evaluate the effectiveness of our proposed method on two RSICC datasets (e.g., LEVIR-CC and LEVIRCCD), and experimental results demonstrate that our method achieves state-of-the-art performance. Yi Wang 0068, Kim-Hui Yap |
ISCAS | 3 |
| 2024 | MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action RecognitionabstractDriver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging, being gloomy and dark, and with the development of sensors, various cameras such as IR and depth cameras have emerged for analyzing drivers' behaviors. Therefore, in this paper, we propose a novel multimodal fusion transformer, named Multi-Fuser, which identifies cross-modal interrelations and interactions among multimodal car cabin videos and adaptively integrates different modalities for improved representations. Specifically, MultiFuser comprises layers of Bi-decomposed Modules to model spatiotemporal features, with a modality synthesizer for multi-modal features integration. Each Bi-decomposed Module includes a Modal Expertise ViT block for extracting modality-specific features and a Patch-wise Adaptive Fusion block for efficient cross-modal fusion. Extensive experiments are conducted on Drive&Act dataset and the results demonstrate the efficacy of our proposed approach. Jianjun Gao 0005, Dan Lin 0008, Kim-Hui Yap |
MMSP | 5 |
| 2024 | Image compressive sensing reconstruction via nonlocal low-rank residual-based ADMM framework
Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu |
Comput. Vis. Image Underst. | 2 |
| 2024 | Top-down framework for weakly-supervised grounded image captioning
Suchen Wang, Kim-Hui Yap, Yi Wang 0068 |
Knowl. Based Syst. | 3 |
| 2024 | Intra- and inter-sector contextual information fusion with joint self-attention for file fragment classificationabstractFile fragment classification (FFC) aims to identify the file type of file fragments in memory sectors, which is of great importance in memory forensics and information security. Existing works focused on processing the bytes within sectors separately and ignoring contextual information between adjacent sectors. In this paper, we introduce a joint self-attention network (JSANet) for FFC to learn intra-sector local features and inter-sector contextual features. Specifically, we propose an end-to-end network with the byte, channel, and sector self-attention modules. Byte self-attention adaptively recognizes the intra-sector significant bytes, and channel self-attention re-calibrates the features between channels. Based on the insight that adjacent memory sectors are most likely to store a file fragment, sector self-attention leverages contextual information in neighboring sectors to enhance inter-sector feature representation. Extensive experiments on seven FFC benchmarks show the superiority of our method compared with state-of-the-art methods. Moreover, we construct VFF-16, a variable-length file fragment dataset to reflect file fragmentation. Integrated with sector self-attention, our method improves accuracy by more than 16.3% against the baseline on VFF-16, and the runtime achieves 5.1 s/GB with GPU acceleration. In addition, we extend our model to malware detection and show its applicability. Yi Wang 0068, Wenyang Liu, Kejun Wu, Kim-Hui Yap, Lap-Pui Chau |
Knowl. Based Syst. | 4 |
| 2024 | Dense Supervision Propagation for Weakly Supervised Semantic Segmentation on 3D Point CloudsabstractSemantic segmentation on 3D point clouds is an important task for 3D scene understanding. While dense labeling on 3D data is expensive and time-consuming, only a few works address weakly supervised semantic point cloud segmentation methods to relieve the labeling cost by learning from simpler and cheaper labels. Meanwhile, there are still huge performance gaps between existing weakly supervised methods and state-of-the-art fully supervised methods. In this paper, we propose Dense Supervision Propagation (DSP) to train a semantic point cloud segmentation network with only a small portion of points being labeled. We argue that we can better utilize the limited supervision information as we densely propagate the supervision signal from the labeled points to other points within and across the input samples. Specifically, we propose a cross-sample feature reallocating module to transfer similar features and therefore re-route the gradients across two samples with common classes and an intra-sample feature redistribution module to propagate supervision signals on unlabeled points across and within point cloud samples. We conduct extensive experiments on public datasets S3DIS and ScanNet. Our weakly supervised method with only 10% and 1% of labels can produce competitive results with the fully supervised counterpart. Jiacheng Wei, Guosheng Lin, Kim-Hui Yap, Fayao Liu, Tzu-Yi Hung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Bitstream-Corrupted JPEG Images are Restorable: Two-stage Compensation and Alignment Framework for Image RestorationabstractIn this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be resolved by existing image restoration methods mainly relying on pre-defined degradation models in the pixel domain. To address these challenges, we propose a robust JPEG decoder, followed by a two-stage compensation and alignment framework to restore bitstream-corrupted JPEC images. Specifically, the robust JPEC decoder adopts an error-resilient mechanism to decode the corrupted JPEG bitstream. The two-stage framework is composed of the self-compensation and alignment (SCA) stage and the guided-compensation and alignment (GCA) stage. The SCA adaptively performs block-wise image color compensation and alignment based on the estimated color and block offsets via image content similarity. The GCA leverages the extracted low-resolution thumbnail from the JPEG header to guide full-resolution pixel-wise image restoration in a coarse-to-fine manner. It is achieved by a coarse-guided pix2pix network and a refine-guided bi-directional Laplacian pyramid fusion network. We conduct experiments on three benchmarks with varying degrees of bit error rates. Experimental results and ablation studies demonstrate the superiority of our proposed method. The code will be released at https://github.com/wenyang001/Two-ACIR. Wenyang Liu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
CVPR | 3 |
| 2023 | TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo SupervisionabstractIn this paper, we investigate an open research task of generating controllable 3D textured shapes from the given textual descriptions. Previous works either require ground truth caption labeling or extensive optimization time. To resolve these issues, we present a novel framework, TAPS3D, to train a text-guided 3D shape generator with pseudo captions. Specifically, based on rendered 2D images, we retrieve relevant words from the CLIP vocabulary and construct pseudo captions using templates. Our constructed captions provide high-level semantic supervision for generated 3D shapes. Further, in order to produce fine-grained textures and increase geometry diversity, we propose to adopt low-level image regularization to enable fake-rendered images to align with the real ones. During the inference phase, our proposed model can generate 3D textured shapes from the given text without any additional optimization. We conduct extensive experiments to analyze each of our proposed components and show the efficacy of our framework in generating high-fidelity 3D textured and text-relevant shapes. Code is available at https://github.com/plusmultiply/TAPS3D Jiacheng Wei, Hao Wang 0094, Jiashi Feng, Guosheng Lin, Kim-Hui Yap |
CVPR | 5 |
| 2023 | A Spatial-Focal Error Concealment Scheme for Corrupted Focal Stack VideoabstractFocal stack image sequences can be regarded as successive frames of videos, which are densely captured by focusing on a stack of focal planes. This type of data is able to provide focus cues for display technologies. Before the displays on the user side, focal stack video is possibly corrupted during compression, storage and transmission chains, generating error frames on the decoder side. The error regions are difficult to be recovered due to the focal changes among frames. Conventional error concealment methods result in sharpness inconsistency between recovered regions and their spatial adjacent regions. Motivated by this, in this paper, we propose a spatial-focal error concealment scheme specialized for focal stack videos. The spatial adjacent regions around an error region are employed to reveal the prediction relations between error frame and focal adjacent frames. Gaussian blur filtering and Lucy-Richardson deblur filtering are applied to simulate the video focal changes. In this way, the error regions can be well recovered by exploiting the spatial-focal information. Experiment results show that the proposed scheme can achieve the highest objective quality in terms of PSNR and SSIM. It can also obtain the best subjective quality with sharpness consistency in recovered regions and without block effect. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
DCC | 4 |
| 2023 | Nonlocal Low-Rank Residual Modeling for Image Compressive Sensing ReconstructionabstractThe nonlocal low-rank (LR) modeling has proven to be an effective approach in image compressive sensing (CS) reconstruction, which starts by clustering similar patches using the nonlocal self-similarity (NSS) prior into nonlocal image groups and then imposes an L-R penalty on each nonlocal image group. However, most existing methods only approximate the LR matrix directly from the degraded nonlocal image group, which may lead to suboptimal LR matrix approximation and thus obtain unsatisfactory reconstruction results. This paper proposes a novel nonlocal low-rank residual (NLRR) approach for image CS reconstruction, which progressively approximates the underlying LR matrix by minimizing the LR residual. To do this, we first use the NSS prior to obtain a good estimate of the original nonlocal image group, and then the LR residual between the degraded nonlocal image group and the estimated nonlocal image group is minimized to derive a more accurate LR matrix. To ensure the optimization is both feasible and reliable, we employ an alternative direction multiplier method (ADMM) to solve the NLRR-based image CS reconstruction problem. Our experimental results show that the proposed NLRR algorithm achieves superior performance against many popular or state-of-the-art image CS reconstruction methods, both in objective metrics and subjective perceptual quality. Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu |
ICIP | 2 |
| 2023 | Image Representation and Deep Inception-Attention for File-type and Malware ClassificationabstractFile-type classification aims to recognize the file types of files/fragments without file-system metadata, which is essential for memory forensics and data recovery. In this paper, we introduce an image representation and deep inception-attention manner for file-type classification. Specifically, we consider file-type classification as an image classification problem. Raw data sequences in the memory block are converted to 2D binary images, enriching the representation ability and visualization while retaining the completeness of the bitstream. With binary images as inputs, we propose a deep inception-attention network to extract discriminate horizontal features and re-calibrate the weights of feature maps, and finally, predict file types. Experiments on a large-scale benchmark show the superiority of the proposed model. Moreover, our method can be extended to a similar application, like malware classification, and achieve outstanding performance. Yi Wang 0068, Kejun Wu, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
ISCAS | 4 |
| 2023 | METFormer: A Motion Enhanced Transformer for Multiple Object TrackingabstractMultiple object tracking (MOT) is an important task in computer vision, especially video analytics. Transformer-based methods are emerging approaches using both tracking and detection queries. However, motion modeling in existing transformer-based methods lacks effective association capability. Thus, this paper introduces a new METFormer model, a Motion Enhanced TransFormer-based tracker with a novel global-local motion context learning technique to mitigate the lack of motion information in existing transformer-based methods. The global-local motion context learning technique first centers on difference-guided global motion learning to obtain temporal information from adjacent frames. Based on global motion, we leverage context-aware local object motion modelling to study motion patterns and enhance the feature representation for individual objects. Experimental results on the benchmark MOT17 dataset show that our proposed method can surpass the state-of-the-art Trackformer [21] by 1.8% on IDF1 and 21.7% on ID Switches under public detection settings. Jianjun Gao 0005, Kim-Hui Yap, Yi Wang 0068, Kratika Garg, Boon Siew Han |
ISCAS | 2 |
| 2023 | Vision-Based Early Fire and Smoke Detection for Smart Factory Applications Using FFS-YOLOabstractEarly-stage fire and smoke detection through visual analysis is crucial for industrial safety and hazard prevention. However, detecting fire and smoke in factories using surveillance cameras poses challenges due to the small size of target objects. To address these challenges, we introduce a refined single-stage detector called FFS-YOLO (Factory Fire Smoke - YOLO). Our approach incorporates the Parameter-Free Attention Module (SimAM) and ResNet-SimMix module into the Backbone and Head of YOLOv7 to enhance key feature extraction. Additionally, we modify the model architecture by adding an extra prediction head to facilitate the fusion of features at multiple scales, specifically for small-scale object detection. Experimental results conducted on our fire and smoke dataset demonstrate the effectiveness of the FFS-YOLO model, achieving an average mAP, Precision, and Recall of 0.92, 0.91, and 0.90, respectively. The performance of the proposed model outperforms existing relevant competitors in the field. The findings of this research contribute significantly to the advancement of early fire detection and prevention in factory settings. Duc Tri Phan, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
MMSP | 2 |
| 2023 | Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and MethodabstractThe past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communication (e.g., telepresence, live streaming, and internet video) and multimedia forensics. To address this, we introduce the bitstream-corrupted video (BSCV) benchmark, the first benchmark dataset with more than 28,000 video clips, which can be used for bitstream-corrupted video recovery in the real world. The BSCV is a collection of 1) a proposed three-parameter corruption model for video bitstream, 2) a large-scale dataset containing rich error patterns, multiple corruption levels, and flexible dataset branches, and 3) a new video recovery framework that serves as a benchmark. We evaluate state-of-the-art video inpainting methods on the BSCV dataset, demonstrating existing approaches' limitations and our framework's advantages in solving the bitstream-corrupted video recovery problem. The benchmark and dataset are released at https://github.com/LIUTIGHE/BSCV-Dataset. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
NeurIPS | 5 |
| 2023 | Reconciliation of statistical and spatial sparsity for robust visual classification
Hao Cheng 0016, Kim-Hui Yap, Bihan Wen |
Neurocomputing | 2 |
| 2023 | SSN: Stockwell Scattering Network for SAR Image Change DetectionabstractRecently, synthetic aperture radar (SAR) image change detection has become an interesting yet challenging direction due to the presence of speckle noise. Although both traditional and modern learning-driven methods attempted to overcome this challenge, deep convolutional neural networks (DCNNs)-based methods are still hindered by the lack of interpretability and the requirement of large computation power. To overcome this drawback, wavelet scattering network (WSN) and Fourier scattering network (FSN) are proposed. Combining respective merits of WSN and FSN, we propose Stockwell scattering network (SSN) based on Stockwell transform (ST), which is widely applied against noisy signals and shows advantageous characteristics in speckle reduction. The proposed SSN provides noise-resilient feature representation and obtains state-of-the-art performance in SAR image change detection as well as high computational efficiency. Experimental results on three real SAR image datasets demonstrate the effectiveness of the proposed method. Gong Chen 0003, Yanan Zhao 0003, Yi Wang 0068, Kim-Hui Yap |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionabstractIt is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi. Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan 0001 |
CVPR | 5 |
| 2022 | Attribute Conditioned Fashion Image CaptioningabstractFashion image captioning aims to automatically generate product descriptions for fashion items. Existing fashion image captioning models predict a fixed caption for a particular fashion item once deployed. Differently, we explore a controllable way of fashion image captioning by taking the semantic attributes as the control signal. We propose a new multi-modal fashion image captioning method that allows the users to specify a few semantic attributes to guide the caption generation to suit unique preferences. To study the problem, we clean, filter, and assemble a new fashion image caption dataset called FACAD170K from the current FACAD dataset to facilitate learning. We investigate the effectiveness of the proposed approach on our assembled FACAD170K dataset. The results demonstrate that our method can outperform existing fashion image captioning models as well as conventional captioning methods. Kim-Hui Yap, Suchen Wang |
ICIP | 2 |
| 2022 | Mixed Membership Generative Adversarial NetworksabstractGANs are designed to learn a single distribution, though multiple distributions can be modeled by treating them separately. However, this naive implementation does not consider overlapping distributions. We propose Mixed Membership Generative Adversarial Networks (MMGAN) analogous to mixed-membership models that model multiple distributions and discover their commonalities and particularities. Each data distribution is modeled as a mixture over a common set of generator distributions, and mixture weights are automatically learned from the data. Mixture weights can give insight into common and unique features of each data distribution. We evaluate our proposed MMGAN and show its effectiveness on MNIST and Fashion-MNIST with various settings. Yasin Yazici, Bruno Lecouat, Kim-Hui Yap, Stefan Winkler 0001, Georgios Piliouras, Vijay Chandrasekhar 0001, Chuan-Sheng Foo |
ICIP | 3 |
| 2022 | Collaborative learning mutual network for domain adaptation in person re-identification
Chiat-Pin Tay, Kim-Hui Yap |
Neural Comput. Appl. | 2 |
| 2021 | A Compact Joint Distillation Network for Visual Food RecognitionabstractVisual food recognition is emerging as an important application in dietary monitoring and management in recent years. Existing works use large backbone networks to achieve good performance. However, these networks are not able to be deployed on personal portable devices due to large size and computation cost. Some compact networks have been developed, however, their performance are usually lower than the large backbone networks. In view of this, this paper proposes a joint distillation framework that targets to achieve a high visual food recognition accuracy using a compact network. As opposed to the more traditional one-directional knowledge distillation methods, the proposed knowledge distillation framework trains both the large teacher network and the compact student network simultaneously. The framework introduces a new Multi-Layer Distillation (MLD) for simultaneous teacher-student learning at multiple layers of different abstraction. A novel Instance Activation Mapping (IAM) is proposed to jointly train the teacher and student networks using generated instance-level activation map that incorporates label information for each training image. Experimental results on the two benchmark datasets UECFood-256 and Food-101 show that the trained compact student network achieves state-of-the-art performance at 83.5% and 90.4%, respectively, while achieving more than 4 times deduction regarding network model size. Zhao Heng, Kim-Hui Yap, Alex Chichung Kot |
ICASSP | 2 |
| 2021 | Discovering Human Interactions with Large-Vocabulary Objects via Query and Multi-Scale DetectionabstractIn this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and interaction classification due to the increasing diversity of objects (e.g., 1000 categories). Different from previous methods, we formulate the HOI detection as a query problem. We propose a unified model to jointly discover the target objects and predict the corresponding interactions based on the human queries, thereby eliminating the need of using generic object detectors, extra steps to associate human-object instances, and multi-stream interaction recognition. This is achieved by a repurposed Transformer unit and a novel cascade detection over multi-scale feature maps. We observe that such a highly-coupled solution brings benefits for both object detection and interaction classification in a large vocabulary setting. To study the new challenges of the large vocabulary HOI detection, we assemble two datasets from the publicly available SWiG and 100 Days of Hands datasets. Experiments on these datasets validate that our proposed method can achieve a notable mAP improvement on HOI detection with a faster inference speed than existing one-stage HOI detectors. Our code is available at https://github.com/scwangdyd/large_vocabulary_hoi_detection. Suchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu, Junsong Yuan 0001, Yap-Peng Tan |
ICCV | 2 |
| 2021 | APNET: Attribute Parsing Network for Person Re-IdentificationabstractMost person re-identification methods rely solely on the pedestrian identity for learning. Person attributes, such as gender, clothing colors, carried bags, etc, are however seldom used. These attributes are highly identity-related and should be capitalized fully. Thus, we propose Attribute Parsing Network (APNet), an architecture designed for both image and person attribute learning and retrievals. To further enhance the re-id performance, we propose to leverage saliency maps and human parsing to boost the foreground features, which when trained together with the global and local networks, resulted in more generic and robust encoded representations. This proposed method achieved state-of-the-art accuracy performance on both Market1501 (87.3% mAP and 95.2% Rank1) and DukeMTMC-reID (78.8% mAP and 89.2% Rank 1) datasets. Chiat-Pin Tay, Kim-Hui Yap |
ICIP | 2 |
| 2021 | Fusion Learning using Semantics and Graph Convolutional Network for Visual Food RecognitionabstractFood-related applications and services are essential for the health and well-being of people. With the rapid development of social networks and mobile devices, food images captured by people can offer rich knowledge about the food and also necessary dietary assistance for people that require special care. Known food recognition frameworks and approaches in computer vision have heavy reliance on many-shot training of a deep network on existing large-scale food datasets. However, it is common for many food categories that it is difficult to collect enough images for training. Traditional few-shot learning is unable to properly address the problem due to the complex characteristics and large variations of food images, and most few-shot frame-works cannot perform classification for many-shot and few-shot categories at the same time. In this paper, we propose a new fusion learning framework for food recognition. It unifies many-shot and few-shot under a single framework, by leveraging on extracted image representations and context sensitive semantic embeddings. Further, considering food categories are often correlated to each other for many commonalities such as same ingredients, cooking methods, the fusion learning framework utilizes a Graph Convolutional Network (GCN) to capture the inter-class relations between both image representations and semantic embeddings of different food categories. The final output fusion classifier will be more robust and discriminative. Comprehensive experimental results on two popular food benchmarks have shown the proposed framework achieves the state-of-the-art fusion performance. Kim-Hui Yap, Alex Chichung Kot |
WACV | 2 |
| 2021 | Attribute saliency network for person re-identification
Chiat-Pin Tay, Kim-Hui Yap |
Image Vis. Comput. | 2 |
| 2020 | Discovering Human Interactions With Novel Objects via Zero-Shot LearningabstractWe aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection task. The core idea is to leverage human visual clues to localize objects which are interacting with humans. We show that our proposed model can outperform existing methods on detecting interacting objects, and generalize well to novel objects. To recognize objects from unseen categories, we devise a zero-shot classification module upon the classifier of seen categories. It utilizes the classifier logits for seen categories to estimate a vector in the semantic space, and then performs nearest search to find the closest unseen category. We validate our method on V-COCO and HICO-DET datasets, and obtain superior results on detecting human interactions with both seen and unseen objects. Suchen Wang, Kim-Hui Yap, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 2 |
| 2020 | Multi-Path Region Mining for Weakly Supervised 3D Semantic Segmentation on Point CloudsabstractPoint clouds provide intrinsic geometric information and surface context for scene understanding. Existing methods for point cloud segmentation require a large amount of fully labeled data. Using advanced depth sensors, collection of large scale 3D dataset is no longer a cumbersome process. However, manually producing point-level label on the large scale dataset is time and labor-intensive. In this paper, we propose a weakly supervised approach to predict point-level results using weak labels on 3D point clouds. We introduce our multi-path region mining module to generate pseudo point-level labels from a classification network trained with weak labels. It mines the localization cues for each class from various aspects of the network feature using different attention modules. Then, we use the point-level pseudo label to train a point cloud segmentation network in a fully supervised manner. To the best of our knowledge, this is the first method that uses cloud-level weak labels on raw 3D space to train a point cloud semantic segmentation network. In our setting, the 3D weak labels only indicate the classes that appeared in our input sample. We discuss both scene- and subcloud-level weakly labels on raw 3D point cloud data and perform in-depth experiments on them. On ScanNet dataset, our result trained with subcloud-level labels is compatible with some fully supervised methods. Jiacheng Wei, Guosheng Lin, Kim-Hui Yap, Tzu-Yi Hung, Lihua Xie 0001 |
CVPR | 3 |
| 2020 | What Does Plate Glass Reveal About Camera Calibration?abstractThis paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image contents. To reduce the impact of a noisy calibration cue estimated from a reflection-contaminated image, we propose two strategies: an optimization-based method that imposes part of though reliable entries on the map and a learning-based method that fully exploits all entries. We collect a dataset containing 320 samples as well as their camera parameters for evaluation. We demonstrate that our method not only facilitates a general single image camera calibration method that leverages image contents but also contributes to improving the performance of single image reflection removal. Furthermore, we show our byproduct output helps alleviate the ill-posed problem of estimating the panorama from a single image. Jinnan Chen, Zhan Lu, Boxin Shi, Xudong Jiang 0001, Kim-Hui Yap, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 6 |
| 2020 | Dynamically Modulated Deep Metric Learning for Visual SearchabstractThis paper proposes dynamically modulated metric learning (DMML) for learning a tiered similarity space to perform visual search. Existing methods often treat the training samples having different degree of information with equal importance which hinders in capturing the underlying granularities in visual similarity. Proposed DMML automatically exploits the informativeness of samples during training by leveraging correlation between image attributes and embedding that are learned jointly. The two tasks are interlinked by supervising signals where the predicted attribute vectors are used to dynamically learn the loss function. To this end, we propose a new soft-binomial deviance loss that helps to capture the feature similarity space at multiple granularities. Compared to recent ensemble and attention based methods, our DMML framework is conceptually simple yet effective, and achieves state-of-the-art performances on standard benchmark datasets; e.g. an improvement of 4% Recall@1 over the SOTA [1] on DeepFashion dataset. Dipu Manandhar, Muhammet Bastan, Kim-Hui Yap |
ICASSP | 3 |
| 2020 | Empirical Analysis Of Overfitting And Mode Drop In Gan TrainingabstractWe examine two key questions in GAN training, namely overfitting and mode drop, from an empirical perspective. We show that when stochasticity is removed from the training procedure, GANs can overfit and exhibit almost no mode drop. Our results shed light on important characteristics of the GAN training procedure. They also provide evidence against prevailing intuitions that GANs do not memorize the training set, and that mode dropping is mainly due to properties of the GAN objective rather than how it is optimized during training. Yasin Yazici, Chuan-Sheng Foo, Stefan Winkler 0001, Kim-Hui Yap, Vijay Chandrasekhar 0001 |
ICIP | 4 |
| 2020 | Semantic granularity metric learning for visual search
Dipu Manandhar, Muhammet Bastan, Kim-Hui Yap |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | Remote detection of idling cars using infrared imaging and deep networks
Muhammet Bastan, Kim-Hui Yap, Lap-Pui Chau |
Neural Comput. Appl. | 2 |
| 2019 | AANet: Attribute Attention Network for Person Re-IdentificationsabstractThis paper proposes Attribute Attention Network (AANet), a new architecture that integrates person attributes and attribute attention maps into a classification framework to solve the person re-identification (re-ID) problem. Many person re-ID models typically employ semantic cues such as body parts or human pose to improve the re-ID performance. Attribute information, however, is often not utilized. The proposed AANet leverages on a baseline model that uses body parts and integrates the key attribute information in an unified learning framework. The AANet consists of a global person ID task, a part detection task and a crucial attribute detection task. By estimating the class responses of individual attributes and combining them to form the attribute attention map (AAM), a very strong discriminatory representation is constructed. The proposed AANet outperforms the best state-of-the-art method \cite{Sun_2018_ECCV} using ResNet-50 by 3.36\% in mAP and 3.12\% in Rank-1 accuracy on DukeMTMC-reID dataset. On Market1501 dataset, AANet achieves 92.38\% mAP and 95.10\% Rank-1 accuracy with re-ranking, outperforming~\cite{kalayeh2018human}, another state of the art method using ResNet-152, by 1.42\% in mAP and 0.47\% in Rank-1 accuracy. In addition, AANet can perform person attribute prediction (e.g., gender, hair length, clothing length etc.), and localize the attributes in the query image. Chiat-Pin Tay, Sharmili Roy, Kim-Hui Yap |
CVPR | 3 |
| 2019 | The Unusual Effectiveness of Averaging in GAN Training
Yasin Yazici, Chuan-Sheng Foo, Stefan Winkler 0001, Kim-Hui Yap, Georgios Piliouras, Vijay Chandrasekhar 0001 |
ICLR (Poster) | 4 |
| 2019 | Convolutional Three-Stream Network Fusion for Driver Fatigue Detection from Infrared VideosabstractWe propose a convolutional three-stream network architecture for driver fatigue detection from infrared videos that are available both in the daytime and in the night time. Specifically, the convolutional three-stream network architecture incorporates current-infrared-frame-based spatial information, optical-flows-based short-term temporal information of two consecutive infrared frames and optical flow-motion history image-based (OF-MHI-based) temporal information within the infrared video sequence. And then these three networks are fused at the last convolutional layer by 3D CNN. Besides, an estimation method to evaluate the current driver fatigue level is proposed based on the fatigue detection results from previous frames, which helps to generate alerts properly in real-life driving applications. We show that the proposed method achieves state-of-the-art performance, 94.68% accuracy, in our driver behavior dataset using the infrared data. Xiaoxi Ma, Lap-Pui Chau, Kim-Hui Yap, Guiju Ping |
ISCAS | 3 |
| 2019 | Multitask Person Re-Identification using Homoscedastic Uncertainty LearningabstractIn this paper, we propose a new multitask neural network called Part Attribute Loss Net (PALNet) for person re-identification (re-id) with homoscedastic uncertainty learning. Currently, many person re-id algorithms use person identity as the main ground truth information to train and to perform prediction task. This single task approach is simple to setup but usually provides poor generalization performance. Some other person re-id works [1] [2] incorporate additional cues such as body parts to improve the learned representations. However, this requires additional body part annotations. Person attributes, which are present in some of the large person re-id datasets like Market1501 [3] and DukeMTMC-reID [4], are not used. Furthermore, the multitask networks found in some of the recent works use heuristic approach in weighing their task losses. Our PALNet seeks to address these issues by leveraging on person identity classification, body part detection and person attribute prediction, all built into a unified network. We also incorporate the homoscedastic uncertainty learning to automatically determine the weighing of the loss function of the different tasks. The PALNet consists of three tasks, with person identity classification as the main task. The body part detection and person attributes are two sub tasks to share different learning experience with the main task. The uncertainty learning provide us good indication of the observable task noises and allows us to optimize the accuracy performance without the time consuming grid search method. When benchmarking on DukeMTMC-reID dataset [4], our approach outperforms other state-of-the-art methods. On Market1501 dataset [3], we are on par with the best state-of-the-art method but outperforms the rest by at least 6.4% mAP and 3.5% Rank-1 accuracy. These results are reported without re-ranking. Chiat-Pin Tay, Sharmili Roy, Kim-Hui Yap |
ISCAS | 3 |
| 2019 | Few-Shot and Many-Shot Fusion Learning in Mobile Visual Food RecognitionabstractMobile visual food recognition is emerging as an important application in food logging and dietary monitoring in recent years. Existing food recognition methods use conventional many-shot learning to train a large backbone network, which refers to the use of sufficient number of training data to train the network. However, these methods firstly do not consider the cases where certain food categories have limited training data. Therefore, they cannot use the conventional training using many-shot learning. Further, existing solutions focus on improving the food recognition performance by implementing state-of-the-art large full networks, and do not pay much attention to reduce the size and computational cost of the network. As a result, they are not amenable for deployment on mobile devices. In this paper, we address these issues by proposing a new few-shot and many-shot fusion learning for mobile visual food recognition, it has a compact framework and is able to learn from existing dataset categories, and also new food categories given only a few sample images. We construct a new Indian food dataset called NTU-IndianFood107 in order to evaluate the performance of the proposed method. The dataset has two parts: (i) a Base Dataset of 83 classes of Indian food images with over 600 images per class to perform many-shot learning, and (ii) a Food Diary of 24 classes captured in restaurants with limited number to simulate the few-shot learning on new food categories. The proposed fusion method achieves a Top-1 classification accuracy of 72.0% on the new dataset. Kim-Hui Yap, Alex Chichung Kot, Ling-Yu Duan, Ngai-Man Cheung |
ISCAS | 2 |
| 2018 | QL-Net: Quantized-by-LookUp CNNabstractConvolutional Neural Networks (CNNs) have achieved a state-of-the-art performance in the different computer vision tasks. However, CNN algorithms are computationally and power intensive, which makes them difficult to run on wearable and embedded systems. One way to address this constraint is to reduce the number of computational operations performed. Recently, several approaches addressed the problem of the computational complexity in the CNNs. Most of these methods, however, require a dedicated hardware. We propose a new method for the computation reduction in CNNs that substitutes Multiply and Accumulate (MAC) operations with a codebook lookup and can be executed on the generic hardware. The proposed method called QL-Net combines several concepts: (i) a codebook construction, (ii) a layer-wise retraining strategy, and (iii) a substitution of the MAC operations with the lookup of the convolution responses at inference time. The proposed QL-Net achieves a 98.6% accuracy on the MNIST dataset with a 5.8x reduction in runtime, when compared to MAC-based CNN model that achieved a 99.2% accuracy. Kamila Abdiyeva, Kim-Hui Yap, Gang Wang 0012, Narendra Ahuja, Martin Lukac |
ICARCV | 2 |
| 2018 | Idling Car Detection with ConvNets in Infrared Image SequencesabstractWe propose a system to detect and localize idling cars in infrared (IR) image sequences for law enforcement to reduce vehicular emission. To this end, we leverage the differences in spatio-temporal heat signatures of idling and stopped cars and monitor car temperatures with a long-wavelength IR camera. We collected a dataset by recording IR image sequences of cars in car parks and trained a ConvNet-based car detector to localize stationary cars in the IR sequences, by utilizing transfer learning and models pre-trained on regular RGB/grayscale images. Then, we used ConvNets with a 3D stack of cropped frames as input to model the spatio-temporal evolution of car temperature over time and detect idling cars. We present promising experimental results on our IR image dataset. Muhammet Bastan, Kim-Hui Yap, Lap-Pui Chau |
ISCAS | 2 |
| 2018 | Hybrid Supervised Deep Learning for Ethnicity Classification using Face ImagesabstractEthnicity information is an integral part of human identity, and a useful identifier for various applications ranging from video surveillance, targeted advertisement to social media profiling. In recent years, Convolutional Neural Networks (CNNs) have shown state-of-the-art performance in many visual recognition problems. Currently, there are a few CNN-based approaches on ethnicity classification [1], [2]. However, the approaches suffer from the following limitations: (i) most face datasets do not include ethnicity information, and those with ethnicity information are typically small to medium in size, thereby they do not provide sufficient samples for training of CNNs from the scratch, and (ii) the CNN methods often treat ethnicity classification as a multi-class classification where the likelihood of each class label is generated. However, it does not utilize the intermediate activation functions of CNNs which provide rich hierarchical features to assist in ethnicity classification. In view of this, this paper proposes a new hybrid supervised learning method to perform ethnicity classification that uses both the strength of CNN as well as the rich features obtained from the network. The method combines the soft likelihood of CNN classification output with an image ranking engine that leverages on matching of the hierarchical features between the query and dataset images. A supervised Support Vector Machine (SVM) hybrid learning is developed to train the combined feature vectors to perform ethnicity classification. The performance of the proposed method is evaluated using a dataset consisting of Bangladeshi, Chinese and Indian ethnicity groups, and it outperforms the state-of-the-art methods [2], [3] by up to 3% in recognition accuracy. Zhao Heng, Dipu Manandhar, Kim-Hui Yap |
ISCAS | 3 |
| 2018 | Brand-Aware Fashion Clothing Search using CNN Feature Encoding and Re-rankingabstractBrand plays a significant role in fashion clothing. Consumers are brand conscious during the clothing search and purchase. Existing visual fashion search methods [1], [8], [17]-[19] often do not explicitly consider the brand information such as logos. Brand logo in clothing images are quite small and often suffer various deformations, and hence pose a significant challenge for branded clothing search. In view of this, this paper presents a new Brand-Aware Fashion Search (BAFS) framework that explores the brand information during the visual search. We construct a new brand fashion dataset which consists of 10K images of branded clothing with trademark logos. The proposed framework first jointly detects the brand logo and clothing items in the images. Next, to extract the rich visual information from the clothing images, we propose a new deep feature encoding known as Principal Component Maximum Activation of Convolutions (PMAC) that leverages hierarchies of CNN activations. The PMAC feature aims to capture both low-level visual information and high-level global abstraction from the images. The proposed method further uses a brand-aware re-ranking technique to improve the search. Experiments conducted on the brand fashion dataset shows that the proposed framework achieves superior performance to other comparative methods. Dipu Manandhar, Kim-Hui Yap, Muhammet Bastan, Zhao Heng |
ISCAS | 2 |
| 2017 | Laplace gradient based Discriminative and Contrast Invertible descriptorabstractThe performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradient based histogram is proposed. Moreover, a robust contrast flipping estimate is proposed based on the divergence of a local region. Experiments on fine-grained object recognition and retrieval applications demonstrate the superior performance of the DCI descriptor to others. Zhenwei Miao, Kim-Hui Yap, Xudong Jiang 0001, Subbhuraam Sinduja |
ICASSP | 2 |
| 2017 | Integrated 3D feature augmentation and view selection in commercial product searchabstractThis paper presents a new integrated 3D feature augmentation and view selection framework for commercial product search. A major challenge in 3D object search is that the object can be captured from different viewpoints and this might result incorrect matching. A solution to this problem is to include every possible view of an object in the database. However this will increase the storage requirement and processing time unnecessarily. One solution to this issue is to include the most important views of an object in the database. This requires careful observation on the significance of the views since not every viewpoint carry relevant information. Existing product databases mainly focus on gathering as many views as possible without analyzing the significance of these views. In view of this, this paper proposes an Integrated Feature augmentation and View selection (IFV) framework that aims to automatically select the salient views to achieve minimal storage requirement and processing time. Anu Susan Skaria, Kim-Hui Yap |
ICIP | 2 |
| 2017 | Lattice-Support repetitive local feature detection for visual search
Dipu Manandhar, Kim-Hui Yap, Zhenwei Miao, Lap-Pui Chau |
Pattern Recognit. Lett. | 2 |
| 2017 | Feature Repetitiveness Similarity Metrics in Visual SearchabstractRepetitive patterns are significant visual cues for matching and detecting objects in images. However, information from repetitive patterns is underutilized in computer vision algorithms as they cause burstiness issue in image similarity scoring. Existing similarity metrics do not take the repetitive patterns into account, and hence cannot handle images with repetitive patterns well. In view of this, this letter presents a new feature repetitiveness similarity (FRS) metric that not only addresses the burstiness issue, but also uses the information from the repetitive patterns to enhance the retrieval performance. The proposed FRS framework detects the repetitive patterns using descriptor and geometric information of local features in the images. Unique and repetitive features are handled separately and then fused at scoring stage using the FRS metric. Experiments conducted on the benchmark Oxford and Paris datasets show that the proposed method outperforms the state-of-the-art methods by a mean average precision of 6%. This demonstrates the effectiveness of FRS metric in matching and retrieval of images with repeated patterns. Dipu Manandhar, Kim-Hui Yap |
IEEE Signal Process. Lett. | 2 |
| 2016 | Contrast Invariant Interest Point Detection by Zero-Norm LoG FilterabstractThe Laplacian of Gaussian (LoG) filter is widely used in interest point detection. However, low-contrast image structures, though stable and significant, are often submerged by the high-contrast ones in the response image of the LoG filter, and hence are difficult to be detected. To solve this problem, we derive a generalized LoG filter, and propose a zero-norm LoG filter. The response of the zero-norm LoG filter is proportional to the weighted number of bright/dark pixels in a local region, which makes this filter be invariant to the image contrast. Based on the zero-norm LoG filter, we develop an interest point detector to extract local structures from images. Compared with the contrast dependent detectors, such as the popular scale invariant feature transform detector, the proposed detector is robust to illumination changes and abrupt variations of images. Experiments on benchmark databases demonstrate the superior performance of the proposed zero-norm LoG detector in terms of the repeatability and matching score of the detected points as well as the image recognition rate under different conditions. Zhenwei Miao, Xudong Jiang 0001, Kim-Hui Yap |
IEEE Trans. Image Process. | 3 |
| 2015 | Hybrid feature-based wallpaper visual searchabstractIn this paper we propose a hybrid feature-based wallpaper visual search system. As opposed to conventional techniques that use global features to perform wallpaper search, this paper proposes to integrate local and global features to support both functions of recognition (identify the product ID of the query images) and retrieval (search wallpapers that are visually similar to the query images). An adaptive SIFT is designed to extract sufficient number of local features from both the query and reference images. The combination of the sparse and dense SIFT features results in a significant improvement of the recognition rate. Global features are further incorporated in the system for the visually similar image retrieval. A new query expansion is proposed to alleviate the problems caused by cluttered background, occlusion, scale change and illumination changes. Experiments on a dataset consisting of 2,208 reference images from 218 different designs show that the proposed method can achieve a recognition rate of more than 90%. Kim-Hui Yap, Zhenwei Miao |
ISCAS | 1 |
| 2015 | Feature weighting in visual product recognitionabstractSignificant progress towards visual search has been made in the past two decades through the development of local invariant features. Among existing local feature detectors, the Scale Invariant Feature Transform (SIFT) is widely used since it is designed to be invariant to minimal illumination changes and certain geometric transformations. However, in practice, the recognition performance is still subject to actual condition. Some keypoints are more stable while others are less stable and can not be repeatedly detected. Besides, in visual object recognition where the foreground object is to be recognized while the background suppressed, the current scalable vocabulary tree (SVT) framework treats each descriptor as equally important, hence restricting its performance. This paper aims to study the effect of SIFT respect to illumination and geometric changes and develop a feature weighting algorithm to incorporate the stability of SIFT and saliency information into weighted scalable vocabulary tree (WSVT) based recognition. Experimental results on a commercial product database show the proposed feature weighting algorithm outperforms the baseline SVT recognition by 5%. Kim-Hui Yap, Dajiang Zhang, Zhenwei Miao |
ISCAS | 2 |
| 2015 | Bijective Weighted Kernel with Connected Component Analysis for Visual Object SearchabstractThis letter proposes a new Bijective Weighted Kernel (BWK) with Connected Component Analysis (CCA) for visual object search. Existing match kernels often employ Term Frequency-Inverse Document Frequency (TF-IDF) weighting which is based on the occurrence frequency of visual words. As opposed to the TF-IDF, the proposed bijective match kernel is designed to exploit the Scalable Vocabulary Tree (SVT) traversal paths to weigh the quantized visual words in image matching. The BWK exploits the corresponding paths between each word in the query and database image to achieve better retrieval performance. The proposed method develops a connected component analysis to detect multiple occurrences of an object with different scales in an image. The method can reduce the computational complexity of geometric verification while achieving accurate object localization. The proposed method is evaluated on the BelgaLogos dataset . Experimental results show that the proposed method outperforms the state-of-the-art methods by a mean Average Precision (mAP) of up to 10%. Subbhuraam Sinduja, Kim-Hui Yap, Dajiang Zhang |
IEEE Signal Process. Lett. | 2 |
| 2014 | Beyond Bag-of-Words: combining generative and discriminative models for scene categorization
Zhen Li 0047, Kim-Hui Yap |
Multim. Tools Appl. | 2 |
| 2014 | Discriminative BoW Framework for Mobile Landmark RecognitionabstractThis paper proposes a new soft bag-of-words (BoW) method for mobile landmark recognition based on discriminative learning of image patches. Conventional BoW methods often consider the patches/regions in the images as equally important for learning. Amongst the few existing works that consider the discriminative information of the patches, they mainly focus on selecting the representative patches for training, and discard the others. This binary hard selection approach results in underutilization of the information available, as some discarded patches may still contain useful discriminative information. Further, not all the selected patches will contribute equally to the learning process. In view of this, this paper presents a new discriminative soft BoW approach for mobile landmark recognition. The main contribution of the method is that the representative and discriminative information of the landmark is learned at three levels: patches, images, and codewords. The patch discriminative information for each landmark is first learned and incorporated through vector quantization to generate soft BoW histograms. Coupled with the learned representative information of the images and codewords, these histograms are used to train an ensemble of classifiers using fuzzy support vector machine. Experimental results on two different datasets show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap |
IEEE Trans. Cybern. | 2 |
| 2014 | Discriminative Soft Bag-of-Visual Phrase for Mobile Landmark RecognitionabstractThis paper proposes a new bag-of-visual phrase (BoP) approach for mobile landmark recognition based on discriminative learning of category-dependent visual phrases. Many previous landmark recognition works adopt a bag-of-words (BoW) method which ignores the co-occurrence relationship between neighboring visual words in an image. Although some works that focus on visual phrase learning have appeared, they mainly construct a generalized phrase dictionary from all categories for recognition, which lacks descriptive capability for a specific category. Another shortcoming of these works is the hard assignment of numerous feature sets to a limited number of phrases, which causes some useful feature sets to be discarded, and yields information loss. In view of this, this paper presents a discriminative soft BoP approach for mobile landmark recognition. The candidate phrases defined as adjacent pairwise codewords are first generated for each category. The important candidates are then selected through a proposed discriminative visual phrase (DVP) selection approach to form the BoP dictionary. Finally, a soft encoding method is developed to quantize each image into a BoP histogram. The context information such as location and direction captured by mobile devices is also integrated with the proposed BoP-based content analysis for landmark recognition. Experimental results on two datasets show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Dajiang Zhang |
IEEE Trans. Multim. | 2 |
| 2013 | An efficient approach for scene categorization based on discriminative codebook learning in bag-of-words framework
Zhen Li 0047, Kim-Hui Yap |
Image Vis. Comput. | 2 |
| 2013 | Context-aware mobile image annotation for media search and sharing
Zhen Li 0047, Kim-Hui Yap, Kiat-Wee Tan |
Signal Process. Image Commun. | 2 |
| 2013 | Context-Aware Discriminative Vocabulary Learning for Mobile Landmark RecognitionabstractThis paper proposes a discriminative vocabulary learning for landmark recognition based on the context information acquired from mobile devices. The vocabulary learning generates a set of discriminative codewords for image representation, which is important for landmark recognition. Many state-of-the-art methods use content analysis alone for vocabulary learning, which underutilizes the context information provided by mobile devices, such as location from the GPS positioner and direction from the digital compass. Although some works start to consider the images' location information for vocabulary learning, the location alone is insufficient since GPS data has significant errors in dense built-up areas. The context analysis techniques that use GPS to shortlist the geographically nearby landmark candidates for subsequent image matching are at times inadequate. In view of this, the paper proposes to employ both direction and location information to learn a discriminative compact vocabulary (DCV) for mobile landmark recognition. Direction information is first considered to supervise image feature clustering to construct direction-dependent scalable vocabulary trees (DSVTs). Location information is then incorporated into the proposed DCV learning algorithm, to select the discriminative codewords of the DSVT to form the DCV. An ImageRank technique and an iterative codeword selection algorithm are developed for DCV learning. Experimental results using the NTU50Landmark database show that the proposed approach achieves 4% improvement over the current method in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Joint Image Registration and Super-Resolution From Low-Resolution Images With Zooming MotionabstractThis paper proposes a new framework for joint image registration and high-resolution (HR) image reconstruction from multiple low-resolution (LR) observations with zooming motion. Conventional super-resolution (SR) methods typically formulate the SR problem as a two-stage process, namely, image registration followed by HR reconstruction. An important step in image SR is the effective estimation of motion parameters. However, the registration algorithms in these two-stage processes experience various degrees of errors. This could degrade the quality of subsequent HR reconstruction. In view of this, this paper presents a new approach that performs joint image registration and SR reconstruction. The proposed iterative SR framework enables the HR image and motion parameters to be estimated simultaneously and progressively. This could increase the potential SR improvement as more accurate estimates of motion parameters could be obtained iteratively. Experimental results show that the proposed method is effective in performing image registration and SR for simulated and real-life images and videos. Yushuang Tian, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Discriminative bag-of-visual phrase learning for landmark recognitionabstractBag-of-visual phrase (BoP) has been proposed and developed for landmark recognition recently. However, existing BoP methods for landmark recognition have two major shortcomings: (i) they try to construct a universal phrase vocabulary for all object categories, which lacks specific descriptive capabilities for a particular category, and (ii) they often adopt simple criterion such as the frequency information to mine the visual phrases, which may cause the selected phrases to be less discriminative or representative for recognition. In view of this, this paper proposes a new discriminative BoP approach for landmark recognition. First, the candidate visual phrases defined as adjacent pairwise words are selected for each category. A phrase-level similarity measure at the latent space is proposed to evaluate the semantic similarity between pairwise phrases. This is then integrated with the phrase frequency information to shortlist the discriminative phrases for each category through a proposed phrase ranking algorithm. Finally, the BoP and bag-of-words (BoW) histograms are combined through a pyramid matching method for recognition. Experimental results on two different datasets demonstrate that the proposed method is effective in landmark recognition. Tao Chen 0003, Kim-Hui Yap, Dajiang Zhang |
ICASSP | 2 |
| 2012 | Multi-frame super-resolution from observations with zooming motionabstractThis paper proposes a new multi-frame super-resolution (SR) approach to reconstruct a high-resolution (HR) image from low-resolution (LR) images with zooming motion. Existing SR methods often assume that the relative motion between acquired LR images consists of only translation and possibly rotation. This restricts the application of these methods in cases when there is zooming motion among the LR images. There are currently only a few methods focusing on zooming SR. Their formulation usually ignores the registration errors or assumes them to be negligible during the HR reconstruction. In view of this, this paper presents a new SR reconstruction approach to handle a more flexible motion model including translation, rotation and zooming. An iterative framework is developed to estimate the motion parameters and HR image progressively. Both simulated and real-life experiments show that the proposed method is effective in performing image SR. Yushuang Tian, Kim-Hui Yap |
ICASSP | 2 |
| 2012 | Efficient mobile landmark recognition based on saliency-aware scalable vocabulary treeabstractIn recent years, the Scalable Vocabulary Tree (SVT) has been shown to be effective in image recognition. However, in mobile landmark image recognition where the foreground is the landmark to be recognized while the background is cluttered, the current SVT framework ignores different local importance of image, hence restricting its performance. In this paper, we propose a new landmark recognition framework that can incorporate saliency information to improve the recognition performance relative to the baseline SVT method. Specifically, the saliency information is incorporated in three phases: image descriptor calculation, vocabulary tree generation, and image representation. We constructed a city-scale landmark dataset in Singapore, and the experimental results show that the proposed mobile landmark recognition by incorporating saliency information outperforms the baseline SVT recognition by about 9%. Kim-Hui Yap, Zhen Li 0047, Dajiang Zhang, Zhan-Ke Ng |
ACM Multimedia | 1 |
| 2012 | Vehicle license plate super-resolution using soft learning prior
Yushuang Tian, Kim-Hui Yap |
Multim. Tools Appl. | 2 |
| 2012 | Content and Context Boosting for Mobile Landmark RecognitionabstractExisting mobile landmark recognition techniques mainly use GPS location information to obtain the candidate images nearby the mobile device, followed by content analysis within the shortlist. This is insufficient since i) GPS often has large errors in dense build-up areas, and ii) direction is underutilized to further improve recognition. In this letter, visual content and two types of mobile context: location and direction, are integrated by the proposed boosting algorithm. Experimental results show that the proposed method outperforms the state-of-the-art methods by about 6%, 11%, and 15% on NTU Landmark-50, PKU Landmark-198, and the large-scale San Francisco landmark dataset, respectively. Zhen Li 0047, Kim-Hui Yap |
IEEE Signal Process. Lett. | 2 |
| 2011 | Beyond bag of words: Combining generative and discriminative models for natural scene categorizationabstractThis paper proposes a simple yet new and effective framework by combining generative model and discriminative model for natural scene categorization. A state-of-the-art approach for scene categorization is the Bag-of-Words (BoW) framework. However, there exist many categories in natural scenes. Often when a new category is considered, the codebook in BoW framework needs to be re-generated, which will involve exhaustive computation. In view of this, this paper tries to ad dress the issue by designing a new framework with the ability of incremental learning. When an additional category is considered, much lower computational cost is needed while the resulting image signatures are still discriminative. The image signatures for training discriminative model are carefully de signed based on the generative model. The effectiveness of the proposed method is validated on UIUC Scene-15 dataset and it is shown to outperform the state-of-the-art method in BoW framework for scene categorization. Zhen Li 0047, Kim-Hui Yap |
ICASSP | 2 |
| 2011 | A discriminative learning technique for mobile landmark recognitionabstractThis paper proposes a discriminative learning bags-of-words (BoW) approach for mobile landmark recognition at patch and image levels. Conventional methods often treat the local patches and images equally important for recognition and do not differentiate their different importance. Although there exist several works that consider the patches' discrimination information, they mainly focus on which patches are to be retained for training and do not incorporate this information when generating the BoW histograms. In view of this, this paper proposes to learn the discriminative information for each landmark category at two levels: local patches and images. At patch level, the patches' discrimination information for each landmark is first discovered using an iterative learning approach. This information is then incorporated into the quantization process to generate the BoW histogram. At image level, the different importance of training images is estimated through a non-parametric density estimator. Finally, fuzzy SVM is used to train the classifier for each category. Experimental results on a landmark database consisting of 3622 training images and 534 testing images show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
ICIP | 2 |
| 2011 | From universal bag-of-words to adaptive bag-of-phrases for mobile scene recognitionabstractThis paper proposes an adaptive bag-of-phrases (BoP) algorithm for mobile scene recognition based on bag-of- words approach. Conventional BoW methods do not consider the dependence and pairwise relationship among different codewords. However, these contextual relations between pairwise codewords play an important role for users to recognize an image. In light of this problem, this paper proposes an effective BoP technique to integrate both the spatial and contextual information between visual words for scene recognition. It first uses hierarchical k-means algorithm to construct a universal codebook for all categories. The contextual (dependence) relationship between pairwise words is then mined for each category based on the mutual information they contain. Subsequently, a visual phrase vocabulary is constructed which is then used to generate a BoP histogram through a proposed quantization method. Finally, support vector machine (SVM) is used to train these histograms into a classifier. Experimental results on the Scene 15 dataset show that the proposed method is effective for mobile scene recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
ICIP | 2 |
| 2011 | A new blind robust image watermarking scheme in SVD-DCT composite domainabstractDigital watermarking has become an important technique for copyright protection, and various watermarking schemes have been proposed. Singular Value Decomposition (SVD) has been used as a valuable transform technique for robust digital watermarking due to some superior characteristics not obtained by DCT, DFT or DWT. In this paper, we present a new robust hybrid image watermarking scheme based on SVD and DCT. After applying SVD to the cover image blocks, we perform DCT on the macro block comprised of the first singular values (SVs) of each image block. We also developed a new method to embed the watermark in the high-frequency band of the SVD-DCT block by imposing a particular relationship between some pseudo-randomly selected pairs of the DCT coefficients. Experimental results show that the proposed watermarking method performs better than state-of-the-art SVD-based methods, and is comparable with the state-of-art wavelet-based robust image watermarking method. Zhen Li 0047, Kim-Hui Yap, Bai Ying Lei |
ICIP | 2 |
| 2011 | Image based approach with k-mean clustering for the compression of human motion sequencesabstractIn this paper, an image format named Virtual Character Animation Image (VCAI) is presented for providing an efficient form of representation for humanoid motion data. By mapping the VCA motion information as 2-D images, characteristics of joint's correlation for the skeletal avatar and temporal coherence within the motion data are jointly reflected as spatial correlation of an image to aid compression. Since the VCA is now encoded as an image, the use of image processing tools and image delivery techniques are now possible. Lastly, a modified motion filter (MMF) is proposed to minimize the visual discontinuity in VCA's motion due to the quantization and transmission noise at high compression rate. The MMF helps to remove high frequency noise components and smoothen the motion signal providing perceptually improved VCA with reduction in distortion. Simulation results demonstrate the effectiveness of the proposed scheme ensuring the minor degradation of VCA quality measured by objective error metric and perceptual loss to the VCA for highly compressed motion stream. Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap |
ISCAS | 3 |
| 2011 | L1-norm multi-frame super-resolution from images with zooming motionabstractThis paper proposes a new image super-resolution (SR) approach to reconstruct a high-resolution (HR) image by fusing multiple low-resolution (LR) images with zooming motion. Most conventional SR image reconstruction methods assume that the motion among different images consists of only translation and possibly rotation. This in-plane motion model, however, is not practical in some applications, when relative zooming exists among the acquired LR images. In view of this, this paper presents a new SR method that addresses a motion model including both in-plane motion (e.g. translation and rotation) and zooming motion. Based on this model, a maximum a posteriori (MAP) based SR algorithm using L1-norm optimization is proposed. Experimental results show that the proposed algorithm based on the new motion model performs well in terms of visual evaluation and quantitative measurement. Yushuang Tian, Kim-Hui Yap, Li Chen 0011 |
MMSP | 2 |
| 2011 | Integrated Content and Context Analysis for Mobile Landmark RecognitionabstractThis paper proposes a new approach for mobile landmark recognition based on integrated content and context analysis. Conventional scene/landmark recognition methods focus mainly on nonmobile desktop/PC platform, where content analysis alone is used to perform landmark recognition. These nonmobile systems, however, do not take unique features of mobile devices into consideration, e.g., limited computational power and fast response time requirement of mobile users. On the contrary, most existing context-aware content mobile landmark recognition methods mainly rely on global positioning system location information for context analysis. In view of this, this paper proposes an effective method that employs an integration of content and context analysis to perform landmark recognition using mobile devices. A new bags-of-words (BoW) framework is developed to perform content analysis. It is then integrated with context analysis involving fusion of location and direction information to perform mobile landmark recognition. Experimental results based on the NTU50Landmark database show that the proposed method can achieve good recognition performance in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | A Fuzzy Clustering Algorithm for Virtual Character Animation RepresentationabstractThe use of realistic humanoid animations generated through motion capture (MoCap) technology is widespread across various 3-D applications and industries. However, the existing compression techniques for such representation often do not consider the implicit coherence within the anatomical structure of a human skeletal model and lacks portability for transmission consideration. In this paper, a novel concept virtual character animation image (VCAI) is proposed. Built upon a fuzzy clustering algorithm, the data similarity within the anatomy structure of a virtual character (VC) model is jointly considered with the temporal coherence within the motion data to achieve efficient data compression. Since the VCA is mapped as an image, the use of image processing tool is possible for efficient compression and delivery of such content across dynamic network. A modified motion filter (MMF) is proposed to minimize the visual discontinuity in VCA's motion due to the quantization and transmission error. The MMF helps to remove high frequency noise components and smoothen the motion signal providing perceptually improved VCA with lessened distortion. Simulation results show that the proposed algorithm is competitive in compression efficiency and decoded VCA quality against the state-of-the-art VCA compression methods, making it suitable for providing quality VCA animation to low-powered mobile devices. Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap |
IEEE Trans. Multim. | 3 |
| 2010 | Bit allocation for scalable video coding of multiple video programsabstractThe problem of joint rate allocation for scalable video coding (SVC) of multiple video programs is addressed in this paper. Most of the existing approaches are based on non-scalable video coding platforms, where computationally expensive encoding or transcoding is demanded to adjust the bit-rate of each video program. Different from all these works, we develop a new statistical multiplexing system, where the scalable video coding technique is applied to compress the video programs. Experiments are carried out to verify the performance of the proposed scheme by comparing it with existing methods. The results demonstrate the merit of the proposed scheme and the variation of quality between different video programs is significantly reduced. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ICIP | 3 |
| 2010 | A Bayesian image annotation framework integrating search and contextabstractConventional approaches to image annotation tackle the problem based on the low-level visual information. Considering the importance of the information on the constrained interaction among the objects in a real world scene, contextual information has been utilized to recognize scene and object categories. In this paper, we propose a Bayesian approach to region-based image annotation, which integrates the content-based search and context into a unified framework. The content-based search selects representative keywords by matching an unlabeled image with the labeled ones followed by a weighted keyword ranking, which are in turn used by the context model to calculate the a prior probabilities of the object categories. Finally, a Bayesian framework integrates the a priori probabilities and the visual properties of image regions. The framework was evaluated using two databases and several performance measures, which demonstrated its superiority to both visual content-based and context-based approaches. Rui Zhang 0010, Kui Wu 0002, Kim-Hui Yap, Ling Guan |
MMSP | 3 |
| 2010 | Adaptive resynchronization approach for scalable video over wireless channel
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Joint Rate Allocation for Multiprogram Video Coding Using FGSabstractIn this paper, we address the problem of joint rate allocation for scalable video coding (SVC) of multiple video programs using the fine granularity scalability (FGS), which is not specified by any current H.264/AVC profile. Most of the existing approaches are based on non-SVC platforms, where computationally expensive encoding or transcoding is demanded to adjust the bit-rate of each video program. Different from all these works, we develop a new statistical multiplexing system, where FGS is applied to compress the video programs. First, we propose an efficient look-ahead approach to distribute the base layer coding bit-rate. Second, a piecewise linear model is applied to accurately estimate the rate-distortion relationship in the FGS layers. Based on this model, a novel algorithm is designed to dynamically allocate the available bandwidth to different programs for rate adaptation in order to minimize the variation of quality of different video programs. Experiments are carried out to verify the performance of the proposed scheme by comparing it with existing methods. The results demonstrate the superiority of the proposed scheme and the quality difference between different programs is greatly reduced. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Progressive Transmission of Motion Capture Data for Scalable Virtual Character AnimationabstractIn this paper, a technique for transmitting level of details motion sequences for virtual character animation is proposed. We have demonstrated that by using progressive level of details (LOD) scheme on the motion capture information, efficient LOD representation of the skeleton motion can be obtained. The proposed virtual character encoder is scalable in nature providing a form of flexibility for the bit-stream to be partially decodable at any bit rate within the bit rate range to address the dynamic bandwidth constraint of a heterogeneous wireless network. Based on the structural characteristic of the virtual character, the packet will be delivered across dynamic channel condition to provide the viewer with the best reconstructed virtual character's motion quality. Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap |
ISCAS | 3 |
| 2009 | A Learning Approach for Single-frame Face Super-resolutionabstractThis paper presents a new learning approach for single-frame face super-resolution (SR). The aim of face SR is to estimate the missing high-resolution (HR) information from a single low-resolution (LR) face image by learning from training samples in the database. A commonly encountered issue in conventional face SR methods is that when the given LR image is a new face significantly different from those in the database, the quality of the reconstructed HR face is usually unsatisfactory. To alleviate this difficulty, we develop a new method to perform face SR based on principal component analysis (PCA) and locally linear embedding (LLE). The reconstructed HR face is able to preserve standard facial features and detailed local information through a residue prediction method using manifold learning. Experimental results show that the proposed method is effective in performing single-frame face SR. Kim-Hui Yap, Lap-Pui Chau |
ISCAS | 2 |
| 2009 | Broadcast of Scalable Video over Wireless NetworksabstractScalable video coding technique has been desired for many years to realize a reliable transmission of video over heterogeneous networks. In this paper, we present a new system of scalable video broadcasting over wireless networks. For each group of pictures, a number of quality layers are produced using scalable video coding. We design the channel protection schemes using the forward error correction codes for the base layer and the enhancement layers, respectively. Given the clients' distribution, we propose a novel algorithm to determine both the source coding bit-rate and the channel coding bit-rate for each layer to maximize a system-defined utility function. Experimental results can demonstrate the superior of the proposed scheme to other schemes and the improvement is up to 2 dB. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ISCAS | 3 |
| 2009 | A Survey on Mobile Landmark Recognition for Information RetrievalabstractThe growing usage of mobile devices has led to proliferation of many mobile applications. A growing trend in mobile applications is centered on mobile landmark recognition. It is a new mobile application that recognizes a captured landmark using the mobile device and retrieves related information. This paper will present a survey on mobile landmark recognition for information retrieval. A general overview of existing mobile landmark recognition systems will be summarized. The techniques and algorithms used in the literatures, including content analysis of landmarks and classification methods for recognition, will be described. Tao Chen 0003, Kui Wu 0002, Kim-Hui Yap, Zhen Li 0047, Flora S. Tsai |
Mobile Data Management | 3 |
| 2009 | A soft MAP framework for blind super-resolution image reconstruction
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
Image Vis. Comput. | 2 |
| 2009 | A Nonlinear L 1 -Norm Approach for Joint Image Registration and Super-ResolutionabstractThis letter proposes a nonlinearL1-norm approach for joint image registration and super-resolution (SR). Image SR is the fusion of multiple low-resolution (LR) images to produce a high-resolution (HR) image. Conventional SR algorithms are sensitive to the initial registration error and outliers in the LR images. In view of this, we present a new SR method to address these problems usingL1-norm optimization in joint image registration and HR image reconstruction. Experimental results show that the proposed method is effective in handling these issues in the HR image reconstruction. Kim-Hui Yap, Yushuang Tian, Lap-Pui Chau |
IEEE Signal Process. Lett. | 1 |
| 2008 | A new color image regularization scheme for blind image deconvolutionabstractThis paper proposes a new regularization scheme to address blind color image deconvolution. Conventional blind monochromatic image deconvolution algorithms handle each color channel independently, thereby ignoring the inter-channel correlation present in the color images. Further, most existing blind color deconvolution algorithms do not take the parametric information of the blurs into consideration. In view of these, a regularization scheme is proposed to perform blind color image deconvolution. A new regularization operator is developed in the blur domain. A reinforcement blur modeling scheme is adopted to evaluate the relevance of manifold parametric blur structures, and the information is integrated into the deconvolution scheme. In addition, a regularization scheme for image is developed to recover edges of color images and reduce color artifacts. Experimental results show that the method is able to achieve satisfactory restored color images under noisy environment. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
ICASSP | 2 |
| 2008 | Spatial resolution decision in scalable bitstream extraction for network and receiver aware adaptationabstractReliable transmission of video over heterogeneous networks requires efficient coding, as well as scalability to different client capabilities, system resources and network conditions. Scalable video coding can provide a full scalability comprising temporal scalability, spatial scalability and quality scalability to increase its adaptability to network and client conditions. It encodes the original video at a full resolution, but enables extracting partial streams to reconstruct the video depending on the specific rate and resolution required by a certain application. This paper addresses the problem of scalable bitstream extraction. Given the bandwidth constraint and the display resolution of the end user, the proposed algorithm will decide the spatial resolution of the sub-stream to be extracted based on the analysis of the content information to maximize the perceptual video quality. Experimental results demonstrate the efficiency of the proposed algorithm. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ICME | 3 |
| 2008 | A Joint Source-Channel Video Coding Scheme Based on Distributed Source CodingabstractRecently, several error resilient schemes have been proposed to tackle the error propagation problem in the motion-compensated predictive video coding based on a promising technique - distributed source coding (DSC). However, these schemes mainly apply the distributed source codes for channel error correction, while under-utilizing their capability for data compression. A channel-aware joint source-channel video coding scheme based on DSC is proposed to eliminate the error propagation problem in predictive video coding in a more efficient way. It is known that near Slepian-Wolf bound DSC is achieved using powerful channel codes, assuming the source and its reference (also known as side-information) are connected by a virtual error-prone channel. In the proposed scheme, the virtual and real error-prone channels are fused so that a unified single channel code is applied to encode the current frame thus accomplishing a joint source-channel coding. Our analysis of the rate efficiency in recovering error propagation shows that the joint scheme can achieve a lower rate compared with performing source and channel coding separately. Simulation results show that the number of bits used for recovering from error propagation can be reduced by up to 10% using the proposed scheme compared to Sehgal-Jagmohan-Ahuja's DSC-based error resilient scheme. Ce Zhu, Kim-Hui Yap |
IEEE Trans. Multim. | 3 |
| 2008 | An Effective Technique for Subpixel Image Registration Under Noisy ConditionsabstractThis paper proposes an effective higher order statistics method to address subpixel image registration. Conventional power spectrum-based techniques employ second-order statistics to estimate subpixel translation between two images. They are, however, susceptible to noise, thereby leading to significant performance deterioration in low signal-to-noise ratio environments or in the presence of cross-correlated channel noise. In view of this, we propose a bispectrum-based approach to alleviate this difficulty. The new method utilizes the characteristics of bispectrum to suppress Gaussian noise. It develops a phase relationship between the image pair and estimates the subpixel translation by solving a set of nonlinear equations. Experimental results show that the proposed technique provides performance improvement over conventional power-spectrum-based methods under different noise levels and conditions. Li Chen 0011, Kim-Hui Yap |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2008 | Subband Synthesis for Color Filter Array DemosaickingabstractThis paper presents a new algorithm for demosaicking images captured through a color filter array (CFA). The objective of the CFA demosaicking is to render a full color image from the mosaicked image. This is commonly achieved by estimating the missing color information from the surrounding observed pixels. In this paper, we integrate the observation that color images have a strong intrachannel spatial correlation in low-frequency components and a dominant interchannel correlation in the high- frequency components. A new framework is proposed to utilize this information where the missing pixels in each color channel are estimated from the wavelet subbands. A modified median-filtering operation is then applied in the subband domain. The algorithm is adaptive and produces superior full-resolution images when compared with other methods. Li Chen 0011, Kim-Hui Yap |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2008 | A Novel Hybrid Model Framework to Blind Color Image DeconvolutionabstractThis paper presents a new hybrid model framework to address blind color image deconvolution. Blind color image deconvolution is a challenging problem due to the limited information on the blurring function. Conventional methods based on the single-input single-output (SISO) model experience suboptimal results as each color channel is processed independently. On the other hand, there are limitations on the practicality of using a multiinput multioutput (MIMO) model in solving this problem as the color channels are usually highly correlated. In view of these constraints, this paper proposes a novel framework to solve blind color image deconvolution by first decomposing the color channels into wavelet subbands, and performing image deconvolution using a hybrid of SISO and single-input multioutput models. The proposed method utilizes the correlation information among different color channels to alleviate the constraints imposed by the MIMO systems. Experimental results show that the method is able to achieve satisfactory restored images under different noise and blurring environments. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2007 | Joint Image Registration and Super-Resolution using Nonlinear Least Squares MethodabstractThis paper proposes a new algorithm to integrate image registration into image super-resolution (SR) by fusing multiple blurred low-resolution (LR) images to render a high-resolution (HR) image. Conventional super-resolution (SR) image reconstruction algorithms assume either the estimated motion (displacement) errors by existing registration methods are negligible or the displacement is known a priori. This assumption, however, is impractical as the performance of existing registration algorithms is still less than perfect. In view of this, we present a new estimation framework that performs joint image registration and HR reconstruction. An iterative scheme based on nonlinear least squares method is developed to estimate the motion shift (displacement) and HR image progressively. The motion model that is considered in this work includes both translation as well as rotation. Experimental results show that the proposed method is effective in performing image super-resolution. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
ICASSP (1) | 2 |
| 2007 | A Motion-Based Selective Error Protection Method for Scalable Video Over Error-Prone ChannelabstractVideo transmission over unreliable networks introduces new challenges in video coding. Due to the predictive coding techniques, the effect of channel errors on the decoded video can be extremely severe when the compressed video is transmitted over error-prone channel. In this paper, the problem of scalable video transmission over error-prone channel is addressed. It is proposed to selectively add forward error correction (FEC) codes to partial information of the compressed bit-stream based on the motion activity of the input video. In addition, unequal error protection (UEP) is applied on the selected data of different temporal layers in a group of pictures (GOP), where the channel rates are optimally allocated. It is shown from the experimental results that our proposed method has a good performance and the improvement is up to 1.2 dB. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ICME | 3 |
| 2007 | Content-based image retrieval using fuzzy perceptual feedback
Kui Wu 0002, Kim-Hui Yap |
Multim. Tools Appl. | 2 |
| 2007 | A Nonlinear Least Square Technique for Simultaneous Image Registration and Super-ResolutionabstractThis paper proposes a new algorithm to integrate image registration into image super-resolution (SR). Image SR is a process to reconstruct a high-resolution (HR) image by fusing multiple low-resolution (LR) images. A critical step in image SR is accurate registration of the LR images or, in other words, effective estimation of motion parameters. Conventional SR algorithms assume either the estimated motion parameters by existing registration methods to be error-free or the motion parameters are known a priori. This assumption, however, is impractical in many applications, as most existing registration algorithms still experience various degrees of errors, and the motion parameters among the LR images are generally unknown a priori. In view of this, this paper presents a new framework that performs simultaneous image registration and HR image reconstruction. As opposed to other current methods that treat image registration and HR reconstruction as disjoint processes, the new framework enables image registration and HR reconstruction to be estimated simultaneously and improved progressively. Further, unlike most algorithms that focus on the translational motion model, the proposed method adopts a more generic motion model that includes both translation as well as rotation. An iterative scheme is developed to solve the arising nonlinear least squares problem. Experimental results show that the proposed method is effective in performing image registration and SR for simulated as well as real-life images. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
IEEE Trans. Image Process. | 2 |
| 2007 | Two-Dimensional Channel Coding Scheme for MCTF-Based Scalable Video CodingabstractThe motion-compensated temporal filtering (MCTF)-based scalable video coding (SVC) provides a full scalability including spatial, temporal and signal-to-noise ratio (SNR) scalability with fine granularity, each of which may result in different visual effect. This paper addresses a novel approach of two-dimensional unequal error protection (2D UEP) for the scalable video with a combined temporal and quality (SNR) scalability over packet-erasure channel. The bit-stream is divided into scalable subbitstreams based on the structure of MCTF. Each subbitstream is further divided into several quality layers. Unequal quantities of bits are allocated to protect different layers to obtain acceptable quality video with smooth degradation under different transmission error conditions. Experimental results are presented to show the advantage of the proposed 2D UEP scheme over the traditional one-dimensional unequal error protection (1D UEP) scheme. Comparing the proposed method with the 1D UEP scheme on SNR layers, our method gives up to 0.81-dB improvement for some video sequences Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
IEEE Trans. Multim. | 4 |
| 2006 | A Bispectrum Technique to Subpixel Image Registration under Noisy ConditionsabstractThis paper proposes an effective higher-order statistics method to address subpixel image registration. Conventional power spectrum-based techniques employ second-order statistics to estimate subpixel translation between two images. They are, however, susceptible to noise, thereby leading to significant performance deterioration in low signal-to-noise (SNR) environments. In view of this, we propose a bispectrum-based approach to alleviate this difficulty. The new method utilizes the characteristics of bispectrum to suppress Gaussian noise. It develops the phase relationship between the image pair, and estimates the subpixel translation by solving a set of nonlinear equations derived from the bispectrum. Experimental results show that the proposed method is effective in identifying subpixel translations under different noise levels and environments. Kim-Hui Yap, Li Chen 0011 |
ICASSP (2) | 1 |
| 2006 | Blind Super-Resolution Image Reconstruction using a Maximum a Posteriori EstimationabstractThis paper proposes a new algorithm to address blind image super-resolution by fusing multiple blurred low-resolution (LR) images to render a high-resolution (HR) image. Conventional super-resolution (SR) image reconstruction algorithms assume either the blurring during the image formation process is negligible or the blurring function is known a priori. This assumption, however, is impractical as it is difficult to eliminate blurring completely in some applications or characterize the blurring function fully. In view of this, we present a new maximum a posteriori (MAP) estimation framework that performs joint blur identification and HR image reconstruction. An iterative scheme based on alternating minimization is developed to estimate the blur and HR image progressively. A blur prior that incorporates the soft parametric blur information and smoothness constraint is introduced in the proposed method. Experimental results show that the new method is effective in performing blind SR image reconstruction where there is limited information about the blurring function. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
ICIP | 2 |
| 2006 | A Novel Resynchronization Method for Scalable Video Over Wireless ChannelabstractA scalable video coder generates scalable compressed bit-stream, which can provide different types of scalability depend on different requirements. This paper proposes a novel resynchronization method for the scalable video with combined temporal and quality (SNR) scalability. The main purpose is to improve the robustness of the transmitted video. In the proposed scheme, the video is encoded into scalable compressed bit-stream with combined temporal and quality scalability. The significance of each enhancement layer unit is estimated properly. A novel resynchronization method is proposed where joint group of picture (GOP) level and picture level insertion of resynchronization marker approach is applied to insert different amount of resynchronization markers in different enhancement layer units for reliable transmission of the video over error-prone channels. It is demonstrated from the experimental results that the proposed method can perform graceful degradation under a variety of error conditions and the improvement can be up to 1 dB compared with the conventional method Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ICME | 3 |
| 2006 | Region-Based Image Retrieval using Radial Basis Function NetworkabstractThis paper presents a new framework that integrates relevance feedback into region-based image retrieval (RBIR) systems based on radial basis function network (RBFN). A modified unsupervised subtractive clustering algorithm is proposed for RBFN center selection according to the characteristics of region-based image representation. A new kernel function of RBFN is introduced for image similarity comparison under region-based representation. The underlying network parameters (weight and width) are then optimized using a supervised gradient-descent training strategy. Experimental results using a database of 10,000 images demonstrate the effectiveness of the proposed hybrid learning approach Kui Wu 0002, Kim-Hui Yap, Lap-Pui Chau |
ICME | 2 |
| 2006 | Two-dimensional channel rate allocation for SVC over error-prone channelabstractThe motion compensated temporal filtering (MCTF) based scalable video coding (SVC) provides a full scalability including spatial, temporal and signal-to-noise ratio (SNR) scalability with fine granularity, each of which may result in different visual effect. This paper addresses a novel approach of two-dimensional unequal error protection (2D UEP) for the scalable video with a combined temporal and quality (SNR) scalability over packet-erasure channel. The bit-stream is divided into scalable sub-bit-streams based on the structure of MCTF. Each sub-bit-stream is further divided into several quality layers. Unequal quantities of bits are allocated to protect different layers to obtain acceptable quality video with smooth degradation under different transmission error conditions. Experimental results are presented to show the advantage of the proposed 2D UEP scheme over the traditional one-dimensional unequal error protection (1D UEP) scheme. Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap |
ISCAS | 4 |
| 2005 | Regularized interpolation using Kronecker product for still imagesabstractIn this paper, we present a new and efficient algorithm for image interpolation. To render high-resolution image from low-resolution image, classical interpolation techniques estimate the missing pixels from the surrounding pixels based on pixel-by-pixel basis. In contrast, this paper proposes an algorithm which is centered on Tikhonov regularization. The regularized solution is derived using the framework of damped least square optimization. Kronecker product and singular value decomposition are employed to reduce the computational cost of the algorithm. Experimental results show that the method produces better interpolation results when compared to other conventional techniques. Li Chen 0011, Kim-Hui Yap |
ICIP (2) | 2 |
| 2005 | Color filter array demosaicking using wavelet-based subband synthesisabstractIn this paper, we present a new and efficient demosaicking algorithm for color filter array (CFA). To render a full-resolution color image using a single-chip camera, the missing color information must be estimated from the surrounding pixels. We take advantage of the observation that the color image has strong inter-channel correlation in the high-frequency subbands. For each color channel, the missing colors are synthesized from the wavelet subbands. The low-frequency subband is estimated by conventional interpolation techniques using intra-channel data. The high-frequency subbands are estimated from the inter-channel data with consideration of the CFA pattern. The algorithm is adaptive in nature and produce superior full-resolution image when compared to other demosaicking techniques. Li Chen 0011, Kim-Hui Yap |
ICIP (2) | 2 |
| 2005 | Blind color image deconvolution based on wavelet decompositionabstractThis paper presents a new framework to address blind color image deconvolution based on wavelet decomposition. Blind color image deconvolution is a challenging problem due to the lack of information available. Conventional methods based on single-input single-output (SISO) model experience significant color artifacts in the restored images. On the other hand, there are limitations on the practicality of using multi-input multi-output (MIMO) model in solving this problem as the color channels are usually highly correlated. In view of this, the paper proposes a new framework to solve blind color image deconvolution by first decomposing the color channels into wavelet subbands, and performing image deconvolution using a combination of SISO and single-input multi-output (SIMO) models. Experimental results show that the proposed method is able to achieve satisfactory restored images. Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau |
ICIP (2) | 2 |
| 2005 | Fuzzy relevance feedback in content-based image retrieval systems using radial basis function networkabstractThis paper presents a new framework called fuzzy relevance feedback in interactive content-based image retrieval (CBIR) systems based on soft-decision. An efficient learning approach is proposed using a fuzzy radial basis function network (FRBFN). Conventional binary labeling schemes require a crisp decision to be made on the relevance of the retrieved images. However, user interpretation varies with respect to different information needs and perceptual subjectivity. In addition, users tend to learn from the retrieval results to further refine their information priority. Therefore, fuzzy relevance feedback is introduced in this paper to integrate the users' fuzzy interpretation of visual content into the notion of relevance feedback. Based on the users' feedbacks, an FRBFN is constructed, and the underlying parameters and network structure are optimized using a gradient-descent training strategy. Experimental results using a database of 10,000 images demonstrate the effectiveness of the proposed method. Kim-Hui Yap, Kui Wu 0002 |
ICME | 1 |
| 2005 | A soft relevance framework in content-based image retrieval systemsabstractThis paper presents a novel framework called fuzzy relevance feedback in interactive content-based image retrieval systems. Conventional binary labeling in relevance feedback requires crisp decisions to be made on the relevance of the retrieved images. This is restrictive as user interpretation of image similarity is imprecise and nonstationary in nature and may vary with respect to different information needs and perceptual subjectivity. It is, therefore, inadequate to model the user perception of image similarity with crisp binary logic. In view of this, we propose a soft relevance notion to integrate the users' fuzzy perception of visual contents into the framework of relevance feedback. A progressive fuzzy radial basis function network is proposed to learn the user information need by optimizing a cost function. An efficient gradient descent-based learning strategy is then employed to estimate the underlying network parameters. Experimental results based on a database of 10 000 images demonstrate the effectiveness of the proposed method. Kim-Hui Yap, Kui Wu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | A soft double regularization approach to parametric blind image deconvolutionabstractThis paper proposes a blind image deconvolution scheme based on soft integration of parametric blur structures. Conventional blind image deconvolution methods encounter a difficult dilemma of either imposing stringent and inflexible preconditions on the problem formulation or experiencing poor restoration results due to lack of information. This paper attempts to address this issue by assessing the relevance of parametric blur information, and incorporating the knowledge into the parametric double regularization (PDR) scheme. The PDR method assumes that the actual blur satisfies up to a certain degree of parametric structure, as there are many well-known parametric blurs in practical applications. Further, it can be tailored flexibly to include other blur types if some prior parametric knowledge of the blur is available. A manifold soft parametric modeling technique is proposed to generate the blur manifolds, and estimate the fuzzy blur structure. The PDR scheme involves the development of the meaningful cost function, the estimation of blur support and structure, and the optimization of the cost function. Experimental results show that it is effective in restoring degraded images under different environments. Li Chen 0011, Kim-Hui Yap |
IEEE Trans. Image Process. | 2 |
| 2004 | An efficient radial basis function network approach for content-based image retrievalabstractIn this paper, an efficient approach using radial basis function network (RBFN) with online learning capability is proposed for interactive content-based image retrieval (CBIR) systems. Based on the users' feedback, an RBFN is constructed, and the underlying parameters and network structure are adjusted adaptively using a training strategy. To capture the users' perceptual consistency in similarity, an error function is expressed in terms of accumulated training samples across all feedback sessions. Experimental results using a database of 10000 images demonstrate the effectiveness of the proposed method. Kui Wu 0002, Kim-Hui Yap |
ICASSP (3) | 2 |
| 2004 | A noisy chaotic neural network approach to image denoisingabstractThis paper presents a new approach to address image denoising based on a new neural network, called noisy chaotic neural network (NCNN). The original Bayesian framework of image denoising is reformulated into a constrained optimization problem using continuous relaxation labeling. The NCNN, which combines the simulated annealing technique with the Hopfield neural network (HNN), is employed to solve the optimization problem. It effectively overcomes the local minima problem which may be incurred by the HNN. The experimental results show that the NCNN could offer good quality solutions. Leipo Yan, Lipo Wang 0001, Kim-Hui Yap |
ICIP | 3 |
| 2003 | A fuzzy K-nearest-neighbor algorithm to blind image deconvolutionabstractThis paper proposes an adaptive blind image deconvolution scheme based on fuzzy K-nearest-neighbor (FKNN) algorithm. It is well known that most point-spread functions (PSFs) satisfy up to a certain degree of parametric structure. The method incorporates such knowledge about the PSF structure by estimating the PSF according to its K nearest neighbors. Through a process of neighbor generation, model matching, and fuzzy weighted mean filtering, FKNN provides a robust estimate for the blur. This further improves the convergence performance in blind deconvolution process. Experimental results show that it is effective in restoring degraded images where there is little prior knowledge about the blur. Li Chen 0011, Kim-Hui Yap |
SMC | 2 |
| 2002 | A fuzzy blur algorithm to adaptive blind image deconvolutionabstractThis paper proposes a new approach to blind image deconvolution based on fuzzy blur interference algorithm. Conventional blind algorithms require a crisp decision to be made on the structure of the blurring function prior to formulation. This creates a dilemma as the complete blur information is usually unknown a priori. Most blind algorithms either employ absolute but inflexible parametric modeling or ignore the parametric blur knowledge completely. This paper presents a fuzzy approach to resolve this difficulty by constructing a soft model set consisting of parametric estimates of the current blur. The relevance of these estimates are evaluated, and integrated to form a fuzzy blur using compositional inference rule. The main feature of the technique lies in its ability to incorporate domain knowledge while preserving the flexibility of the scheme. Experimental results show that the technique is effective in restoring blurred, noisy images without prior knowledge of the blur. Kim-Hui Yap, Ling Guan |
ICARCV | 1 |
| 2002 | A computational reinforced learning scheme to blind image deconvolutionabstractThis paper presents a new approach to adaptive blind image deconvolution based on computational reinforced learning in an attractor-embedded solution space. The new technique develops an evolutionary strategy that generates the improved blur and image populations progressively. A dynamic attractor space is constructed by integrating the knowledge domain of the blur structures into the algorithm. The attractors are predicted using a maximum a posteriori estimator and their relevance is evaluated with respect to the computed blurs. We develop a novel reinforced mutation scheme that combines stochastic search and pattern acquisition throughout the blur identification. It enhances the algorithmic convergence and reduces the computational cost significantly. The new technique is robust in alleviating the constraints and difficulties encountered by most conventional methods. Experimental results show that the new algorithm is effective in restoring the degraded images and identifying the blurs. Kim-Hui Yap, Ling Guan |
IEEE Trans. Evol. Comput. | 1 |
| 2001 | Blind adaptive detection for CDMA systems based on regularized independent component analysisabstractWe present a new approach to blind adaptive detection for CDMA systems based on regularized independent component analysis (ICA). Classical ICA algorithms are effective in separating linearly weighted signal mixtures consisting of subGaussian and superGaussian signals. However, they do not incorporate any information of the weighting matrix, in this case, the user's signature sequence into the formulation. This results in underutilization of the information available. To address this difficulty, we propose a new ICA algorithm that combines a contrast function and a regularization functional to integrate the information of the user's signature. A blind adaptive detector based on stochastic gradient optimization of the new cost function is derived. Simulation results show that the new technique provides good interference suppression, fast convergence and low BER performance when compared with other blind detectors. Kim-Hui Yap, Ling Guan, Jamie S. Evans |
GLOBECOM | 1 |
| 2001 | An attractor space approach to blind image deconvolutionabstractWe present a new approach to adaptive blind image deconvolution based on computational reinforced learning in attractor-embedded solution space. A new subspace optimization technique is developed to restore the image and identify the blur. Conjugate gradient optimization is employed to provide an adaptive image restoration while a new evolutionary scheme is devised to generate the high-performance blur estimates. The new technique is flexible as it does not suffer from various image or blur constraints imposed by most traditional blind methods. Experimental results show that the new algorithm is effective in blind deconvolution of images degraded under different blur structures and noise levels. Kim-Hui Yap, Ling Guan |
ICASSP | 1 |
| 2000 | A Recursive Soft-Decision PSF and Neural Network Approach to Adaptive Blind Image RegularizationabstractWe present a new approach to adaptive blind image regularization based on a neural network and soft-decision blur identification. We formulate blind image deconvolution into a recursive scheme by projecting and optimizing a novel cost function with respect to its image and blur subspaces. The new algorithm provides a continual blur adaptation towards the best-fit parametric structure throughout the restoration. It integrates the knowledge of real-life blur structures without compromising its flexibility in restoring images degraded by other nonstandard blurs. A nested neural network, called the hierarchical cluster model is employed to provide an adaptive, perception-based restoration. On the other hand, conjugate gradient optimization is adopted to identify the blur. Experimental results show that the new approach is effective in restoring the degraded image without the prior knowledge of the blur. Kim-Hui Yap, Ling Guan |
ICIP | 1 |
| 1999 | Image Restoration Based on Hierarchical Cluster Model with Evolutionary OptimizationabstractIn this paper, a new approach to adaptive image regularization based on Hierarchical Cluster Model (HCM) with evolutionary optimization is proposed. HCM is a hierarchical neural network with distributed clusters. Its sparse synaptic connections and parallel structure reduce the computational cost of restoration. Adaptive restoration is achieved by assigning entries of an optimized regularization vector to each homogeneous cluster of the image. The clusters are restored in the order of smooth to texture and edge regions to minimize the regularization error. An evolutionary scheme is employed to improve the performance profile of the restored image and optimize the regularization vector. Experimental results show that the new approach is superior in suppressing noise and ringing at the smooth background while preserving fine details at the texture and edge regions effectively. Kim-Hui Yap, Ling Guan |
ICIP (1) | 1 |