VLDB 2026 Research / reviewers in the wild / expert
Joo-Hwee Lim
dblp:56/432 · also Joo Hwee Lim
· DBLP profile ↗
168ranked-venue papers
22as first author
39since 2021 · last 2026
0000-0002-4103-3824ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 114 · 15 first-author · 28 since 2021Artificial intelligence and machine learning · 75 · 10 first-author · 20 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorHuman-computer interaction and ubiquitous computing · 5Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If CalibratedabstractDespite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world. Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002 |
AAAI | 9 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | SPASCA: Social Presence and Support with Conversational Agent for Persons Living with DementiaabstractWe present SPASCA - a conversational AI system that promotes psychological and cognitive well-being of persons living with dementia (PLWD). This system features an AI agent that provides social presence and support to PLWD through verbal communications, without physical presence of human caregivers. The system integrates (1) a novel dialogue model that generates dialogue items relevant to the user's experiences and lifestyle, (2) a digital avatar in the form of a talking head with the identity of a caregiver who is familiar to the demented user. We develop prototypes that adopt various interaction modalities and conversational styles and report the pros and cons of different system configurations through expert review. Our system shows the potential of conversational AI for personalized and affordable healthcare services. Ali Koksal, Jingjing Gu, Kotaro Hara, Joo-Hwee Lim, Qianli Xu |
AAAI | 5 |
| 2025 | Visual Prompting for One-shot Controllable Video Editing without InversionabstractOne-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made – using any image editing tool – on the first frame of a video to all subsequent frames, while ensuring content consistency between edited frames and source frames. To achieve this, prior methods employ DDIM inversion to transform source frames into latent noise, which is then fed into a pre-trained diffusion model, conditioned on the user-edited first frame, to generate the edited video. However, the DDIM inversion process accumulates errors, which hinder the latent noise from accurately reconstructing the source frames, ultimately compromising content consistency in the generated edited frames. To overcome it, our method eliminates the need for DDIM inversion by performing OCVE through a novel perspective based on visual prompting. Furthermore, inspired by consistency models that can perform multi-step consistency sampling to generate a sequence of content-consistent images, we propose a content consistency sampling (CCS) to ensure content consistency between the generated edited frames and the source frames. Moreover, we introduce a temporal-content consistency sampling (TCS) based on Stein Variational Gradient Descent to ensure temporal consistency across the edited frames. Extensive experiments validate the effectiveness of our approach. Zhengbo Zhang, Duo Peng, Joo-Hwee Lim, Zhigang Tu 0001, De Wen Soh, Lin Geng Foo |
CVPR | 4 |
| 2025 | Decreasing Word Error Rates in Paragraph Handwritten Text Recognition with Synthetic DataabstractHandwritten Text Recognition (HTR) faces a persistent challenge with the scarcity of data at the paragraph level, arising from the difficulty of acquiring diverse, cost-efficient, and cleanly labeled datasets for training. As such, works in HTR leverage segmentation, regularization techniques, and language modeling to excel in a low-data environment. While synthetic generation methods gain traction on the word and line-level recognition, this success has not translated to the paragraph level. Hence, our work seeks to mimic the nuances of paragraph-level text images with a custom synthetic data engine using Wikipedia texts. Experiments show that by using our synthetic dataset in tandem with a simple encoder-decoder Transformer, we can achieve the best Word Error Rate (WER) amongst the state-of-the-art methods for handwriting recognition on the IAM dataset. Additionally, we show the model pretrained on English texts can also recognize French and German texts with minimal finetuning. Ernest Yu Kai Chew, Adams Wai-Kin Kong, Joo-Hwee Lim |
ICASSP | 3 |
| 2025 | A Unified Model for Paragraph and Line-Level Handwritten Text Recognition
Ernest Yu Kai Chew, Adams Wai-Kin Kong, Joo-Hwee Lim |
ICDAR (2) | 3 |
| 2025 | NeuroViG - Integrating Event Cameras for Resource-Efficient Video GroundingabstractSpatio-Temporal Video Grounding (STVG) - the task of identifying the target object in the field-of-view, that the language instruction refers to - is a fundamental vision-language task. Current STVG approaches typically utilize feeds from an RGB camera that is assumed to be always-on and process the video frames using complex neural network pipelines. As a result, they often impose prohibitive system overheads (energy, latency) on pervasive devices. To address this, we propose NeuroViG with two key innovations: (a) leveraging on event streams from a low-power neuromorphic event camera sensor to perform selective triggering of the more energy-hungry RGB camera for STVG, and (b) augmenting the STVG model with a lightweight Adaptive Frame Selector (AFS) that bypasses complex transformer-based operations for a majority of video frames, thereby enabling its execution on a pervasive Jetson AGX device. We have also introduced modifications to the neural network processing pipeline such that the system can offer tunable tradeoffs between accuracy and energy/latency. Our proposed NeuroViG system allows us to reduce the STVG energy overhead and latency by ~ 4x and ~ 3.8x, respectively, for less than 1% loss in accuracy. Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo-Hwee Lim, Archan Misra |
WACV | 3 |
| 2025 | Unveiling the Tapestry: The Interplay of Generalization and Forgetting in Continual LearningabstractIn artificial intelligence (AI), generalization refers to a model's ability to perform well on out-of-distribution data related to the given task, beyond the data it was trained on. For an AI agent to excel, it must also possess the continual learning capability, whereby an agent incrementally learns to perform a sequence of tasks without forgetting the previously acquired knowledge to solve the old tasks. Intuitively, generalization within a task allows the model to learn underlying features that can readily be applied to novel tasks, facilitating quicker learning and enhanced performance in subsequent tasks within a continual learning framework. Conversely, continual learning methods often include mechanisms to mitigate catastrophic forgetting, ensuring that knowledge from earlier tasks is retained. This preservation of knowledge over tasks plays a role in enhancing generalization for the ongoing task at hand. Despite the intuitive appeal of the interplay of both abilities, existing literature on continual learning and generalization has proceeded separately. In the preliminary effort to promote studies that bridge both fields, we first present empirical evidence showing that each of these fields has a mutually positive effect on the other. Next, building upon this finding, we introduce a simple and effective technique known as shape-texture consistency regularization (STCR), which caters to continual learning. STCR learns both shape and texture representations for each task, consequently enhancing generalization and thereby mitigating forgetting. Remarkably, extensive experiments validate that our STCR, can be seamlessly integrated with existing continual learning methods, including replay-free approaches. Its performance surpasses these continual learning methods in isolation or when combined with established generalization techniques by a large margin. Our data and source code are available at https://github.com/ZhangLab-DeepNeuroCogLab/distillation-style-cnn. Zenglin Shi, Jie Jing 0001, Ying Sun 0001, Joo-Hwee Lim, Mengmi Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Diffusion Time-step Curriculum for One Image to 3D GenerationabstractScore distillation sampling (SDS) has been widely adopted to overcome the absence of unseen views in reconstructing 3D objects from a single image. It leverages pretrained 2D diffusion models as teacher to guide the reconstruction of student 3D models. Despite their remarkable success, SDS-based methods often encounter geometric artifacts and texture saturation. We find out the crux is the overlooked indiscriminate treatment of diffusion time-steps during optimization: it unreasonably treats the student-teacher knowledge distillation to be equal at all time-steps and thus entangles coarse-grained and fine-grained modeling. Therefore, we propose the Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both the teacher and student models collaborating with the time-step curriculum in a coarse-to-fine manner. Extensive experiments on NeRF4, RealFusion15, GSO and Level50 benchmark demonstrate that DTC123 can produce multiview consistent, high-quality, and diverse 3D assets. Codes and more generation demos will be released in https://github.com/yxymessi/DTC123. Xuanyu Yi, Zike Wu, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Hanwang Zhang |
CVPR | 5 |
| 2024 | Poster: Towards Efficient Spatio-Temporal Video Grounding in Pervasive Mobile DevicesabstractAs the use of pervasive devices expands into complex collaborative tasks such as cognitive assistants and interactive AR/VR companions, they are equipped with a myriad of sensors facilitating natural interactions, such as voice commands. Spatio-Temporal Video Grounding (STVG), the task of identifying the target object in the field-of-view referred to in a language instruction, is a key capability needed for such systems. However, current STVG models tend to be resource-intensive, relying on multiple cross-attentional transformers applied to each video frame. This results in runtime complexity that increases linearly with video length. Furthermore, deploying these models on mobile devices while maintaining a low-latency poses additional challenges. Hence, this paper explores the latency and energy requirements for implementing STVG models on a pervasive device. Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo-Hwee Lim, Archan Misra |
MobiSys | 3 |
| 2024 | MVGamba: Unify 3D Content Generation as State Space Sequence ModelingabstractRecent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (\eg, Transformers).
To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity.
With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0.1\times$ of the model size. The codes are available at \url{https://github.com/SkyworkAI/MVGamba}. Xuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Shuicheng Yan, Xinchao Wang, Hanwang Zhang |
NeurIPS | 6 |
| 2024 | Enhancing Representation Learning With Spatial Transformation and Early Convolution for Reinforcement Learning-Based Small Object DetectionabstractAlthough object detection has achieved significant progress in the past decade, detecting small objects is still far from satisfactory due to the high variability of object scales and complex backgrounds. The common way to enhance small object detection is to use high-resolution (HR) images. However, this method incurs huge computational resources which grow squarely with the resolution of images. To achieve both accuracy and efficiency, we propose a novel reinforcement learning framework that employs an efficient policy network consisting of a Spatial Transformation Network to enhance the state representation learning and a Transformer model with early convolution to improve feature extraction. Our method has two main steps: (1) coarse location query (CLQ), where an RL agent is trained to predict the locations of small objects on low-resolution (LR) (down-sampled version of HR) images; (2) context-sensitive object detection where HR image patches are used to detect objects on the selected coarse locations and LR image patches on background areas (containing no small objects). In this way, we can obtain high detection performance on small objects while avoiding unnecessary computation on background areas. The proposed method has been tested and benchmarked on various datasets. On the Caltech Pedestrians Detection and Web Pedestrians datasets, the proposed method improves the detection accuracy by 2%, while reducing the number of processed pixels. On the Vision meets Drone object detection dataset and the Oil and Gas Storage Tank dataset, the proposed method outperforms the state-of-the-art (SotA) methods. On MS COCO mini-val set, our method outperforms SotA methods on small object detection, while also achieving comparable performance on medium and large objects. Fen Fang, Wenyu Liang, Qianli Xu, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question AnsweringabstractThe main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in questions and employ a simple graph to aggregate features without considering relative relations between objects, which may lead to inferior performance. In this paper, we propose a Keyword-aware Relative Spatio-Temporal (KRST) graph network for VideoQA. First, to make question features aware of keywords, we employ an attention mechanism to assign high weights to keywords during question encoding. The keyword-aware question features are then used to guide video graph construction. Second, because relations are relative, we integrate the relative relation modeling to better capture the spatio-temporal dynamics among object nodes. Moreover, we disentangle the spatio-temporal reasoning into an object-level spatial graph and a frame-level temporal graph, which reduces the impact of spatial and temporal relation reasoning on each other. Extensive experiments on the TGIF-QA, MSVD-QA and MSRVTT-QA datasets demonstrate the superiority of our KRST over multiple state-of-the-art methods. Hehe Fan, Dongyun Lin, Ying Sun 0001, Mohan Kankanhalli, Joo-Hwee Lim |
IEEE Trans. Multim. | 6 |
| 2024 | Controllable Video Generation With Text-Based InstructionsabstractMost of the existing studies on controllable video generation either transfer disentangled motion to an appearance without detailed control over motion or generate videos of simple actions such as the movement of arbitrary objects conditioned on a control signal from users. In this study, we introduce Controllable Video Generation with text-based Instructions (CVGI) framework that allows text-based control over action performed on a video. CVGI generates videos where hands interact with objects to perform the desired action by generating hand motions with detailed control through text-based instruction from users. By incorporating the motion estimation layer, we divide the task into two sub-tasks: (1) control signal estimation and (2) action generation. In control signal estimation, an encoder models actions as a set of simple motions by estimating low-level control signals for text-based instructions with given initial frames. In action generation, generative adversarial networks (GANs) generate realistic hand-based action videos as a combination of hand motions conditioned on the estimated low control level signal. Evaluations on several datasets (EPIC-Kitchens-55, BAIR robot pushing, and Atari Breakout) show the effectiveness of CVGI in generating realistic videos and in the control over actions. Ali Koksal, Kenan E. Ak, Ying Sun 0001, Deepu Rajan, Joo-Hwee Lim |
IEEE Trans. Multim. | 5 |
| 2023 | Counterfactual Dynamics Forecasting - a New Setting of Quantitative ReasoningabstractRethinking and introspection are important elements of human intelligence. To mimic these capabilities, counterfactual reasoning has attracted attention of AI researchers recently, which aims to forecast the alternative outcomes for hypothetical scenarios (“what-if”). However, most existing approaches focused on qualitative reasoning (e.g., casual-effect relationship). It lacks a well-defined description of the differences between counterfactuals and facts, as well as how these differences evolve over time. This paper defines a new problem formulation - counterfactual dynamics forecasting - which is described in middle-level abstraction under the structural causal models (SCM) framework and derived as ordinary differential equations (ODEs) as low-level quantitative computation. Based on it, we propose a method to infer counterfactual dynamics considering the factual dynamics as demonstration. Moreover, the evolution of differences between facts and counterfactuals are modelled by an explicit temporal component. The experimental results on two dynamical systems demonstrate the effectiveness of the proposed method. Yanzhu Liu, Ying Sun 0001, Joo-Hwee Lim |
AAAI | 3 |
| 2023 | Towards Debiasing Frame Length Bias in Text-Video Retrieval via Causal Intervention
Burak Satar, Hongyuan Zhu 0002, Hanwang Zhang, Joo-Hwee Lim |
BMVC | 4 |
| 2023 | Invariant Training 2D-3D Joint Hard Samples for Few-Shot Point Cloud RecognitionabstractWe tackle the data scarcity challenge in few-shot point cloud recognition of 3D objects by using a joint prediction from a conventional 3D model and a well-trained 2D model. Surprisingly, such an ensemble, though seems trivial, has hardly been shown effective in recent 2D-3D models. We find out the crux is the less effective training for the "joint hard samples", which have high confidence prediction on different wrong labels, implying that the 2D and 3D models do not collaborate well. To this end, our proposed invariant training strategy, called INVJOINT, does not only emphasize the training more on the hard samples, but also seeks the invariance between the conflicting 2D and 3D ambiguous predictions. INVJOINT can learn more collaborative 2D and 3D representations for better ensemble. Extensive experiments on 3D shape classification with widely adopted ModelNet10/40, ScanObjectNN and Toys4K, and shape retrieval with ShapeNet-Core validate the superiority of our INVJOINT. Codes will be publicly Available1. Xuanyu Yi, Jiajun Deng, Qianru Sun, Xian-Sheng Hua 0001, Joo-Hwee Lim, Hanwang Zhang |
ICCV | 5 |
| 2023 | Data Augmentation Using Corner CutMix and an Auxiliary Self-Supervised LossabstractDeep convolutional neural networks (CNNs) have achieved remarkable success in computer vision tasks, but their training is susceptible to overfitting when the training sample size is insufficient. In this paper, we introduce Corner CutMix, a novel data augmentation technique for CNN training. During training, Corner CutMix randomly selects a region from one of four corner areas in an image and replaces it with a randomly chosen region from a distractor image. Additionally, we design an auxiliary self-supervised loss function to learn the position of the selected corner region, thereby improving the transferability and generalizability of the learned representation. Corner CutMix is easy to implement, adding little computational overhead, and can be combined with other augmentation methods such as random cropping, color distortion, and flipping. Our extensive classification task experiments in self-supervised learning on public datasets (e.g., CIFAR10, CIFAR100, and STL10) demonstrate the effectiveness of Corner CutMix, which consistently outperforms strong baselines such as CutOut and CutMix. Fen Fang, Nhat M. Hoang, Qianli Xu, Joo-Hwee Lim |
ICIP | 4 |
| 2023 | Learning by Imagination: A Joint Framework for Text-Based Image Manipulation and Change CaptioningabstractImage and text are dual modalities of our semantic interpretation. Changing images based on text descriptions allows us to imagine and visualize the world (a.k.a. text-based image manipulation (TIM)). In this paper, we introduce a framework that combines TIM with change captioning (CC) and utilizes the benefits of co-training. CC aims to describe what has changed in a scene and can be regarded as the inverse version of TIM where both tasks rely on generative networks. These generative networks can be regarded as data producers of each other and unlike previous methods, we discover that integrating their learning procedures can benefit both. Since the CC module describes differences between two images as text, the CC module can be used as evaluation criteria and provide feedback. Furthermore, we utilize a shared attention mechanism in TIM and CC modules to localize towards prominent regions as well as enabling a change-aware discriminator. In the opposite direction, the output image synthesized by the TIM module can be assessed with the CC module, by checking whether the ground truth text description can be redescribed. Following this insight, not only do we boost the training of the TIM module, but we also utilize the TIM module as additional supervision for the CC training. Experimental results show that our framework outperforms existing TIM methods on several datasets substantially and we achieve marginal improvements in the CC module. To our best knowledge, this is the first study dedicated to the joint training of TIM and CC tasks. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Multim. | 3 |
| 2022 | Identifying Hard Noise in Long-Tailed Sample Distribution
Xuanyu Yi, Kaihua Tang, Xian-Sheng Hua 0001, Joo-Hwee Lim, Hanwang Zhang |
ECCV (26) | 4 |
| 2022 | Improving Generalization of Reinforcement Learning Using a Bilinear Policy NetworkabstractIn deep reinforcement learning (DRL), the agent is usually trained on seen environments by optimizing a policy network. However, it is difficult to be generalized to unseen environments properly, even when the environmental variations are insignificant. This is partly because the policy network cannot effectively learn the representation of visual difference that is subtle among highly similar states in the environments. Because a bilinear structured model containing two feature extractors allows pairwise feature interactions in a translation-ally invariant manner which makes it particularly useful for subtle difference recognition among highly similar states, in this work, a bilinear policy network is employed to enhance representation learning, and thus to improve generalization of the DRL. The proposed bilinear policy network is tested on various DRL task, including a control task on path planning for active object detection, and Grid World, an AI game task. The test results show that the generalization of DRL can be improved by the proposed network. Fen Fang, Wenyu Liang, Yan Wu 0002, Qianli Xu, Joo-Hwee Lim |
ICIP | 5 |
| 2022 | Hierarchical Defect Detection Based On Reinforcement LearningabstractIn this paper, we propose a novel reinforcement learning (RL) based method for defects detection in high-resolution (HR) images. e.g. cracks and scratches on the surfaces of buildings, constructions, and products. Our innovation leverages RL to explore challenging images in progressive manner, using pre-trained deep learning (DL) detection as feedback mechanism. First, The DL model is pre-trained on low resolution (LR) images with relatively high defect background ratio (DBR). The RL agent is trained by optimizing a policy network according to feedback of DL model on selected regions of HR images with fairly low DBR to coarsely predict defective region by executing two actions: defective region selection and region refinement. Then, the selected defective regions are evaluated using the DL model to generate final defect region which will be mapped back to the HR images. Experimental results on HR crack and scratch images indicate that our method is able to achieve state-of-the-art performance with 0.976 and 0.965 F1-score respectively. Fen Fang, Qianli Xu, Joo-Hwee Lim |
ICIP | 3 |
| 2022 | Portmanteauing Features for Scene Text RecognitionabstractScene text images have different shapes and are subjected to various distortions, e.g. perspective distortions. To handle these challenges, the state-of-the-art methods rely on a rectification network, which is connected to the text recognition network. They form a linear pipeline which uses text rectification on all input images, even for images that can be recognized without it. Undoubtedly, the rectification network improves the overall text recognition performance. However, in some cases, the rectification network generates unnecessary distortions on images, resulting in incorrect predictions in images that would have otherwise been correct without it. In order to alleviate the unnecessary distortions, the portmanteauing of features is proposed. The portmanteau feature, inspired by the portmanteau word, is a feature containing information from both the original text image and the rectified image. To generate the portmanteau feature, a non-linear input pipeline with a block matrix initialization is presented. In this work, the transformer is chosen as the recognition network due to its utilization of attention and inherent parallelism, which can effectively handle the portmanteau feature. The proposed method is examined on 6 benchmarks and compared with 13 state-of-the-art methods. The experimental results show that the proposed method outperforms the state-of-the-art methods on various of the benchmarks. Yew Lee Tan, Ernest Yu Kai Chew, Adams Wai-Kin Kong, Jung-Jae Kim 0001, Joo-Hwee Lim |
ICPR | 5 |
| 2022 | Entropy guided attention network for weakly-supervised action localization
Ying Sun 0001, Hehe Fan, Tao Zhuo, Joo-Hwee Lim, Mohan Kankanhalli |
Pattern Recognit. | 5 |
| 2022 | EEG-Video Emotion-Based Summarization: Learning With EEG Auxiliary SignalsabstractVideo summarization is the process of selecting a subset of informative keyframes to expedite storytelling with limited loss of information. In this article, we propose an EEG-Video Emotion-based Summarization (EVES) model based on a multimodal deep reinforcement learning (DRL) architecture that leverages neural signals to learn visual interestingness to produce quantitatively and qualitatively better video summaries. As such, EVES does not learn from the expensive human annotations but the multimodal signals. Furthermore, to ensure the temporal alignment and minimize the modality gap between the visual and EEG modalities, we introduce a Time Synchronization Module (TSM) that uses an attention mechanism to transform the EEG representations onto the visual representation space. We evaluate the performance of EVES on the TVSum and SumMe datasets. Based on the rank order statistics benchmarks, the experimental results show that EVES outperforms the unsupervised models and narrows the performance gap with supervised models. Furthermore, the human evaluation scores show that EVES receives a higher rating than the state-of-the-art DRL model DR-DSN by 11.4% on the coherency of the content and 7.4% on the emotion-evoking content. Thus, our work demonstrates the potential of EVES in selecting interesting content that is both coherent and emotion-evoking. Wai-Cheong Lincoln Lew, Di Wang 0004, Kai Keng Ang, Joo-Hwee Lim, Hiok Chai Quek, Ah-Hwee Tan |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Image Understanding With Reinforcement Learning: Auto-Tuning Image Attributes and Model Parameters for Object Detection and SegmentationabstractModels for image semantics understanding, such as deep learning (DL) models and mathematical models, are often trained on specific dataset or configured with specific parameters. When deploying such models on new tasks in a different test environment, it requires considerable effort to re-train the model or extensive expertise to tune the parameters. In this paper, we propose a smart reinforcement learning (RL) agent that could learn to tune parameters automatically to enhance model performance. The learning process is formulated as a generic control task for parameter adjustment, and applied to two use scenarios: (1) image attributes tuning to improve object detection performance on fixed DL model, and (2) parameter tuning of the mathematical model (Level Set) for image segmentation. We design a novel dynamic threshold mechanism in a multi-branch RL agent to effectively tune parameters of image qualities (for object detection) and Level Set models (for object segmentation). We conduct experiments on Pascal-VOC testing set, MS COCO validation set and a proprietary dataset of industrial components, where we achieve substantial improvement on object detection accuracy. We also perform experiments on the automatic parameter tuning of Level Set models. Results show that our method facilitates considerable performance improvement on public datasets compared with baseline method. Fen Fang, Qianli Xu, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Align R-CNN: A Pairwise Head Network for Visual Relationship DetectionabstractScene graphs connect individual objects with visual relationships. They serve as a comprehensive scene representation for downstream multimodal tasks. However, by exploring recent progress in Scene Graph Generation (SGG), we find that the performance of recent works is highly limited by the pairwise relationship modeling by naive feature concatenation. Such pairwise features lack sufficient object interaction due to the mis-aligned object parts, resulting in non-discriminative pairwise features for visual relationship prediction. For example, naive concatenated pairwise feature usually make the model fail to discriminate betweenridingandfeedingfor object pairpersonandhorse. To this end, we design a meta-architecture— learning-to-align — for dynamic object feature concatenation. We call our model:Align R-CNN. Specifically, we introduce a novel attention-based multiple region alignment module that can be jointly optimized with SGG. Experiments on the large-scale SGG benchmark Visual Genome show that the proposedAlign R-CNNcan replace the naive feature concatenation and thus boost all the existing SGG methods. Mitra Tajrobehkar, Kaihua Tang, Hanwang Zhang, Joo-Hwee Lim |
IEEE Trans. Multim. | 4 |
| 2021 | TAILOR: Teaching with Active and Incremental Learning for Object RegistrationabstractWhen deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions. Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
AAAI | 9 |
| 2021 | Robust Multi-Frame Future Prediction By Leveraging View SynthesisabstractIn this paper, we focus on the problem of video prediction, i.e., future frame prediction. Most state-of-the-art techniques focus on synthesizing a single future frame at each step. However, this leads to utilizing the model’s own predicted frames when synthesizing multi-step prediction, resulting in gradual performance degradation due to accumulating errors in pixels. To alleviate this issue, we propose a model that can handle multi-step prediction. Additionally, we employ techniques to leverage from view synthesis for future frame prediction, where both problems are treated independently in the literature. Our proposed method employs multiview camera pose prediction and depth-prediction networks to project the last available frame to desired future frames via differentiable point cloud renderer. For the synthesis of moving objects, we utilize an additional refinement stage. In experiments, we show that the proposed framework outperforms state-of-theart methods in both KITTI and Cityscapes datasets. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 3 |
| 2021 | Action Relational Graph for Weakly-Supervised Temporal Action LocalizationabstractThe task of weakly-supervised temporal action localization (WTAL) is to recognize plentiful unstructured actions in untrimmed videos with only video-level class labels. As various actions may occur in an untrimmed video, it is desirable to capture the correlation among different actions to effectively identify the target actions. In this paper, we propose a novel Action Relational Graph Network (ARG-Net) to model the correlation between action labels. Specifically, we build a co-occurrence graph using Graph Convolutional Network (GCN), where the graph nodes and edges are represented by word embedding of action labels and relations between two labels, respectively. Then we apply the GCNs to project the action label embeddings into a set of correlated action classifiers which are multiplied with the learned video representations for video-level classification. To facilitate discriminative video representation learning, we employ the attention mechanism to model the probability of a frame containing action instances. A new Action Normalization Loss (ANL) is proposed to further alleviate the confusion from irrelevant background frames (i.e., frames containing no actions). Experimental results on THUMOS14 and ActivityNet1.2 datasets demonstrate that our ARG-Net outperforms the state-of-the-art methods. Ying Sun 0001, Dongyun Lin, Joo-Hwee Lim |
ICIP | 4 |
| 2021 | Enhancing Multi-Step Action Prediction for Active Object DetectionabstractActive vision for robots is one promising solution to open world visual detection problems. A fundamental issue is view planning, i.e., predicting next best views to capture images of interest to reduce uncertainty. While multi-step action in a reinforcement learning (RL) setup can boost the efficiency of view planning, existing methods suffer from unstable detection outcome when the Q-values of multiple branches of action advantages (i.e., action range and action type) are combined naively. To tackle this issue, we propose a novel mechanism to disentangle action range from action type through a two-stage training strategy on a deep Q-network. It combines well-crafted loss functions with respect to action range and action type to enforce separated training of these two branches. We evaluate our method on two public datasets and show that it facilitates substantial gain in view planning efficiency, while enhancing detection accuracy. Fen Fang, Qianli Xu, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 5 |
| 2021 | A Diagnostic Study Of Visual Question Answering With Analogical ReasoningabstractThe deep learning community has made rapid progress in low-level visual perception tasks such as object localization, detection and segmentation. However, for tasks such as Visual Question Answering (VQA) and visual language grounding that require high-level reasoning abilities, huge gaps still exist between artificial systems and human intelligence. In this work, we perform a diagnostic study on recent popular VQA in terms of analogical reasoning. We term it as Analogical VQA, where a system needs to reason on a group of images to find analogical relations among them in order to correctly answer a natural language question. To study the task in depth, we propose an initial diagnostic synthetic dataset CLEVR-Analogy, which tests a range of analogical reasoning abilities (e.g. reasoning on object attributes, spatial relationships, existence, and arithmetic analogies). We benchmark various recent state-of-the-art methods on our dataset and compare the results against human performance, and discover that existing systems fall shorts when facing analogical reasoning involving spatial relationships. The dataset and code will be publicly available to facilitate future research. Hongyuan Zhu 0002, Ying Sun 0001, Dongkyu Choi, Cheston Tan, Joo-Hwee Lim |
ICIP | 6 |
| 2021 | Joint Learning on the Hierarchy Representation for Fine-Grained Human Action RecognitionabstractFine-grained human action recognition is a core research topic in computer vision. Inspired by the recently proposed hierarchy representation of fine-grained actions in FineGym and SlowFast network for action recognition, we propose a novel multi-task network which exploits the FineGym hierarchy representation to achieve effective joint learning and prediction for fine-grained human action recognition. The multi-task network consists of three pathways of SlowOnly networks with gradually increased frame rates for events, sets and elements of fine-grained actions, followed by our proposed integration layers for joint learning and prediction. It is a two-stage approach, where it first learns deep feature representation at each hierarchical level, and is followed by feature encoding and fusion for multi-task learning. Our empirical results on the FineGym dataset achieve a new state-of-the-art performance, with 91.80% Top-1 accuracy and 88.46% mean accuracy for element actions, which are 3.40% and 7.26% higher than the previous best results. Mei Chee Leong, Hui Li Tan, Haosong Zhang 0001, Liyuan Li, Feng Lin 0002, Joo-Hwee Lim |
ICIP | 6 |
| 2021 | Semantic Role Aware Correlation Transformer For Text To Video RetrievalabstractWith the emergence of social media, voluminous video clips are uploaded every day, and retrieving the most relevant visual content with a language query becomes critical. Most approaches aim to learn a joint embedding space for plain textual and visual contents without adequately exploiting their intra-modality structures and inter-modality correlations. This paper proposes a novel transformer that explicitly disentangles the text and video into semantic roles of objects, spatial contexts and temporal contexts with an attention scheme to learn the intra- and inter-role correlations among the three roles to discover discriminative features for matching at different levels. The preliminary results on popular YouCook2 indicate that our approach surpasses a current state-of-the-art method, with a high margin in all metrics. It also overpasses two SOTA methods in terms of two metrics. Burak Satar, Hongyuan Zhu 0002, Xavier Bresson, Joo-Hwee Lim |
ICIP | 4 |
| 2021 | Towards Efficient Multiview Object Detection with Adaptive Action PredictionabstractActive vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection. We formulate the multi-object detection task as an active multiview object detection problem given the initial location of the objects. Next, we propose a novel adaptive action prediction (A2P) method built on a deep Q-learning network with a dueling architecture. The A2P method is able to perform view planning based on visual information of multiple objects; and adjust action ranges according to the task status. Evaluated on the AVD dataset, A2P leads to 21.9% increase in detection accuracy in unfamiliar environments, while improving efficiency by 22.7%. On the T-LESS dataset, multi-object detection boosts efficiency by more than 30% while achieving equivalent detection accuracy. Qianli Xu, Fen Fang, Nicolas Gauthier, Wenyu Liang, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
ICRA | 7 |
| 2021 | Predicting Event Memorability from Contextual Visual SemanticsabstractEpisodic event memory is a key component of human cognition. Predicting event memorability,i.e., to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorability according to a cued recall process. Specifically, we explore whether event memorability is contingent on the event context, as well as the intrinsic visual attributes of image cues. We design a novel experiment protocol and conduct a large-scale experiment with 47 elder subjects over 3 months. Subjects’ memory of life events is tested in a cued recall process. Using advanced visual analytics methods, we build a first-of-its-kind event memorability dataset (called R3) with rich information about event context and visual semantic features. Furthermore, we propose a contextual event memory network (CEMNet) that tackles multi-modal input to predict item-wise event memorability, which outperforms competitive benchmarks. The findings inform deeper understanding of episodic event memory, and open up a new avenue for prediction of human episodic memory. Source code is available at https://github.com/ffzzy840304/Predicting-Event-Memorability. Qianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju, Joo-Hwee Lim |
NeurIPS | 5 |
| 2021 | A comprehensive survey of procedural video datasets
Hui Li Tan, Hongyuan Zhu 0002, Joo-Hwee Lim, Cheston Tan |
Comput. Vis. Image Underst. | 3 |
| 2021 | Single-Image Dehazing via Compositional Adversarial NetworkabstractSingle-image dehazing has been an important topic given the commonly occurred image degradation caused by adverse atmosphere aerosols. The key to haze removal relies on an accurate estimation of global air-light and the transmission map. Most existing methods estimate these two parameters using separate pipelines which reduces the efficiency and accumulates errors, thus leading to a suboptimal approximation, hurting the model interpretability, and degrading the performance. To address these issues, this article introduces a novel generative adversarial network (GAN) for single-image dehazing. The network consists of a novel compositional generator and a novel deeply supervised discriminator. The compositional generator is a densely connected network, which combines fine-scale and coarse-scale information. Benefiting from the new generator, our method can directly learn the physical parameters from data and recover clean images from hazy ones in an end-to-end manner. The proposed discriminator is deeply supervised, which enforces that the output of the generator to look similar to the clean images from low-level details to high-level structures. To the best of our knowledge, this is the first end-to-end generative adversarial model for image dehazing, which simultaneously outputs clean images, transmission maps, and air-lights. Extensive experiments show that our method remarkably outperforms the state-of-the-art methods. Furthermore, to facilitate future research, we create the HazeCOCO dataset which is currently the largest dataset for single-image dehazing. Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Zhao Kang 0001, Shijian Lu, Zhiwen Fang, Liyuan Li, Joo-Hwee Lim |
IEEE Trans. Cybern. | 9 |
| 2021 | Lifelog Image Retrieval Based on Semantic Relevance MappingabstractLifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method. Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2020 | Learning Cross-Modal Representations for Language-Based Image ManipulationabstractIn this paper, we propose a generative architecture for manipulating images/scenes with natural language descriptions. This is a challenging task as the generative network is expected to perform the given text instruction without changing the non-affiliating contents of the input image. Two main drawbacks of the existing methods are their limitation of performing changes that would affect only a limited region and the inability of handling complex instructions. The proposed approach, designed to address these limitations initially uses two sets of networks to extract the image and text features respectively. Rather than a simple combination of these two modalities during the image manipulation process, we use an improved technique to compose image and text features. Additionally, the generative network utilizes similarity learning to improve text manipulation which also enforces only the text-relevant changes on the input image. Our experiments on CSS and Fashion Synthesis datasets show that the proposed approach performs remarkably well and outperforms the baseline frameworks in terms of R-precision and FID. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 3 |
| 2020 | Task-Oriented Multi-Modal Question Answering For Collaborative ApplicationsabstractCobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task. Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 9 |
| 2020 | Active Image Sampling on Canonical Views for Novel Object DetectionabstractTo alleviate the costly data annotation problem in deep learning-based object detection, we leverage the canonical view model for active sample selection to improve the effectiveness of learning. Inspired by the view-approximation model, we hypothesize that visual features learned from canonical views denote better representations of objects, thus boosting the effectiveness of object learning. We validate the hypothesis empirically in the context of robot learning for novel object detection. Based on this, we propose a novel on-line viewpoint exploration (OLIVE) method that (1) defines goodness-of-view by combining informativeness of visual features and consistency of model-based object detection, and (2) systematically explores and selects viewpoints to boost learning efficiency. Furthermore, we train a legacy Faster R-CNN model with a data augmentation method while leveraging data samples generated by the OLIVE pipeline. We test our method on the T-LESS dataset and show that the proposed method outperforms competitive benchmarking methods, especially when the samples are few. Qianli Xu, Fen Fang, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 5 |
| 2020 | Gesture Enhanced Comprehension of Ambiguous Human-to-Robot InstructionsabstractThis work demonstrates the feasibility and benefits of using pointing gestures, a naturally-generated additional input modality, to improve the multi-modal comprehension accuracy of human instructions to robotic agents for collaborative tasks.We present M2Gestic, a system that combines neural-based text parsing with a novel knowledge-graph traversal mechanism, over a multi-modal input of vision, natural language text and pointing. Via multiple studies related to a benchmark table top manipulation task, we show that (a) M2Gestic can achieve close-to-human performance in reasoning over unambiguous verbal instructions, and (b) incorporating pointing input (even with its inherent location uncertainty) in M2Gestic results in a significant (30%) accuracy improvement when verbal instructions are ambiguous. Dulanga Weerakoon, Vigneshwaran Subbaraju, Nipuni Karumpulli, Qianli Xu, U-Xuan Tan, Joo-Hwee Lim, Archan Misra |
ICMI | 7 |
| 2020 | 6D Pose Estimation with Correlation Fusionabstract6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illumination, so it is important to complement them with depth information. However, existing methods using RGB-D data cannot adequately exploit consistent and complementary information between RGB and depth modalities. In this paper, we present a novel method to effectively consider the correlation within and across both modalities with attention mechanism to learn discriminative and compact multi-modal features. Then, effective fusion strategies for intra- and inter-correlation modules are explored to ensure efficient information flow between RGB and depth. To our best knowledge, this is the first work to explore effective intra- and inter-modality fusion in 6D pose estimation. The experimental results show that our method can achieve the state-of-the-art performance on LineMOD and YCB-Video dataset. We also demonstrate that the proposed method can benefit a real-world robot grasping task by providing accurate object pose estimation. Hongyuan Zhu 0002, Ying Sun 0001, Cihan Acar, Yan Wu 0002, Liyuan Li, Cheston Tan, Joo-Hwee Lim |
ICPR | 9 |
| 2020 | Detecting Objects with High Object Region PercentageabstractObject shape is a subtle but important factor for object detection. It has been observed that the object-region-percentage (ORP) can be utilized to improve detection accuracy for elongated objects, which have much lower ORPs than other types of objects. In this paper, we propose an approach to improve the detection performance for objects with high ORPs. Our method consists of three steps. First, we adjust the ground truth bounding boxes of high-ORP objects to an optimal range. Second, we train an object detector, Faster R-CNN, based on adjusted bounding boxes to achieve high recall. Finally, we train a DCNN to learn the adjustment ratios towards four directions and adjust detected bounding boxes of objects to get better localization for higher precision. We evaluate the effectiveness of our method on 12 high-ORP objects in COCO and 8 objects in a proprietary gearbox dataset. The experimental results show that our method can achieve state-of-the-art performance on these objects while costing less resources in training and inference stages. Fen Fang, Qianli Xu, Liyuan Li, Joo-Hwee Lim |
ICPR | 5 |
| 2020 | A novel hybrid approach for crack detection
Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
Pattern Recognit. | 5 |
| 2020 | Semantically consistent text to fashion image synthesis with an enhanced attentional generative adversarial network
Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim |
Pattern Recognit. Lett. | 2 |
| 2020 | Combining Faster R-CNN and Model-Driven Clustering for Elongated Object DetectionabstractWhile analyzing the performance of state-of-the-art R-CNN based generic object detectors, we find that the detection performance for objects with low object-region-percentages (ORPs) of the bounding boxes are much lower than the overall average. Elongated objects are examples. To address the problem of low ORPs for elongated object detection, we propose a hybrid approach which employs a Faster R-CNN to achieve robust detections of object parts, and a novel model-driven clustering algorithm to group the related partial detections and suppress false detections. First, we train a Faster R-CNN with partial region proposals of suitable and stable ORPs. Next, we introduce a deep CNN (DCNN) for orientation classification on the partial detections. Then, on the outputs of the Faster R-CNN and DCNN, the algorithm of adaptive model-driven clustering first initializes a model of an elongated object with a data-driven process on local partial detections, and refines the model iteratively by model-driven clustering and data-driven model updating. By exploiting Faster R-CNN to produce robust partial detections and model-driven clustering to form a global representation, our method is able to generate a tight oriented bounding box for elongated object detection. We evaluate the effectiveness of our approach on two typical elongated objects in the COCO dataset, and other typical elongated objects, including rigid objects (pens, screwdrivers and wrenches) and non-rigid objects (cracks). Experimental results show that, compared with the state-of-the-art approaches, our method achieves a large margin of improvements for both detection and localization of elongated objects in images. Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
IEEE Trans. Image Process. | 4 |
| 2019 | Singe Image Rain Removal with Unpaired Information: A Differentiable Programming PerspectiveabstractSingle image rain-streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. Previous works solve this problem using various hand-designed priors or by explicitly mapping synthetic rain to paired clean image in a supervised way. In practice, however, the pre-defined priors are easily violated and the paired training data are hard to collect. To overcome these limitations, in this work, we propose RainRemoval-GAN (RRGAN), the first end-to-end adversarial model that generates realistic rain-free images using only unpaired supervision. Our approach alleviates the paired training constraints by introducing a physical-model which explicitly learns a recovered images and corresponding rain-streaks from the differentiable programming perspective. The proposed network consists of a novel multiscale attention memory generator and a novel multiscale deeply supervised discriminator. The multiscale attention memory generator uses a memory with attention mechanism to capture the latent rain streaks context at different stages to recover the clean images. The deeply supervised multiscale discriminator imposes constraints at the recovered output in terms of local details and global appearance to the clean image set. Together with the learned rainstreaks, a reconstruction constraint is employed to ensure the appearance consistent with the input image. Experimental results on public benchmark demonstrates our promising performance compared with nine state-of-the-art methods in terms of PSNR, SSIM, visual qualities and running time. Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Songfan Yang, Vijay Chanderasekh, Liyuan Li, Joo-Hwee Lim |
AAAI | 7 |
| 2019 | Attribute Manipulation Generative Adversarial Networks for Fashion ImagesabstractRecent advances in Generative Adversarial Networks (GANs) have made it possible to conduct multi-domain image-to-image translation using a single generative network. While recent methods such as Ganimation and SaGAN are able to conduct translations on attribute-relevant regions using attention, they do not perform well when the number of attributes increases as the training of attention masks mostly rely on classification losses. To address this and other limitations, we introduce Attribute Manipulation Generative Adversarial Networks (AMGAN) for fashion images. While AMGAN's generator network uses class activation maps (CAMs) to empower its attention mechanism, it also exploits perceptual losses by assigning reference (target) images based on attribute similarities. AMGAN incorporates an additional discriminator network that focuses on attribute-relevant regions to detect unrealistic translations. Additionally, AMGAN can be controlled to perform attribute manipulations on specific regions such as the sleeve or torso regions. Experiments show that AMGAN outperforms state-of-the-art methods using traditional evaluation metrics as well as an alternative one that is based on image retrieval. Kenan E. Ak, Ashraf A. Kassim, Joo-Hwee Lim, Jo Yew Tham |
ICCV | 3 |
| 2019 | Towards Real-Time Crack Detection Using a Deep Neural Network With a Bayesian Fusion AlgorithmabstractSurface cracks can represent very small and thin objects in images. With irregular shapes and sizes, and non-fixed texture patterns, the detection of cracks can be a challenging problem in computer vision. Prior work has been undertaken on detecting cracks for images using a sliding window mode. However, such methods can be time consuming, and result in high false alarms. To help address this problem, a new crack detection and segmentation method is proposed in this paper. Specifically, our method includes three main features: (1) a Faster R-CNN model to detect crack patches in images; (2) the use of a Bayesian fusion algorithm to suppress false alarms based on detected patch orientation; and (3) image processing functions to obtain final segmentation masks, such as for Gaussian blur, erosion, etc. Experimental results show that our method can achieve high detection accuracy on sampled images in real-time. Fen Fang, Liyuan Li, Mark D. Rice, Joo-Hwee Lim |
ICIP | 4 |
| 2019 | An Adaptive Fitting Approach for the Visual Detection and Counting of Small Circular Objects in Manufacturing ApplicationsabstractDetecting, localizing and counting small circular objects in machine parts is an important task in many applications for manufacturing. Existing methods of circle detection face difficulties due to the high-curvature and limited edge points of circles. As a result, in this paper we propose a novel two-stage circle detection method, which integrates bottom-up coarse detection and top-down circle fitting. First, a circle detector combining low-level feature descriptors and a linear SVM is developed. This is used to scan an input image in a sliding window mode to detect small circles with coarse estimates of locations and scales. Next, a hierarchical Bayesian model performs a top-down adaptive circle fitting, with the ability to achieve a maximum a posteriori probability to fit circles to local image features. The evaluation of our approach with manufacturing images has demonstrated to be efficient in detecting small circles in machine parts. Liyuan Li, Fen Fang, Mark D. Rice, Jamie Ng, Wei Xiong 0001, Joo-Hwee Lim |
ICIP | 7 |
| 2019 | Towards Robust Retrieval for Imperfectly Scanned Point Cloud ObjectsabstractWith the development of 3D model analysis and particularly on classification challenge, the algorithms are getting better and better. Since no large dataset of scanned models is available, evaluating the algorithms in real-life scenarios is not straightforward. For now, these studies rely on ModelNet, a dataset of CAD models. Moreover, no studies considered the robustness to common recording data corruptions like occlusion or noise. In this paper, we present a preliminary study to assess the retrieval performances of point cloud-based algorithms for occluded or noisy objects. The experiment shows very promising results, even in the case of deep learning networks which has been pre-trained with CAD models. Indeed, more than 90% of the objects are retrieved when even only 40% of the query object is visible. Justin Lev, Joo-Hwee Lim, Nizar Ouarti, Mounir Mokhtari |
ICIP | 2 |
| 2019 | Which Body Is Mine?abstractIn the light of the human studies that report a strong correlation between head circumference and body size, we propose a new research problem: head-body matching. Given an image of a person's head, we want to match it with his body (headless) image. We propose a dual-pathway framework which computes head and body discriminating features independently, and learns the correlation between such features. We introduce a comprehensive evaluation of our proposed framework for this problem using different features including anthropometric features and deep-CNN features, different experimental setting such as head-body scale variations, and different body parts. We demonstrate the usefulness of our framework with two novel applications: head/body recognition, and T-shirt sizing from a head image. Our evaluations for head/body recognition application on the challenging large scale PIPA dataset (contains high variations of pose, viewpoint, and occlusion) show up to 53% of performance improvement using deep-CNN features, over the global model features in which head and body features are not separated or correlated. For T-shirt sizing application, we use anthropometric features for head-body matching. We achieve promising experimental results on small and challenging datasets. Mona Ragab, Terence Sim, Joo-Hwee Lim, Keng Teck Ma |
WACV | 3 |
| 2019 | Anticipating Where People will Look Using Adversarial NetworksabstractWe introduce a new problem of gaze anticipation on future frames which extends the conventional gaze prediction problem to go beyond current frames. To solve this problem, we propose a new generative adversarial network based model, Deep Future Gaze (DFG), encompassing two pathways: DFG-P is to anticipate gaze prior maps conditioned on the input frame which provides task influences; DFG-G is to learn to model both semantic and motion information in future frame generation. DFG-P and DFG-G are then fused to anticipate future gazes. DFG-G consists of two networks: a generator and a discriminator. The generator uses a two-stream spatial-temporal convolution architecture (3D-CNN) for explicitly untangling the foreground and background to generate future frames. It then attaches another 3D-CNN for gaze anticipation based on these synthetic frames. The discriminator plays against the generator by distinguishing the synthetic frames of the generator from the real frames. Experimental results on the publicly available egocentric and third person video datasets show that DFG significantly outperforms all competitive baselines. We also demonstrate that DFG achieves better performance of gaze prediction on current frames in egocentric and third person videos than state-of-the-art methods. Mengmi Zhang, Keng Teck Ma, Joo-Hwee Lim, Qi Zhao 0001, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Learning Attribute Representations With Localization for Flexible Fashion SearchabstractIn this paper, we investigate ways of conducting a detailed fashion search using query images and attributes. A credible fashion search platform should be able to (1) find images that share the same attributes as the query image, (2) allow users to manipulate certain attributes, e.g. replace collar attribute from round to v-neck, and (3) handle region-specific attribute manipulations, e.g. replacing the color attribute of the sleeve region without changing the color attribute of other regions. A key challenge to be addressed is that fashion products have multiple attributes and it is important for each of these attributes to have representative features. To address these challenges, we propose the FashionSearchNet which uses a weakly supervised localization method to extract regions of attributes. By doing so, unrelated features can be ignored thus improving the similarity learning. Also, FashionSearchNet incorporates a new procedure that enables region awareness to be able to handle region-specific requests. FashionSearchNet outperforms the most recent fashion search techniques and is shown to be able to carry out different search scenarios using the dynamic queries. Kenan E. Ak, Ashraf A. Kassim, Joo-Hwee Lim, Jo Yew Tham |
CVPR | 3 |
| 2018 | DehazeGAN: When Image Dehazing Meets Differential ProgrammingabstractSingle image dehazing has been a classic topic in computer vision for years. Motivated by the atmospheric scattering model, the key to satisfactory single image dehazing relies on an estimation of two physical parameters, i.e., the global atmospheric light and the transmission coefficient. Most existing methods employ a two-step pipeline to estimate these two parameters with heuristics which accumulate errors and compromise dehazing quality. Inspired by differentiable programming, we re-formulate the atmospheric scattering model into a novel generative adversarial network (DehazeGAN). Such a reformulation and adversarial learning allow the two parameters to be learned simultaneously and automatically from data by optimizing the final dehazing performance so that clean images with faithful color and structures are directly produced. Moreover, our reformulation also greatly improves the GAN’s interpretability and quality for single image dehazing. To the best of our knowledge, our method is one of the first works to explore the connection among generative adversarial models, image dehazing, and differentiable programming, which advance the theories and application of these areas. Extensive experiments on synthetic and realistic data show that our method outperforms state-of-the-art methods in terms of PSNR, SSIM, and subjective visual quality. Hongyuan Zhu 0002, Xi Peng 0001, Vijay Chandrasekhar 0001, Liyuan Li, Joo-Hwee Lim |
IJCAI | 5 |
| 2018 | Egocentric Spatial MemoryabstractEgocentric spatial memory (ESM) defines a memory system with encoding, storing, recognizing and recalling the spatial information about the environment from an egocentric perspective. We introduce an integrated deep neural network architecture for modeling ESM. It learns to estimate the occupancy state of the world and progressively construct top-down 2D global maps from egocentric views in a spatially extended environment. During the exploration, our proposed ESM model updates belief of the global map based on local observations using a recurrent neural network. It also augments the local mapping with a novel external memory to encode and store latent representations of the visited places over longterm exploration in large environments which enables agents to perform place recognition and hence, loop closure. Our proposed ESM network contributes in the following aspects: (1) without feature engineering, our model predicts free space based on egocentric views efficiently in an end-to-end manner; (2) different from other deep learning-based mapping system, ESMN deals with continuous actions and states which is vitally important for robotic control in real applications. In the experiments, we demonstrate its accurate and robust global mapping capacities in 3D virtual mazes and realistic indoor environments by comparing with several competitive baselines. Mengmi Zhang, Keng Teck Ma, Shih-Cheng Yen, Joo-Hwee Lim, Qi Zhao 0001, Jiashi Feng |
IROS | 4 |
| 2018 | Predicting Visual Context for Unsupervised Event Segmentation in Continuous Photo-streamsabstractSegmenting video content into events provides semantic structures for indexing, retrieval, and summarization. Since motion cues are not available in continuous photo-streams, and annotations in lifelogging are scarce and costly, the frames are usually clustered into events by comparing the visual features between them in an unsupervised way. However, such methodologies are ineffective to deal with heterogeneous events, e.g. taking a walk, and temporary changes in the sight direction, e.g. at a meeting. To address these limitations, we propose Contextual Event Segmentation (CES), a novel segmentation paradigm that uses an LSTM-based generative network to model the photo-stream sequences, predict their visual context, and track their evolution. CES decides whether a frame is an event boundary by comparing the visual context generated from the frames in the past, to the visual context predicted from the future. We implemented CES on a new and massive lifelogging dataset consisting of more than 1.5 million images spanning over 1,723 days. Experiments on the popular EDUB-Seg dataset show that our model outperforms the state-of-the-art by over 16% in f-measure. Furthermore, CES' performance is only 3 points below that of human annotators. Ana Garcia del Molino, Joo-Hwee Lim, Ah-Hwee Tan |
ACM Multimedia | 2 |
| 2018 | Personalized Serious Games for Cognitive Intervention with Lifelog Visual AnalyticsabstractThis paper presents a novel serious game app and a method to cre- ate and integrate personalized game content based on lifelog visual analytics. The main objective is to extract personalized content from visual lifelogs, integrate it into mobile games, and evaluate the effect of personalization on user experience. First, a suite of visual analysis methods is proposed to extract semantic informa- tion from visual lifelogs and discover the association among the lifelog entities. The outcome is dataset that contains augmented and personal lifelog images. Next, a mobile game app is developed that makes use of the dataset as game content. Finally, an experiment is conducted to evaluate user gameplay behaviors in the wild over three months, where a mixture of generic and personalized game content is deployed. It is observed that user adherence is heightened by personalized game content as compared to generic content. Also observed is a higher enjoyment level in personalized than generic game content. The result provides the first empirical evidence of the effect of personalized games on user adherence and preference for cognitive intervention. This work paves the way for effective cognitive training with user-generated content. Qianli Xu, Vigneshwaran Subbaraju, Chee How Cheong, Aijing Wang, Kathleen Kang, Munirah Bashir, Yanhong Dong, Liyuan Li, Joo-Hwee Lim |
ACM Multimedia | 9 |
| 2018 | Efficient Multi-attribute Similarity Learning Towards Attribute-Based Fashion SearchabstractIn this paper, we propose an attribute-based query & retrieval system designed for fashion products. Our system addresses the problem of carrying out fashion searches by the query image and attribute manipulation, e.g. replacing long sleeve attribute of a dress to sleeveless. We present the attributes in two groups: (1) general attributes (category, gender etc.) and (2) special attributes (sleeve length, collar etc.). The special attributes are more suitable for the attribute manipulation and thus conducting searches. In order to solve the mentioned fashion search problem, it is crucial for the deep neural networks to understand attribute similarities. To facilitate more specific similarity learning, clothing items are represented by their structural subcomponents or "parts". The parts are estimated using an unsupervised segmentation method and used inside the proposed Convolutional Neural Network (CNN) as an attention mechanism. Meaning, different parts are connected to the special attributes, e.g. sleeve part is connected with sleeve length attribute. With this mechanism, part-based triplet ranking constraint is applied to learn similarity of each special attribute independently from one another in a single network. In the end, the well-defined features are used to conduct the fashion search. Additionally, an adaptive relevance feedback module is used to personalize the fashion search process with the feature descriptions. For our experiments, a new dataset is constructed containing 101,021 images which consist of pure clothing items. Besides achieving decent retrieval results in our dataset, the experiments show that proposed technique outperforms different baselines and is able to adapt towards user's requests. Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim |
WACV | 2 |
| 2018 | Which shirt for my first date? Towards a flexible attribute-based fashion query system
Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim |
Pattern Recognit. Lett. | 2 |
| 2018 | A Probabilistic Model of Social Working Memory for Information Retrieval in Social InteractionsabstractSocial working memory (SWM) plays an important role in navigating social interactions. Inspired by studies in psychology, neuroscience, cognitive science, and machine learning, we propose a probabilistic model of SWM to mimic human social intelligence for personal information retrieval (IR) in social interactions. First, we establish a semantic hierarchy as social long-term memory to encode personal information. Next, we propose a semantic Bayesian network as the SWM, which integrates the cognitive functions of accessibility and self-regulation. One subgraphical model implements the accessibility function to learn the social consensus about IR-based on social information concept, clustering, social context, and similarity between persons. Beyond accessibility, one more layer is added to simulate the function of self-regulation to perform the personal adaptation to the consensus based on human personality. Two learning algorithms are proposed to train the probabilistic SWM model on a raw dataset of high uncertainty and incompleteness. One is an efficient learning algorithm of Newton's method, and the other is a genetic algorithm. Systematic evaluations show that the proposed SWM model is able to learn human social intelligence effectively and outperforms the baseline Bayesian cognitive model. Toward real-world applications, we implement our model on Google Glass as a wearable assistant for social interaction. Liyuan Li, Qianli Xu, Tian Gan 0002, Cheston Tan, Joo-Hwee Lim |
IEEE Trans. Cybern. | 5 |
| 2017 | Active Video Summarization: Customized Summaries via On-line Interaction with the UserabstractTo facilitate the browsing of long videos, automatic video summarization provides an excerpt that represents its content. In the case of egocentric and consumer videos, due to their personal nature, adapting the summary to specific user's preferences is desirable. Current approaches to customizable video summarization obtain the user's preferences prior to the summarization process. As a result, the user needs to manually modify the summary to further meet the preferences. In this paper, we introduce Active Video Summarization (AVS), an interactive approach to gather the user's preferences while creating the summary. AVS asks questions about the summary to update it on-line until the user is satisfied. To minimize the interaction, the best segment to inquire next is inferred from the previous feedback. We evaluate AVS in the commonly used UTEgo dataset. We also introduce a new dataset for customized video summarization (CSumm) recorded with a Google Glass. The results show that AVS achieves an excellent compromise between usability and quality. In 41% of the videos, AVS is considered the best over all tested baselines, including summaries manually generated. Also, when looking for specific events in the video, AVS provides an average level of satisfaction higher than those of all other baselines after only six questions to the user. Ana Garcia del Molino, Xavier Boix, Joo-Hwee Lim, Ah-Hwee Tan |
AAAI | 3 |
| 2017 | Deep Future Gaze: Gaze Anticipation on Egocentric Videos Using Adversarial NetworksabstractWe introduce a new problem of gaze anticipation on egocentric videos. This substantially extends the conventional gaze prediction problem to future frames by no longer confining it on the current frame. To solve this problem, we propose a new generative adversarial neural network based model, Deep Future Gaze (DFG). DFG generates multiple future frames conditioned on the single current frame and anticipates corresponding future gazes in next few seconds. It consists of two networks: generator and discriminator. The generator uses a two-stream spatial temporal convolution architecture (3D-CNN) explicitly untangling the foreground and the background to generate future frames. It then attaches another 3D-CNN for gaze anticipation based on these synthetic frames. The discriminator plays against the generator by differentiating the synthetic frames of the generator from the real frames. Through competition with discriminator, the generator progressively improves quality of the future frames and thus anticipates future gaze better. Experimental results on the publicly available egocentric datasets show that DFG significantly outperforms all well-established baselines. Moreover, we demonstrate that DFG achieves better performance of gaze prediction on current frames than state-of-the-art methods. This is due to benefiting from learning motion discriminative representations in frame generation. We further contribute a new egocentric dataset (OST) in the object search task. DFG also achieves the best performance for this challenging dataset. Mengmi Zhang, Keng Teck Ma, Joo-Hwee Lim, Qi Zhao 0001, Jiashi Feng |
CVPR | 3 |
| 2017 | Principal curvature of point cloud for 3D shape recognitionabstractIn the recent years, we experienced the proliferation of sensors for retrieving depth information on a scene, such as LIDAR or RGBD sensors (Kinect). However, it is still a challenge to identify the meaning of a specific point cloud to recognize the underlying object. Here, we wonder if it is possible to define a global feature for an object that is robust to noise, sampling and occlusion. We propose a local measure based on curvature. We called it Principal Curvature because rather than using the Gaussian curvature we keep the information of the two principal curvatures. In our approach, this local information is then aggregated as histograms that are compared with a Chi-2 metric. Results show the robustness of the method particularly when only few points are available. This means that our approach can be very suitable to match objects even with a limited resolution and possible occlusions. It could be particularly adapted to recognize objects with LIDAR inputs. Justin Lev, Joo-Hwee Lim, Nizar Ouarti |
ICIP | 2 |
| 2017 | Multi-layer linear model for top-down modulation of visual attention in natural egocentric visionabstractTop-down attention plays an important role in guidance of human attention in real-world scenarios, but less efforts in computational modeing of visual attention has been put on it. Inspired by the mechanisms of top-down attention in human visual perception, we propose a multi-layer linear model of top-down attention to modulate bottom-up saliency maps actively. The first layer is a linear regression model which combines the bottom-up saliency maps on various visual features and objects. A contextual dependent upper layer is introduced to tune the parameters of the lower layer model adaptively. Finally, a mask of selection history is applied to the fused attention map to bias the attention selection towards the task related regions. Efficient learning algorithm with single-pass polynomial complexity is derived. We evaluate our model on a set of natural egocentric videos captured from a wearable glass in real-world environments. Our model outperforms the baseline and state-of-the-art bottom-up saliency models. Keng Teck Ma, Liyuan Li, Peilun Dai, Joo-Hwee Lim, Chengyao Shen, Qi Zhao 0001 |
ICIP | 4 |
| 2017 | Foveated neural network: Gaze prediction on egocentric videosabstractA novel deep convolution neural network is proposed to predict gaze on current frames in egocentric videos. Inspired by human visual system, we introduce a fovea module responsible for sharp central vision and name our model as Foveated Neural Network (FNN). The retina-like visual inputs from the region of interest on the previous frame are analysed and encoded. The fusion of the hidden representations of the previous frame and the feature maps of the current frame guides the gaze prediction on the current frame. In order to simulate motion, we also include the dense optical flow between these adjacent frames as additional input. Experimental results show that FNN outperforms the state-of-the-art algorithms in the publicly available egocentric dataset. The analysis of FNN demonstrates that the hidden representations of the foveated visual input from the previous frame as well as the motion information between adjacent frames are efficient in improving gaze prediction performance in egocentric videos. Mengmi Zhang, Keng Teck Ma, Joo-Hwee Lim, Qi Zhao 0001 |
ICIP | 3 |
| 2017 | The effect of different types of navigation assistance on indoor scene memorabilityabstractWith the rapid growing of wearable computing devices, indoor navigation guidance will become popular in the near future like the GPS-based navigation tools for drivers today. However, how the guided indoor navigation affects human’s memory of a novel environment has not been well studied. In this paper, we investigate route memory with three types of navigation assistance, that is, 2D map, wearable navigation assistant, and human usher. Twenty participants were asked to remember the route while being guided through a novel indoor environment. Our results show that the participants have similar patterns in remembering visual scenes, even using different types of assistance. These findings support previous work on scene memorability and provide the new insight that scene memorability is not affected by the type of navigation guidance. This may indicate that spatial working memory and visual memory are dissociated. We also show that scenes with navigation information are more memorable than scenes without such information. Finally, we provide some evidence that the location of a scene is linked to its memorability. In general, our findings provide valuable information about indoor scene memorability. Michal Mukawa, Cheston Tan, Joo-Hwee Lim, Qianli Xu, Liyuan Li |
Behav. Inf. Technol. | 3 |
| 2017 | A Wearable Virtual Usher for Vision-Based Cognitive Indoor NavigationabstractInspired by progresses in cognitive science, artificial intelligence, computer vision, and mobile computing technologies, we propose and implement a wearable virtual usher for cognitive indoor navigation based on egocentric visual perception. A novel computational framework of cognitive wayfinding in an indoor environment is proposed, which contains a context model, a route model, and a process model. A hierarchical structure is proposed to represent the cognitive context knowledge of indoor scenes. Given a start position and a destination, a Bayesian network model is proposed to represent the navigation route derived from the context model. A novel dynamic Bayesian network (DBN) model is proposed to accommodate the dynamic process of navigation based on real-time first-person-view visual input, which involves multiple asynchronous temporal dependencies. To adapt to large variations in travel time through trip segments, we propose an online adaptation algorithm for the DBN model, leading to a self-adaptive DBN. A prototype system is built and tested for technical performance and user experience. The quantitative evaluation shows that our method achieves over 13% improvement in accuracy as compared to baseline approaches based on hidden Markov model. In the user study, our system guides the participants to their destinations, emulating a human usher in multiple aspects. Liyuan Li, Qianli Xu, Vijay Chandrasekhar 0001, Joo-Hwee Lim, Cheston Tan, Michal Mukawa |
IEEE Trans. Cybern. | 4 |
| 2017 | Summarization of Egocentric Videos: A Comprehensive SurveyabstractThe introduction of wearable video cameras (e.g., GoPro) in the consumer market has promoted video life-logging, motivating users to generate large amounts of video data. This increasing flow of first-person video has led to a growing need for automatic video summarization adapted to the characteristics and applications of egocentric video. With this paper, we provide the first comprehensive survey of the techniques used specifically to summarize egocentric videos. We present a framework for first-person view summarization and compare the segmentation methods and selection algorithms used by the related work in the literature. Next, we describe the existing egocentric video datasets suitable for summarization and, then, the various evaluation methods. Finally, we analyze the challenges and opportunities in the field and propose new lines of research. Ana Garcia del Molino, Cheston Tan, Joo-Hwee Lim, Ah-Hwee Tan |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2015 | Whole space subclass discriminant analysis for face recognitionabstractIn this work, we propose to divide each class (a person) into subclasses using spatial partition trees which helps in better capturing the intra-personal variances arising from the appearances of the same individual. We perform a comprehensive analysis on within-class and within-subclass eigen-spectrums of face images and propose a novel method of eigen-spectrum modeling which extracts discriminative features of faces from both within-subclass and total or between-subclass scatter matrices. Effective low-dimensional face discriminative features are extracted for face recognition (FR) after performing discriminant evaluation in the entire eigenspace. Experimental results on popular face databases (AR, FERET) and the challenging unconstrained YouTube Face database show the superiority of our proposed approach on all three databases. Bappaditya Mandal, Liyuan Li, Vijay Chandrasekhar 0001, Joo-Hwee Lim |
ICIP | 4 |
| 2015 | Scene text extraction based on edges and support vector regression
Shijian Lu, Tao Chen 0003, Shangxuan Tian, Joo-Hwee Lim, Chew Lim Tan |
Int. J. Document Anal. Recognit. | 4 |
| 2014 | A Three-Color Coupled Level-Set Algorithm for Simultaneous Multiple Cell Segmentation and Tracking
Jierong Cheng, Wei Xiong 0001, Shue-Ching Chia, Yue Wang 0005, Joo-Hwee Lim |
ACCV (3) | 6 |
| 2014 | Incremental Graph Clustering for Efficient Retrieval from Streaming Egocentric Video DataabstractWith wearable devices like Google Glass, it will soon become possible to record everything we see. We envision a system where one's entire visual memory is captured, stored and indexed. One of the biggest challenges is the scale of the retrieval problem. In this work, we focus on how to organize streaming egocentric video data. Egocentric video data is highly redundant, in that, we see several objects and scenes repeatedly as we go about our lives. To exploit this redundancy, we propose an evolving sparse-graph representation for egocentric video data. We propose an incremental local density clustering scheme, which learns salient objects and scenes for streaming egocentric video data. We use the density clustering scheme to prune redundant data in the database. For image-retrieval applications, by retaining only representative nodes from dense sub graphs in the streaming data source, we show we can achieve 90% of peak recall by retaining only 1% of data, with a significant 18% improvement in absolute recall over naive uniform sub sampling of the egocentric video data. Vijay Chandrasekhar 0001, Cheston Tan, Wu Min, Liyuan Li, Xiaoli Li 0001, Joo-Hwee Lim |
ICPR | 6 |
| 2014 | Character Recognition in Natural Scenes Using Convolutional Co-occurrence HOGabstractRecognition of characters in natural images is a challenging task due to the complex background, variations of text size and perspective distortion, etc. Traditional optical character recognition (OCR) engine cannot perform well on those unconstrained text images. A novel technique is proposed in this paper that makes use of convolutional cooccurrence histogram of oriented gradient (ConvCoHOG), which is more robust and discriminative than both the histogram of oriented gradient (HOG) and the co-occurrence histogram of oriented gradients (CoHOG). In the proposed technique, a more informative feature is constructed by exhaustively extracting features from every possible image patches within character images. Experiments on two public datasets including the ICDAr 2003 Robust Reading character dataset and the Street View Text (SVT) dataset, show that our proposed character recognition technique obtains superior performance compared with state-of-the-art techniques. Bolan Su, Shijian Lu, Shangxuan Tian, Joo-Hwee Lim, Chew Lim Tan |
ICPR | 4 |
| 2014 | A wearable virtual guide for context-aware cognitive indoor navigationabstractIn this paper, we explore a new way to provide context-aware assistance for indoor navigation using a wearable vision system. We investigate how to represent the cognitive knowledge of wayfinding based on first-person-view videos in real-time and how to provide context-aware navigation instructions in a human-like manner. Inspired by the human cognitive process of wayfinding, we propose a novel cognitive model that represents visual concepts as a hierarchical structure. It facilitates efficient and robust localization based on cognitive visual concepts. Next, we design a prototype system that provides intelligent context-aware assistance based on the cognitive indoor navigation knowledge model. We conducted field tests and evaluated the system's efficacy by benchmarking it against traditional 2D maps and human guidance. The results show that context-awareness built on cognitive visual perception enables the system to emulate the efficacy of a human guide, leading to positive user experience. Qianli Xu, Liyuan Li, Joo-Hwee Lim, Cheston Tan, Michal Mukawa, Gang S. Wang |
Mobile HCI | 3 |
| 2014 | Robust and Efficient Saliency Modeling from Image Co-Occurrence HistogramsabstractThis paper presents a visual saliency modeling technique that is efficient and tolerant to the image scale variation. Different from existing approaches that rely on a large number of filters or complicated learning processes, the proposed technique computes saliency from image histograms. Several two-dimensional image co-occurrence histograms are used, which encode not only "how many" (occurrence) but also "where and how" (co-occurrence) image pixels are composed into a visual image, hence capturing the "unusualness" of an object or image region that is often perceived by either global "uncommonness" (i.e., low occurrence frequency) or local "discontinuity" with respect to the surrounding (i.e., low co-occurrence frequency). The proposed technique has a number of advantageous characteristics. It is fast and very easy to implement. At the same time, it involves minimal parameter tuning, requires no training, and is robust to image scale variation. Experiments on the AIM dataset show that a superior shuffled AUC (sAUC) of 0.7221 is obtained, which is higher than the state-of-the-art sAUC of 0.7187. Shijian Lu, Cheston Tan, Joo-Hwee Lim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Extended Spectral Regression for efficient scene recognition
Liyuan Li, Weixun Goh, Joo-Hwee Lim, Sinno Jialin Pan |
Pattern Recognit. | 3 |
| 2014 | Learning Deep Hierarchical Visual Feature CodingabstractIn this paper, we propose a hybrid architecture that combines the image modeling strengths of the bag of words framework with the representational power and adaptability of learning deep architectures. Local gradient-based descriptors, such as SIFT, are encoded via a hierarchical coding scheme composed of spatial aggregating restricted Boltzmann machines (RBM). For each coding layer, we regularize the RBM by encouraging representations to fit both sparse and selective distributions. Supervised fine-tuning is used to enhance the quality of the visual representation for the categorization task. We performed a thorough experimental evaluation using three image categorization data sets. The hierarchical coding scheme achieved competitive categorization accuracies of 79.7% and 86.4% on the Caltech-101 and 15-Scenes data sets, respectively. The visual representations learned are compact and the model's inference is fast, as compared with sparse coding methods. The low-level representations of descriptors that were learned using this method result in generic features that we empirically found to be transferrable between different image data sets. Further analysis reveal the significance of supervised fine-tuning when the architecture has two layers of representations as opposed to a single layer. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | Visual Recognition using a Combination of Shape and Color Features
Sepehr Jalali, Cheston Tan, Joo-Hwee Lim, Jo Yew Tham, Sim Heng Ong, Paul J. Seekings, Elizabeth A. Taylor |
CogSci | 3 |
| 2013 | Encoding Co-occurrence of Features in the HMAX Model
Sepehr Jalali, Cheston Tan, Joo-Hwee Lim, Jo Yew Tham, Sim Heng Ong, Paul J. Seekings, Elizabeth A. Taylor |
CogSci | 3 |
| 2013 | A Wearable Cognitive Vision System for Navigation Assistance in Indoor Environment
Liyuan Li, Gang S. Wang, Weixun Goh, Joo-Hwee Lim, Cheston Tan |
ICONIP (3) | 4 |
| 2013 | The use of optical and sonar images in the human and dolphin brain for image classificationabstractIn this paper we propose a new biologically inspired model which simulates the visual pathways in the human brain used for classification of matching optical and sonar derived images. Marine mammals, such as dolphins, that live in waters with poor optical clarity and low light levels such as littoral zones, use a combination of optical vision and biosonar to navigate and hunt for prey. Given that dolphins have evolved a synergistic combination of optical visual input and acoustic/sonar input, the primary focus of this paper is on reaching a similar level of synergy for a diver or Autonomous Underwater Vehicle (AUV) platform equipped with a system to extend the range and resolution of vision in poor ambient visibility. We propose a biologically inspired model that combines and processes visual images acquired via optical and acoustic pathways and show that the combined model enhances the accuracy of automatic classification of target objects in underwater images. Sepehr Jalali, Paul J. Seekings, Cheston Tan, Aiswarya Ratheesh, Joo-Hwee Lim, Elizabeth A. Taylor |
IJCNN | 5 |
| 2013 | Classification of marine organisms in underwater images using CQ-HMAX biologically inspired color approachabstractIn many coastal environments, particularly in tropical zones, coral reef ecosystems have exceptional biodiversity, contribute to coastal defense, provide unique and important habitats and valuable commercial resources. Assessment of environmental impacts on biodiversity in such areas are increasingly important to mitigate potential adverse effects on specific ecosystems. Visual classification of marine organisms is necessary for population estimates of individual species of corals or other benthic organisms. In this paper, we introduce a new image dataset of benthic organisms that are of different colors, shapes, scales, visibility and are taken from different viewpoints. We evaluate several different classification approaches on this dataset, and show that CQ-HMAX, our new biologically inspired approach to utilizing color information for object and scene recognition, that is inspired by the characteristics of color- and object-selective neurons in the high-level inferotemporal (IT) cortex of the primate visual system, results in better classification results in comparison with existing computational models such as support vectors machines, SIFT based approaches and the HMAX biologically inspired approach. We show that concatenating our model which encodes color information with the HMAX model which encodes grayscale shape information results in the highest classification accuracy. Sepehr Jalali, Paul J. Seekings, Cheston Tan, Hazel Z. W. Tan, Joo-Hwee Lim, Elizabeth A. Taylor |
IJCNN | 5 |
| 2013 | Top-Down Regularization of Deep Belief NetworksabstractDesigning a principled and effective algorithm for learning deep architectures is a challenging problem. The current approach involves two training phases: a fully unsupervised learning followed by a strongly discriminative optimization. We suggest a deep learning strategy that bridges the gap between the two phases, resulting in a three-phase learning procedure. We propose to implement the scheme using a method to regularize deep belief networks with top-down information. The network is constructed from building blocks of restricted Boltzmann machines learned by combining bottom-up and top-down sampled signals. A global optimization procedure that merges samples from a forward bottom-up pass and a top-down pass is used. Experiments on the MNIST dataset show improvements over the existing algorithms for deep belief networks. Object recognition results on the Caltech-101 dataset also yield competitive results. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
NIPS | 4 |
| 2012 | Visual Attention is Attracted by Text Features Even in Scenes without Text
Hsueh-Cheng Wang, Shijian Lu, Joo-Hwee Lim, Marc Pomplun |
CogSci | 3 |
| 2012 | Unsupervised and Supervised Visual Codes with Restricted Boltzmann Machines
Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
ECCV (5) | 4 |
| 2012 | Saliency Modeling from Image Histograms
Shijian Lu, Joo-Hwee Lim |
ECCV (7) | 2 |
| 2012 | Segmentation of neural stem cells/neurospheres in high content brightfield microscopy images using localized level setsabstractNeural stem cells and neural progenitors are early nervous system cells that form neurospheres when propagated in vitro. We study changes in growth using brightfield images to understand the effects of drugs. The image quality is generally poor, imposing challenges for automatic analysis. Level-set segmentation methods are able to handle topology changes but require close initializations for accurate and efficient results. Global level-set methods using single image-wide optimization objective functions are difficult to cope with large illumination and shading changes. We propose to adopt Hough transform to initialize localized level-sets for cell segmentation. Experimental results on 480 images with 738 neurospheres show that our proposed method performs best over existing level-set methods without appropriate initial contours. Wei Xiong 0001, Shue-Ching Chia, Joo-Hwee Lim, Hwee Kuan Lee, Shvetha Sankaran, Sohail Ahmed |
ICIP | 3 |
| 2012 | Segmentation of neural stem cells/neurospheres in unevenly illuminated brightfield images with shading reduction
Wei Xiong 0001, Shue-Ching Chia, Joo-Hwee Lim, Hwee Kuan Lee, Shvetha Sankaran, Sohail Ahmed |
ICPR | 3 |
| 2012 | Clustering and use of spatial and frequency information in a biologically inspired approach to image classificationabstractIn this paper, we explore the use of spatial and frequency information of features in the biologically inspired model of HMAX. We discuss and refine previous models which use a similar framework and build specialized features which are better tuned to image structures by using unsupervised methods of clustering and picking the most frequent features using the statistics of the occurrence of the features in different spatial zones. Our classification results on the Caltech 101 dataset show significant improvements of up to 6% compared to previous improvements of the biologically inspired model of HMAX. Sepehr Jalali, Joo-Hwee Lim, Jo Yew Tham, Sim Heng Ong |
IJCNN | 2 |
| 2012 | Neurosphere fate prediction: An analysis-synthesis approach for feature extractionabstractThe study of stem cells is one of the current most important biomedical research field. Understanding their development could allow multiple applications in regenerative medicine. For this purpose, we need automated methods for the segmentation and the modeling of neural stem cell development process into a neurosphere colony from phase contrast microscopy. We use such methods to extract relevant structural and textural features like cell division dynamism and cell behavior patterns for biological interpretation. The combination of phase contrast imaging, high fragility and complex evolution of neural stem cells pose many challenges in image processing and image analysis. This study introduces an on-line analysis method for the modeling of neurosphere evolution during the first three days of their development. From the corresponding time-lapse sequences, we extract information from the neurosphere using a combination of fast level set and curve detection for segmenting the cells. Then, based on prior biological knowledge, we generate possible and optimal 3-dimensional configuration using registration and evolutionary optimisation algorithm. Stephane Ulysse Rigaud, Nicolas Loménie, Shvetha Sankaran, Sohail Ahmed, Joo-Hwee Lim, Daniel Racoceanu |
IJCNN | 5 |
| 2012 | Topic Based Query Suggestions for Video Search
Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
MMM | 3 |
| 2012 | Visual graph modeling for scene recognition and mobile robot localization
Trong-Ton Pham, Philippe Mulhem, Loïc Maisonnasse, Éric Gaussier, Joo-Hwee Lim |
Multim. Tools Appl. | 5 |
| 2012 | A non-parametric visual-sense model of images - extending the cluster hypothesis beyond text
Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
Multim. Tools Appl. | 3 |
| 2011 | Extended Visual Memory for Computer-Aided Vision
Joo-Hwee Lim |
CogSci | 1 |
| 2011 | Computer-aided cataract detection using enhanced texture features on retro-illumination lens imagesabstractCataract is a leading cause of blindness worldwide. Computer-aided cataract detection is two-fold significant. Firstly, it will be helpful in mass screening. Secondly, it can be used as the preprocessing step for computer-aided grading. In this paper, the enhanced texture feature is proposed based on the graders' expertise of cataract and the characteristics of the retro-illumination lens images. The statistics of the enhanced texture feature is used to train the linear discriminant analysis to detect the cataract. The accuracy of 84.8% is achieved on a clinical database that contains 4545 pairs of images. It demonstrates that the proposed method is promising for mass screening and as the preprocessing step for computer-aided grading. Xinting Gao, Huiqi Li, Joo-Hwee Lim, Tien Yin Wong |
ICIP | 3 |
| 2011 | Learning invariant color features with sparse topographic restricted Boltzmann machinesabstractOur objective is to learn invariant color features directly from data via unsupervised learning. In this paper, we introduce a method to regularize restricted Boltzmann machines during training to obtain features that are sparse and topographically organized. Upon analysis, the features learned are Gabor-like and demonstrate a coding of orientation, spatial position, frequency and color that vary smoothly with the topography of the feature map. There is also differentiation between monochrome and color filters, with some exhibiting color-opponent properties. We also found that the learned representation is more invariant to affine image transformations and changes in illumination color. Hanlin Goh, Lukasz Kusmierz, Joo-Hwee Lim, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2011 | Automated nuclei clump decomposition for image analysis in neuronal cell fluorescent microscopyabstractAutomated clump decomposition is crucial for the analysis of neuronal cell fluorescent microscopic images due to inevitable nuclei and cell clumping. Existing techniques are not satisfactory in terms of accuracy and efficiency. In the current work, we propose a new method to automatically decompose nuclei clumps, using multiple visual cues from seeds, skeletons, and object boundaries and nuclei shape models to deliver robust and efficient analysis. The use of shape models enhances its robustness in handling complicated clumps against those bottom-up approaches without prior knowledge. We adopt a strategy of model verification in allowable local shape changes to improve the computational efficiency over template matching. Validation experiments in the analysis of 50 images containing over 2000 nuclei have demonstrated accuracy improvements over existing techniques. Wei Xiong 0001, Shue-Ching Chia, Joo-Hwee Lim |
ICIP | 3 |
| 2011 | Advertisement Image Recognition for a Location-Based Reminder System
Aiyuan Guo, Joo-Hwee Lim |
MMM (2) | 4 |
| 2011 | A Computer Assisted Method for Nuclear Cataract Grading From Slit-Lamp Images Using RankingabstractIn clinical diagnosis, a grade indicating the severity of nuclear cataract is often manually assigned by a trained ophthalmologist to a patient after comparing the lens' opacity severity in his/her slit-lamp images with a set of standard photos. This grading scheme is often subjective and time-consuming. In this paper, a novel computer-aided diagnosis method via ranking is proposed to facilitate nuclear cataract grading following conventional clinical decision-making process. The grade of nuclear cataract in a slit-lamp image is predicted using its neighboring labeled images in a ranked image list, which is achieved using a learned ranking function. This ranking function is learned via direct optimization on a newly proposed approximation to a ranking evaluation measure. Our proposed method has been evaluated by a large dataset composed of 1000 different cases, which are collected from an ongoing clinical population-based study. Both experimental results and comparison with several existing methods demonstrate the benefit of grading via ranking by our proposed method. Wei Huang 0013, Kap Luk Chan, Huiqi Li, Joo-Hwee Lim, Jiang Liu 0001, Tien Yin Wong |
IEEE Trans. Medical Imaging | 4 |
| 2010 | Automatic optic disc detection through background estimationabstractThis paper presents an automatic optic disc (OD) detection technique. Given a retinal image, the proposed method first estimates a retinal background surface through an iterative Savitzky-Golay smoothing procedure. The OD is then detected through the global thresholding of the difference between the retinal image and the estimated background surface. Finally, an OD boundary is determined after a pair of morphological post-processing operations. The proposed technique has been tested over three public datasets that are composed of 130, 89, and 40 retinal images, respectively. Experiments show that an average OD detection accuracy of 96.91% is attained. In addition, 84.37% OD pixels are correctly located compared with the manually labeled ones. Shijian Lu, Joo-Hwee Lim |
ICIP | 2 |
| 2010 | Automatic macula detection from retinal images by a line operatorabstractThis paper presents an automatic macula detection technique that makes use of the circular brightness profile of the macula: the macula is usually darker than the surrounding pixels whose intensities increase gradually with their distances from the macula center. A line operator is designed to capture the macula circular brightness profile, which evaluates the image brightness variation along multiple line segments of specific orientations that pass through each retinal image pixel. The orientation of the line segment with the minimum/ maximum variation has specific patterns that indicate the position of the macula efficiently. The proposed technique has been tested over DRIVE project's dataset and the STARE project's dataset. Experiments show that the accuracies reach up to 100% and 95.45%, respectively, based on 35 and 44 retinal images having discernible macula within the two public datasets. Shijian Lu, Joo-Hwee Lim |
ICIP | 2 |
| 2010 | Automatic cell classification and population estimation in blastocystis autophagy imagesabstractBlastocystis is a unicellular but polymorphic protozoan parasite causing digestive diseases in humans. Autophagy, a self-degradation process, is only recently found in Blastocystis. Identifying and enumerating autophagic Blastocystis cells using fluorescent microscopy are important in biology. Doing this manually is laborious and error-prone. This paper proposes image analysis techniques to automate the process. The difficulties are poor image quality and large variations in illumination and cell morphology. We divide the cells into several sub-classes of different morphology. Support vector machines are used to learn domain knowledge and classify the cells. Validation experiments on separate data sets show reliable performance for manually segmented cells with sensitivity 82.2% and specificity 86.7%. For automatically segmented cells, the sensitivity is the same. However, the specificity drops down to 68.4%. To our knowledge, this is the first attempt in automatic processing these images. Wei Xiong 0001, Joo-Hwee Lim, Sim Heng Ong, Jiang Liu 0001, Yin Jing, Kevin S. W. Tan |
ICIP | 2 |
| 2010 | Learning cell geometry models for cell image simulation: An unbiased approachabstractComputer generation of cell images can provide annotated data to simulate various imaging conditions with controllable parameters. Synthesized images based on simple models cannot reflect the complicated parameter constraints in simulating real objects in terms of their deformation with appropriate probabilities. Learning-based techniques can provide insight to these properties and impose constraints on deformation selections. In this work, we discuss the simulation of gray level images of healthy red blood cell populations. Different from existing techniques, we learn the unbiased average shape and deformation models of the cells. Both models are used to guide the selection of possible deformations. We also learn cell color models to govern the texture generation of simulated cells. We apply this technique to simulate cell populations and validate the results using cell segmentation and counting algorithms. The proposed learning and simulation technique is generic and can be applied to other types of cells as well. Wei Xiong 0001, Sim Heng Ong, Joo-Hwee Lim, Lijun Jiang |
ICIP | 4 |
| 2010 | Faceted topic retrieval of news video using joint topic modeling of visual features and speech transcriptsabstractBecause of the inherent ambiguity in user queries, an important task of modern retrieval systems is faceted topic retrieval (FTR), which relates to the goal of returning diverse or novel information elucidating the wide range of topics or facets of the query need. We introduce a generative model for hypothesizing facets in the (news) video domain by combining the complementary information in the visual keyframes and the speech transcripts. We evaluate the efficacy of our multimodal model on the standard TRECVID-2005 video corpus annotated with facets. We find that: (1) the joint modeling of the visual and text (speech transcripts) information can achieve significant F-score improvement over a text-alone system; (2) our model compares favorably with standard diverse ranking algorithms such as the MMR. Our FTR model has been implemented on a news search prototype that is undergoing commercial trial. Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
ICME | 3 |
| 2010 | Dictionary of Features in a Biologically Inspired Approach to Image Classification
Sepehr Jalali, Joo-Hwee Lim, Sim Heng Ong, Jo Yew Tham |
ICONIP (2) | 2 |
| 2010 | A Recursive and Model-Constrained Region Splitting Algorithm for Cell Clump DecompositionabstractDecomposition of cells in clumps is a difficult segmentation task requiring region splitting techniques. Techniques that do not employ prior shape constraints usually fail to achieve accurate segmentation. Those using shape constraints are unable to cope with large clumps and occlusions. In this work, we propose a model-constrained region splitting algorithm for cell clump decomposition. We build the cell model using joint probability distribution of invariant shape features. The shape model, the contour smoothness and the gradient information along the cut are used to optimize the splitting in a recursive manner. The short cut rule is also adopted as a strategy to speed up the process. The algorithm performs well in validation experiments using 60 images with 4516 cells and 520 clumps. Wei Xiong 0001, Sim Heng Ong, Joo-Hwee Lim |
ICPR | 3 |
| 2010 | Epitomized Summarization of Wireless Capsule Endoscopic Videos for Efficient Visualization
Xinqi Chu, Chee Khun Poh, Liyuan Li, Kap Luk Chan, Shuicheng Yan, Weijia Shen, That Mon Htwe, Jiang Liu 0001, Joo-Hwee Lim, Eng Hui Ong |
MICCAI (2) | 9 |
| 2009 | A Latent Model for Visual Disambiguation of Keyword-based Image SearchabstractThe problem of polysemy in keyword-based image search arises mainly from the inherent ambiguity in user queries. We propose a latent model based approach that resolves user search ambiguity by allowing sense specific diversity in search results. Given a query keyword and the images retrieved by issuing the query to an image search engine, we first learn a latent visual sense model of these polysemous images. Next, we use Wikipedia to disambiguate the word sense of the original query, and issue these Wiki-senses as new queries to retrieve sense specific images. A sense-specific image classifier is then learnt by combining information from the latent visual sense model, and used to cluster and re-rank the polysemous images from the original query keyword into its specific senses. Results on a ground truth of 17K image set returned by 10 keyword searches and their 62 word senses provides empirical indications that our method can improve upon existing keyword based search engines. Our method learns the visual word sense models in a totally unsupervised manner, effectively filters out irrelevant images, and is able to mine the long tail of image search. Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia, Sujoy Roy |
BMVC | 3 |
| 2009 | Selecting representative and distinctive descriptors for efficient landmark recognitionabstractTo have a robust and informative image content representation for image categorization, we often need to extract as many as possible visual features at various locations, scales and orientations. Thus it is not surprised that an image has a few hundreds or even thousands of visual descriptors. This raises huge cost of computation and memory. To eliminate the problem, we can only select the most representative and distinctive descriptors and discard the other non-informative features when training the image category models. This paper will present a Markov chain based algorithm to learn a measure of the descriptor importance in order to weigh the degree of representativeness and distinctiveness. From the measures the descriptor selection algorithm is derived. The presented approach starts from constructing a graph with each node being a descriptor to characterize the pair-wise descriptor similarity and then the PageRank algorithm is exploited to estimate the stationary distribution of the graph whose values are the indicator of the descriptor importance. We evaluate the proposed approach on the STOIC-101 landmark dataset. Our experiments demonstrate the Markov chain based descriptor selection can select the most informative descriptors to distinguish the landmarks. Even with the large reduction of the size of descriptors, the classification accuracy is still competitive or overcomes compared with the system without any descriptor selection. Joo-Hwee Lim |
ICIP | 2 |
| 2009 | Photometric correction of retinal images by polynomial interpolationabstractThis paper presents a photometric restoration technique that automatically corrects shading within retinal images taken with a fundus camera. The proposed technique is based on the observation that the background of retinal images usually shows flat reflectance variations due to its high similarity in color and texture. It estimates shading through an iterative polynomial interpolation procedure that first estimates a shading image through a horizontal interpolation process and then improves the shading estimation by a vertical interpolation process. Once the shading image is estimated, a reflectance image can accordingly be determined based on the luminance of the retina image under study. Experiments on 161 retinal images of different qualities show promising results. Jiang Liu 0001, Shijian Lu, Joo-Hwee Lim, Zhuo Zhang 0001, Ngan Meng Tan, Damon Wing Kee Wong, Huiqi Li, Tien Yin Wong |
ICIP | 3 |
| 2009 | Showroom introduction using mobile phone based on scene image recognitionabstractIn this paper, image recognition is used to link a picture to relevant information. As mobile camera phone becoming more and more popular, applications for these phones will bring value to the phone and convenience to the users. We develop a system to perform audio introduction for our technology showroom using images taken by the built-in camera of mobile phone. Visitors to the showroom can take a picture of a poster, a physical prototype or any materials using their mobile phone. An audio introduction in his favorite language will be delivered to the visitor. A few schemes are used to improve the accuracy and system reliability, which include discriminative feature selection based on feature distance; repeating queries if result is not correct; user's verification on delivered content; employment of multiple classifiers. Experimental results show that the system performance is promising. Joo-Hwee Lim, Kart-Leong Lim, Yilun You |
ICME | 2 |
| 2009 | A Bayesian approach integrating regional and global features for image semantic learningabstractIn content-based image retrieval, the ldquosemantic gaprdquo between visual image features and user semantics makes it hard to predict abstract image categories from low-level features. We present a hybrid system that integrates global features (G-features) and region features (R-features) for predicting image semantics. As an intermediary between image features and categories, we introduce the notion of mid-level concepts, which enables us to predict an image's category in three steps. First, a G-prediction system uses G-features to predict the probability of each category for an image. Simultaneously, a R-prediction system analyzes R-features to identify the probabilities of mid-level concepts in that image. Finally, our hybrid H-prediction system based on a Bayesian network reconciles the predictions from both R-prediction and G-prediction to produce the final classifications. Results of experimental validations show that this hybrid system outperforms both G-prediction and R-prediction significantly. Luong-Dong Nguyen, Ghim-Eng Yap, Ying Liu 0026, Ah-Hwee Tan, Liang-Tien Chia, Joo-Hwee Lim |
ICME | 6 |
| 2009 | A Computer-Aided Diagnosis System of Nuclear Cataract via Ranking
Wei Huang 0013, Huiqi Li, Kap Luk Chan, Joo-Hwee Lim, Jiang Liu 0001, Tien Yin Wong |
MICCAI (1) | 4 |
| 2009 | Fuzzy Associative Conjuncted Maps NetworkabstractThe fuzzy associative conjuncted maps (FASCOM) is a fuzzy neural network that associates data of nonlinearly related inputs and outputs. In the network, each input or output dimension is represented by a feature map that is partitioned into fuzzy or crisp sets. These fuzzy sets are then conjuncted to form antecedents and consequences, which are subsequently associated to form if-then rules. The associative memory is encoded through an offline batch mode learning process consisting of three consecutive phases. The initial unsupervised membership function initialization phase takes inspiration from the organization of sensory maps in our brains by allocating membership functions based on uniform information density. Next, supervised Hebbian learning encodes synaptic weights between input and output nodes. Finally, a supervised error reduction phase fine-tunes the network, which allows for the discovery of the varying levels of influence of each input dimension across an output feature space in the encoded memory. In the series of experiments, we show that each phase in the learning process contributes significantly to the final accuracy of prediction. Further experiments using both toy problems and real-world data demonstrate significant superiority in terms of accuracy of nonlinear estimation when benchmarked against other prominent architectures and exhibit the network's suitability to perform analysis and prediction on real-world applications, such as traffic density prediction as shown in this paper. Hanlin Goh, Joo-Hwee Lim, Hiok Chai Quek |
IEEE Trans. Neural Networks | 2 |
| 2009 | Mobile phone-based mixed reality: the Snap2Play game
Tat-Jun Chin, Yilun You, Céline Coutrix, Joo-Hwee Lim, Jean-Pierre Chevallet, Laurence Nigay |
Vis. Comput. | 4 |
| 2008 | Using densely recorded scenes for place recognitionabstractWe investigate the task of efficiently modeling a scene to build a robust place recognition system. We propose an approach which involves densely capturing a place with video recordings to greedily cover as many viewpoints of the place as possible. Our contribution is a framework to (1) effectively exploit the temporal continuity intrinsic in the video sequences to reduce the amount of data to process without losing the unique visual information which describes a place, and (2) train discriminative classifiers with the reduced data for place recognition. We show that our method is more efficient and effective than straightforwardly applying scene or object category recognition methods on the video frames. Tat-Jun Chin, Hanlin Goh, Joo-Hwee Lim |
ICASSP | 3 |
| 2008 | Similarity Learning for Nearest Neighbor ClassificationabstractIn this paper, we propose an algorithm for learning a general class of similarity measures for kNN classification. This class encompasses, among others, the standard cosine measure, as well as the Dice and Jaccard coefficients. The algorithm we propose is an extension of the voted perceptron algorithm and allows one to learn different types of similarity functions (either based on diagonal, symmetric or asymmetric similarity matrices). The results we obtained show that learning similarity measures yields significant improvements on several collections, for two prediction rules: the standard kNN rule, which was our primary goal, and a symmetric version of it. Ali Mustafa Qamar, Éric Gaussier, Jean-Pierre Chevallet, Joo-Hwee Lim |
ICDM | 4 |
| 2008 | Automatic opacity detection in retro-illumination images for cortical cataract diagnosisabstractComputer aided analysis of medical images, a unique type of non-text media, can facilitate clinical diagnosis. As an example, an automatic opacity detection approach is proposed in this paper to grade cortical cataract more objectively. The automatic pupil detection is performed by detecting the strongest edges on the convex hull and ellipse fitting using nonlinear least square method. The cortical opacity is detected by radial edge detection and post-processing. The automatic grades are assigned following Wisconsin cataract grading protocol. The accuracy of pupil detection is 98.2%. The mean error of opacity area detection is 7 percent compared with the result of human grader. And 86.3% accurate grades of cortical cataract are achieved. This is the first time that the spoke-like feature is utilized in the automatic detection of cortical cataract to separate from other opacity types. The encouraging results show that it is probable to apply the proposed approach to clinical diagnosis later. Huiqi Li, Liling Ko, Joo-Hwee Lim, Jiang Liu 0001, Damon Wing Kee Wong, Tien Yin Wong, Ying Sun 0001 |
ICME | 3 |
| 2008 | Cascaded classification with optimal candidate selection for effective place recognitionabstractA two-stage cascaded classification approach with an optimal candidate selection scheme is proposed to recognize places using images taken by camera phones. An optimal acceptance threshold is chosen to maximize the probability of accepting more positives and rejecting more negatives at the first stage so that an optimal number of candidates are selected. The first classifier is trained using simple color and texture features. The second classifier is trained by Scale Invariant Feature Transform (SIFT). For a query image, a number of matching candidates are selected using k nearest neighbor at the first stage and passed on to the second stage for a refining classification to select the best matching result. The searching range is narrowed down dynamically at the second stage depending on the output of the first stage. Experimental results show that this method is promising by improving the recognition accuracy and reducing the computation time. Joo-Hwee Lim, Hanlin Goh |
ICME | 2 |
| 2008 | Automatic working area classification in peripheral blood smears using spatial distribution features across scalesabstractAutomatic classification of working areas in peripheral blood smears can provide objective and reproducible quality control for the evaluation of smears and smear maker devices. However, it has drawn little research attention. In this paper we study this topic using image analysis and statistical pattern recognition methods. We employ generic features without requiring the extraction of individual cells. Two new spatial distribution features across scales are defined and utilized to classify working areas. We demonstrate that the only feature and method proposed in a similar work by others is insufficient to characterize the goodness of working areas, particularly the cell distribution. However, by utilizing it together with the features developed in this paper, we can achieve much better results. Our method has been tested on about 150 labeled images acquired from three malaria-infected Giemsa-stained blood smears using an oil immersion 100x objective lens. Wei Xiong 0001, Sim Heng Ong, Joo-Hwee Lim, Nn Tung, Jiang Liu 0001, Daniel Racoceanu, Kevin S. W. Tan, Alvin G. L. Chong, Kelvin Weng Chiong Foong |
ICPR | 3 |
| 2008 | Learning associations of conjuncted fuzzy sets for data predictionabstractFuzzy associative conjuncted maps (FASCOM) is a fuzzy neural network that represents information by conjuncting fuzzy sets and associates them through a combination of unsupervised and supervised learning. The network first quantizes input and output feature maps using fuzzy sets. They are subsequently conjuncted to form antecedents and consequences, and associated to form fuzzy if-then rules. These associations are learnt through a learning process consisting of three consecutive phases. First, an unsupervised phase initializes based on information density the fuzzy membership functions that partition each feature map. Next, a supervised Hebbian learning phase encodes synaptic weights of the input-output associations. Finally, a supervised error reduction phase fine-tunes the fine-tunes the network and discovers the varying influence of an input dimension across output feature space. FASCOM was benchmarked against other prominent architectures using data taken from three nonlinear data estimation tasks and a real-world road traffic density prediction problem. The promising results compiled show significant improvements over the state-of-the-art for all four data prediction tasks. Hanlin Goh, Joo-Hwee Lim, Hiok Chai Quek |
IJCNN | 2 |
| 2008 | Deploying and evaluating a mixed reality mobile treasure hunt: Snap2PlayabstractWith the current trend, we can anticipate that future mobile phones will have ever-increasing computational power and be able to embed several captors/effectors including cameras, GPS, orientation sensors, tactile surfaces and vibro-tactile display. Such powerful mobile platforms enable us to deploy mixed reality systems. Many studies on mobile mixed reality focus on games. In this paper, we describe the deployment and a user study of a mixed reality location-based mobile treasure hunt, Snap2Play[1], using technologies such as place recognition, accelerometers and GPS tracking for enhancing the interaction with the game and therefore the game playability. The game that we deployed and tested is running on an off-the-shelf camera phone. Yilun You, Tat-Jun Chin, Joo-Hwee Lim, Jean-Pierre Chevallet, Céline Coutrix, Laurence Nigay |
Mobile HCI | 3 |
| 2008 | Snap2Play: A Mixed-Reality Game Based on Scene Identification
Tat-Jun Chin, Yilun You, Céline Coutrix, Joo-Hwee Lim, Jean-Pierre Chevallet, Laurence Nigay |
MMM | 4 |
| 2008 | Bi-modal Conceptual Indexing for Medical Image Retrieval
Joo-Hwee Lim, Jean-Pierre Chevallet, Diem Thi Hoang Le, Hanlin Goh |
MMM | 1 |
| 2008 | Rich representation and ranking for photographic image retrieval in ImageCLEF 2007abstractThe task of ad hoc photographic image retrieval in ImageCLEF 2007 international benchmark is to retrieve relevant images in the database to the user query formulated as keywords and image examples. This paper presents rich representation and indexing technologies exploited in our system that participated in ImageCLEF 2007. It uses diverse visual content representation, text representation, pseudo-relevance feedback and fusion, which make our system, with mean average precision 0.2833, in the 4th place among 457 automatic runs submitted from 20 participants to photographic ImageCLEF 2007 and in the 2nd place in terms of participants. Our systematic analysis in the paper demonstrates that 1) combing diverse low-level visual features and ranking technologies significantly improves the content-based image retrieval (CBIR) system; 2) cross-modality pseudo-relevance feedback improves the system performance; and 3) fusion of CBIR and TBIR outperforms individual modality based system. Jean-Pierre Chevallet, Joo-Hwee Lim |
MMSP | 3 |
| 2007 | Domain knowledge conceptual inter-media indexing: application to multilingual multimedia medical reportsabstractConceptual Indexing is a way to produce only one index for many multilingual documents. Inter-Media conceptual indexing promotes the use of common concepts between two media in order to use a single index for several media. In this paper we explore such an advance indexing point of view. We show the benefit of an automatic conceptual indexing for texts and its extension for text and image documents. Tests are conducted on the multilingual image and text medical document corpus of the CLEF initiative, where we obtain best results on text in 2005 and 2006, and show promising results on images, and best results for the combination of image and text. Jean-Pierre Chevallet, Joo-Hwee Lim, Diem Thi Hoang Le |
CIKM | 2 |
| 2007 | Latent semantic fusion model for image retrieval and annotationabstractThis paper studies the effect of Latent Semantic Analysis (LSA) on two different tasks: multimedia document retrieval (MDR) and automatic image annotation (AIA). The contributions of this paper are twofold. First, to the best of our knowledge, this work is the first study of the influence of LSA on the retrieval of a significant number of multimedia documents (i.e. collection of 20000 tourist images). Second, it shows how different image representations (region-based and keypoint-based) can be combined by LSA to improve automatic image annotation. The document collections used for these experiments are the Corel photo collection and ImageCLEF 2006 collection. Trong-Ton Pham, Nicolas Maillot, Joo-Hwee Lim, Jean-Pierre Chevallet |
CIKM | 3 |
| 2007 | Outlier Detection from Pooled Data for Image Retrieval System EvaluationabstractWidely used in the evaluation of retrieval systems, the pooling method collects top ranked images from submitted retrieval systems resulting in possibly a very large pool of images. Inevitably, the pool may contain outliers. Human experts then manually annotate the relevance of them to create a ground truth for evaluation. Studies show that this annotation is time-consuming, tedious and inconsistent. To reduce human workload, this paper introduces an automatic method to detect outliers. Different from traditional detection methods using unsupervised techniques only, we utilize both supervised and unsupervised techniques sequentially as both positive and negative examples are (partially) available in this context. Specifically, support vector machines (SVMs) and fuzzy c-means clustering are used to predict data relevance and "outlierness". Performance improvements using our method after outlier removal have been validated on the medical image retrieval task in ImageCLEF 2004. Wei Xiong 0001, Sim Heng Ong, Joo-Hwee Lim, Qi Tian 0002, Changsheng Xu, Kelvin Weng Chiong Foong |
ICASSP (1) | 3 |
| 2007 | Propagating Image-Level Part Statistics to Enhance Object DetectionabstractThe bag-of-words approach has become increasingly attractive in the fields of object category recognition and scene classification, witnessed by some successful applications [5, 7, 11]. Its basic idea is to quantize an image using visual terms and exploit the image-level statistics for classification. However, the previous work still lacks the capability of modeling the spatial dependency and the correspondence between patches and object parts. Moreover, quantization always deteriorates the descriptive power of the patch feature. This paper proposes the hidden maximum entropy (HME) approach for modeling the object category. Each object is modeled by the parts, each having a Gaussian distribution. The spatial dependency and image-level statistics of parts are modeled through the maximum entropy approach. The model is learned by an EM-IIS (expectation maximum embedded with improved iterative scaling) algorithm. Our experiments on the Caltech 101 dataset show that the relative reduction of equal error rate of 23.5 % and relative improvement of AUC (area under ROC) of 22.0 % are obtained when comparing the HME based system with the ME based baseline system. Joo-Hwee Lim, Qibin Sun |
ICIP (6) | 2 |
| 2007 | Hidden Maximum Entropy Approach for Visual Concept ModelingabstractRecently, the bag-of-words approach has been successfully applied to automatic image annotation, object recognition, etc. The method needs to first quantize an image using the visual terms and then extract the image-level statistics for classification. Although successful applications have been reported, it lacks the capability to model the spatial dependency and the correspondence between the patches and visual parts. Moreover, quantization deteriorates the descriptive power of patch feature. This paper proposes the hidden maximum entropy (HME) approach for modeling visual concepts. Each concept is composed of a set of visual parts, each part having a Gaussian distribution. The spatial dependency and image-level statistics of parts are modeled through the maximum entropy. The model is learned using the developed EM-IIS algorithm. We report the preliminary results on the 260 concepts in the Corel dataset and compared with the maximum entropy (ME) approach. Our experiments on concept detection show that (1) a relative increment of 10.3% is observed when comparing the average AUC value of HME approach with that of the ME approach and (2) the HME approach reduces the average equal error rate from 0.412 for the ME approach to 0.354. Joo-Hwee Lim, Qibin Sun |
ICME | 2 |
| 2007 | Knowledge-Assisted Medical Image RetrievalabstractIn this paper, we present a knowledge-assisted approach to index and retrieve large volume of medical images. Both images and associated texts are indexed using medical concepts from the Unified Medical Language System (UMLS) meta-thesaurus. We propose a structured learning framework for modular acquisition of medical semantics from images with complementary global and local image indexing schemes. Two fusion approaches are also developed to improve text retrieval using the UMLS-based image indexing: a simple post-query fusion and a visual modality filtering to remove visually aberrant images according to the query modality concepts. On the ImageCLEFmed 2005 database, our framework outperformed our previous result which ranked top in the ImageCLEFmed 2005 Medical Image Retrieval task benchmark. Joo-Hwee Lim, Caroline Lacoste, Jean-Pierre Chevallet, Diem Thi Hoang Le |
ICME | 1 |
| 2007 | Scene Recognition with Camera Phones for Tourist Information AccessabstractCamera phones present new opportunities and challenges for mobile information association and retrieval. The visual input in the real environment is a new and rich interaction modality between a mobile user and vast information base connected to a user's device via rapidly advancing communication infrastructure. We have developed a system for tourist information access to provide scene description based on an image taken of the scene. In this paper, we describe the working system, the STOIC 101 database, and a new pattern discovery algorithm to learn image patches that are recurrent within a scene class and discriminative across others. We report preliminary scene recognition results on 90 scenes, trained on 5 images per scene, with an accuracy of 92% and 88% on a test set of 110 images, with and without location priming. Joo-Hwee Lim, Yilun You, Jean-Pierre Chevallet |
ICME | 1 |
| 2007 | An integrated statistical model for multimedia evidence combinationabstractGiven the rich content-based features of multimedia (e.g., visual, text, or audio) and the development of various approaches to automatic detectors (e.g., SVM, Adaboost, HMM or GMM, etc), can we find an efficient approach to combine these evidences? In the paper, we address this issue by proposing an Integrated Statistical Model (ISM) to combine diverse evidences extracted from the domain knowledge of detectors, the intrinsic structure of modality distribution and inter-concept associations. The ISM provides a unified framework for evidence fusion, having the following unique advantages: 1) the intrinsic modes in the modality distribution are discovered and modeled by a generative model; 2) each mode is a partial description of structure of the modality and the mode configuration, i.e. a set of modes, and is a new representation of the document content; 3) mode discrimination is automatically learned; 4) prior knowledge such as detector correlations and inter-concept relations can be explicitly described and integrated. More importantly, an efficient pseudo-EM algorithm is realized for training the statistical model. The learning algorithm relaxes the computational cost due to the normalized factor and latent variables in the graphical model. We evaluate system performance of our multimedia semantic concept detection with the TRECVID 2005 development dataset, in terms of efficiency and capacity. Our experimental results demonstrate that the ISM fusion outperforms the SVM based discriminative fusion method. Joo-Hwee Lim, Qibin Sun |
ACM Multimedia | 2 |
| 2007 | An image-based outdoor place recognition and information retrieval systemabstractIn image-based place recognition, an image is used to deduce the location of the viewer during image acquisition. The identified place can subsequently be used to provide additional information to the user. This demonstration paper briefly describes a system that performs image-based place recognition on outdoor images, made possible by viewer-centric data sampling and local feature-based scene identification. It also explains our proposed demonstration that displays the recognized place on a map and provides information of amenities in its vicinity. Hanlin Goh, Joo-Hwee Lim |
ACM Multimedia | 3 |
| 2007 | Outdoor place recognition using compact local descriptors and multiple queries with user verificationabstractIn this paper, we propose a novel method to model outdoor places with compact local descriptors extracted from images taken around geographical places. Region-based and clustering-based methods are used to reduce the number of feature vectors to represent the natural scene images. A Multiple Queries with User Verification (MQUV) scheme is proposed to improve the recognition accuracy and the system reliability. In our application, a mobile phone camera is used to take images around a place and send them back to the server to get relevant information about the place. The MQUV scheme calculates the maximum confidence level of all top 5 matching places and returns the best matching result to the user together with a typical sample image of the recognized place for the user's visual verification. User is suggested to take more images if the system is not confident enough to provide a result. The user can also make one's own decision by visually matching the returned image with the scenery of the place. Experimental results show that the number of feature vectors is significantly reduced with the compact place modeling and the recognition accuracy is improved with the MQUV scheme. Joo-Hwee Lim |
ACM Multimedia | 2 |
| 2007 | Metadata Management, Reuse, Inference and Propagation in a Collection-Oriented Metadata Framework for Digital Images
William Ku, Mohan Kankanhalli, Joo-Hwee Lim |
MMM (2) | 3 |
| 2007 | Object identification and retrieval from efficient image matching. Snap2Tell with the STOIC dataset
Jean-Pierre Chevallet, Joo-Hwee Lim, Mun-Kew Leong |
Inf. Process. Manag. | 2 |
| 2007 | Medical-Image Retrieval Based on Knowledge-Assisted Text and Image IndexingabstractVoluminous medical images are generated daily. They are critical assets for medical diagnosis, research, and teaching. To facilitate automatic indexing and retrieval of large medical-image databases, both images and associated texts are indexed using medical concepts from the Unified Medical Language System (UMLS) meta-thesaurus. We propose a structured learning framework based on support vector machines to facilitate modular design and learning of medical semantics from images. We present two complementary visual indexing approaches within this framework: a global indexing to access image modality and a local indexing to access semantic local features. Two fusion approaches are developed to improve textual retrieval using the UMLS-based image indexing. First, a simple fusion of the textual and visual retrieval approaches is proposed, improving significantly the retrieval results of both text and image retrieval. Second, a visual modality filtering is designed to remove visually aberrant images according to the query modality concept(s). Using the ImageCLEFmed database, we demonstrate the effectiveness of our framework which is superior when compared with the automatic runs evaluated in 2005 on the same medical-image retrieval task. Caroline Lacoste, Joo-Hwee Lim, Jean-Pierre Chevallet, Diem Thi Hoang Le |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Towards Automatic Mobile BloggingabstractWeblog (usually shortened as blog) has gained its popularity lately. There are about 70,000 new blogs a day and about 29,100 blog updates an hour. As an emerging blogging phenomenon, with the proliferation of camera phones, mobile bloggers can write their blogs almost instantaneously. But how much further can current mobile blogging tools enhance the experience? In this paper, we propose a Mobilog framework to automate context-relevant annotation and synthesise personalised content for mobile blogging. In particular, we describe a system implementation of the framework, Travelog, adapted for tourism applications. Finally, we discuss the challenges and future possibilities for mobile blogging Pujianto Cemerlang, Joo-Hwee Lim, Yilun You, Jean-Pierre Chevallet |
ICME | 2 |
| 2006 | A Collection-Oriented Metadata Framework for Digital ImagesabstractA digital photo can "tell a thousand words" through the use of its metadata and as it is usually part of a collection, metadata management, reuse, propagation&inference could be achieved via its association with a collection. However, there is not much work on metadata management, reuse, propagation&inference, particularly on a group basis. In this paper, we proposed a collection-oriented metadata framework which provides a basis for metadata management, reuse, propagation&inference and demonstrated the utility of such a framework William Ku, Mohan Kankanhalli, Joo-Hwee Lim |
ICME | 3 |
| 2006 | Combining Textual and Visual Ontologies to Solve Medical Multimodal QueriesabstractIn order to solve medical multimodal queries, we propose to split the queries in different dimensions using ontology. We extract both textual and visual terms depending on the ontology dimension they belong to. Based on these terms, we build different sub queries each corresponds to one query dimension. Then we use Boolean expressions on these sub queries to filter the entire document collection. The filtered document set is ranked using the techniques in vector space model. We also combine the ranked lists generated using both text and image indexes to further improve the retrieval performance. We have achieved the best overall performance for the medical image retrieval task in CLEF 2005. These experimental results show that while most queries are better handled by the text query processing as most semantic information are contained in the medical text cases, both textual and visual ontology dimensions are complementary in improving the results during media fusion Saïd Radhouani, Joo-Hwee Lim, Jean-Pierre Chevallet, Gilles Falquet |
ICME | 2 |
| 2005 | SnapToTell: A Singapore Image Test Bed for Ubiquitous Information Access from Camera
Jean-Pierre Chevallet, Joo-Hwee Lim, Ramnath Vasudha |
ECIR | 2 |
| 2005 | A structured learning framework for content-based image indexing and visual query
Joo-Hwee Lim, Jesse S. Jin |
Multim. Syst. | 1 |
| 2005 | Combining intra-image and inter-class semantics for consumer image retrieval
Joo-Hwee Lim, Jesse S. Jin |
Pattern Recognit. | 1 |
| 2004 | Semantics Discovery for Image Indexing
Joo-Hwee Lim, Jesse S. Jin |
ECCV (1) | 1 |
| 2004 | Goal detection in soccer video using audio/visual keywords
Yu-Lin Kang, Joo-Hwee Lim, Mohan Kankanhalli, Changsheng Xu, Qi Tian 0002 |
ICIP | 2 |
| 2004 | Combining local class patterns and discovered semantics for image retrievalabstractDetecting meaningful visual entities (e.g. faces, sky, foliage, buildings, etc.) based on supervised pattern classifiers has become a trend in content-based image retrieval. However, a drawback of the supervised learning approach is the need for manually labeled regions as training samples. We propose a semi-supervised framework to discover local semantic patterns and generate their samples for training with minimal human intervention. Image classifiers are first trained on local image blocks from a small number of labeled images. Then, local semantic patterns are discovered from clustering the image blocks with high classification output. Training samples are induced from cluster memberships for support vector learning to form local semantic pattern detectors. During retrieval, similarities based on local class pattern indexes and discovered pattern indexes are combined to rank images. Query-by-example experiments on 2400 unconstrained consumer photos with 16 semantic queries show that the combined matching approach outperformed the fusion of color and texture features significantly in average precision by 37%. Joo-Hwee Lim, Jesse S. Jin |
ICIP | 1 |
| 2004 | Unifying local and global content-based similarities for home photo retrievalabstractUnlike professional or domain-specific images, home photos vary significantly. They pose great challenge for content-based image retrieval. In this paper, we propose a Bayesian formulation to unify both local and global content-based similarities for image matching and demonstrate its superior retrieval performance on 2400 genuine home photos. Our proposed framework uses support vector machines to extract and combine intraimage and interclass semantics. Support vector detectors are first trained on semantically meaningful regions and used to form detection image indexes. The indexes then serve as input for support vector learning of image classifiers to generate class-relative indexes. During image retrieval, similarities based on both detection-based and class-relative indexes are combined to rank images. Query-by-example experiments on 2400 home photos with 16 semantic queries show that the combined matching approach is better than matching with single index. It also outperformed the fusion of color and texture features by 55% in average precision. Joo-Hwee Lim, Jesse S. Jin |
ICIP | 1 |
| 2004 | A generic mid-level representation for semantic video analysisabstractThe paper presents a generic, mid-level representation for efficient semantic video analysis, which adopts a frame-by-frame scheme using P-frames rather than shot-based schemes. Each P-frame is partitioned into an m/spl times/n grid (row by column), and each cell is called a 'block'. The representation can bridge the semantic gap and build an intermediate description of video features across frames and blocks. Soccer video is used to showcase the potential of the framework for real video processing. Experiments with tennis video and news video have also been conducted. Results demonstrate the excellent performance of the framework in semantic analysis and also indicate its further potential for automatic video analysis. Joo-Hwee Lim, Jesse S. Jin, Haiping Sun, Qi Tian 0002 |
ICIP | 2 |
| 2004 | Image retrieval using spatial iconsabstractWe propose an adaptive semantic image indexing framework based on the notion of semantic support regions. Semantic support regions are salient image regions that exhibit semantic meanings and that can be derived statistically to span a new indexing space. We perform the learning and indexing of semantic support regions for 2400 heterogeneous consumer photos using support vector machines. We demonstrate the uniqueness and effectiveness of query by spatial icons, supported by our semantic indexing framework, with 12 queries. Joo-Hwee Lim, Jesse S. Jin |
ICME | 1 |
| 2004 | Learning and Integrating Semantics for Image Indexing
Joo-Hwee Lim, Jesse S. Jin |
PRICAI | 1 |
| 2003 | Real-time camera field-view tracking in soccer videoabstractSoccer video content-based analysis remains a challenging problem due to the lack of structure in a soccer game. To automate game and tactic analysis, we need to detect and track important activities such as ball possession in a soccer video that is highly correlated to the camera's field-view. In this paper, we present a system that tracks the camera's field-view in a soccer video in real-time. It utilizes a host of content-based visual cues that are obtained by independent threads running in parallel. The result is visualized as an active rectangular bounding box that approximates the camera's field of view superimposed on a virtual soccer field. Experimental results show that the system can reliably track the camera field-view as the game progresses. Kong-Wah Wan, Joo-Hwee Lim, Changsheng Xu, Xinguo Yu |
ICASSP (3) | 2 |
| 2003 | Support regions and images for photo event retrievalabstractThe notion of event conveys rich semantics to consumers in their collection of photos. Our user studies confirm that consumers would like to organize and access their digital photo memories by events, complemented by other semantic axes such as people, time, and place. In this paper, we address photo event retrieval by means of semantic support regions and images. Semantic support regions are salient image regions that exhibit semantic meanings and that can be learned from sample images to span a new indexing space. Semantic support images are key photos derived statistically to model photo events. Retrieval of a photo event is performed using a winner-take-all approach to compute the relevance measure of a given photo to the event models. This novel approach for photo event retrieval is experimented on 2400 heterogeneous consumer photos with very promising results. Joo-Hwee Lim, Jesse S. Jin |
ICIP (2) | 1 |
| 2003 | Event-based home photo retrievalabstractWith rapid advances in sensor, storage, processor, and communication technologies, consumers can now afford to create, store, process, and share large digital photo collections. With more and more digital photos accumulated, consumers need effective and efficient tools to organize and access photos in a semantically meaningful way without too much manual annotation effort. From user studies, we confirm that users prefer to organize and access photos along semantic axes such as event, people, time, and place. In this paper, we propose a computational learning framework to construct event models from sample photos with event labels given by a user and to compute relevance measures of unlabeled photos to the event models. We demonstrate event-based retrieval on 2400 genuine home photos using our proposed approach. Joo-Hwee Lim, Philippe Mulhem, Qi Tian 0002 |
ICME | 1 |
| 2003 | Learning Consumer Photo Categories for Semantic Retrieval
Joo-Hwee Lim, Jesse S. Jin |
IJCAI | 1 |
| 2002 | Symbolic photograph content-based retrievalabstractPhotograph retrieval systems face the difficulty to deal with the different ways to apprehend the content of images. We consider and demonstrate here the use of multiple index representations of photographs to achieve effective retrieval. The use of multiple indexes allows integration of the complementary strengths of different indexing and retrieval models. The proposed representation supports multiple labels for regions and attributes, and handles inferences and relationships. We define links between indexing levels and the related query modes. The experiment conducted on 2400 home photographs shows the behavior of the multiple indexing levels during retrieval. Philippe Mulhem, Joo-Hwee Lim |
CIKM | 2 |
| 2002 | Semantic indexing and retrieval of home photosabstractWith rapid advances in sensor, storage, processor, and communication technologies, consumers can now afford to create, store, process, and share large digital photo collections. With more and more digital photos accumulated, consumers need effective and efficient tools to organize and access photos in a semantically meaningful way without too much manual annotation effort. From user studies, we confirm that users prefer to organize and access photos along semantic axes such as event, people, time, and place. As a matter of fact, research on content-based image retrieval in the last decade is yet to bridge the semantic gap between feature-based indexes computed automatically and human preferences on query and retrieval. In this paper, we attempt to address this semantic gap by focusing on the notion of "event" in home photos. First we propose event taxonomy for home photos. Next we propose a computational learning framework to construct event models from sample photos with event labels given by a user and to compute relevance measures of unlabeled photos to the event models, Last but not least, we demonstrate event-based retrieval on 2400 genuine home photos using our proposed approach. Joo-Hwee Lim, Jesse S. Jin |
ICARCV | 1 |
| 2002 | Image indexing and retrieval using visual keyword histogramsabstractWe propose a novel image representation called visual keyword histogram (VKH) for content-based indexing and retrieval. Visual keywords are domain-relevant visual prototypes (e.g. faces, foliage, buildings etc) with both perceptual appearance and textual semantics. Collectively, VKHs axe computed over spatial tessellation to represent the distribution of visual keywords in various parts of an image. To construct a vocabulary of visual keywords, an incremental neural network is deployed to learn visual keywords from examples. This allows us to build domain-specific visual vocabularies rapidly and incrementally. Last but not least, we propose a new visual query language called Query by Spatial Icons (QBSI) that allows a user to specify a query in terms of "what" and "where". A visual query term constrains whether a visual keyword should be present and a query formals chains these terms into a disjunctive normal form via logical operators. We show our approach on real and complex home photos with very promising results. Joo-Hwee Lim, Jesse S. Jin |
ICME (1) | 1 |
| 2001 | Fuzzy Object Patterns for Visual Indexing and SegmentationabstractIn this paper, we propose a fuzzy object pattern (FOP) for the representation of image content. The FOP is derived from view-based object recognition against a pre-defined vocabulary of visual object classes. Tessellation of FOPs over an image is further aggregated spatially to summarize the image content. This description scheme has been deployed in image indexing and retrieval on home photographs with very promising results. Furthermore, the FOP spans a new fuzzy pattern space in which incremental clustering is carried out to aggregate adjacent FOPs into larger regions. As a consequence, dominant regions can be segmented from an image. Joo-Hwee Lim |
FUZZ-IEEE | 1 |
| 2001 | Building Visual Vocabulary for Image Indexation and Query Formulation
Joo-Hwee Lim |
Pattern Anal. Appl. | 1 |
| 2001 | Learning Similarity Matching in Multimedia Content-Based RetrievalabstractMany multimedia content-based retrieval systems allow query formulation with the user setting the relative importance of features (e.g., color, texture, shape, etc.) to mimic the user's perception of similarity. However, the systems do not modify their similarity matching functions, which are defined during the system development. We present a neural network-based learning algorithm for adapting the similarity matching function toward the user's query preference based on his/her relevance feedback. The relevance feedback is given as ranking errors (misranks) between the retrieved and desired lists of multimedia objects. The algorithm is demonstrated for facial image retrieval using the NIST Mugshot Identification Database with encouraging results. Joo-Hwee Lim, Jian-Kang Wu, Sumeet Singh, Arcot Desai Narasimhalu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2000 | Explicit query formulation with visual keywords
Joo-Hwee Lim |
ACM Multimedia | 1 |
| 1996 | An Application of Hierarchical Knowledge Integration in Hand-Written Form Processing
Fon-Lin Lai, Joo-Hwee Lim, Ah-Hwee Tan, Ho-Chung Lui |
PRICAI | 2 |
| 1996 | Stochastic topology with elastic matching for off-line handwritten character recognition
Joo-Hwee Lim, Hoon Heng Teh, Ho-Chung Lui, Pei-Zhuang Wang |
Pattern Recognit. Lett. | 1 |
| 1992 | A Framework for Integrating Fault Diagnosis and Incremental Knowledge Acquisition in Connectionist Expert Systems
Joo-Hwee Lim, Ho-Chung Lui, Pei-Zhuang Wang |
AAAI | 1 |