Cheston Tan

dblp:136/9366 · also Cheston Yin Chet Tan · DBLP profile ↗
← Back
43ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0003-1248-4906ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 10 Open Challenges Steering the Future of Vision-Language-Action Models
abstract
Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly preva- lent in the embodied AI arena, following the widespread suc- cess of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing develop- ment of VLA models—multimodality, reasoning, data, eval- uation, cross-robkot action generalization, efficiency, whole- body coordination, safety, agents, and coordination with hu- mans. Furthermore, we discuss the emerging trends of us- ing spatial understanding, modeling world dynamics, post training, and data synthesis—all aiming to reach these mile- stones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability.
Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Kenneth Kwok, Ziwei Wang 0010, Cheston Tan, Jiajun Wu 0001, David Hsu
AAAI8
2026 Unveiling and Mitigating Untargeted Poisoning Attacks on Federated Knowledge Graph Embedding
Wenzheng Jiang, Ke Liang 0006, Wenke Huang 0003, Xiongtao Zhang, Guancheng Wan, Cheston Tan, Flint Xiaofeng Fan, Ji Wang 0002
WWW7
2025 Theory of Mind in Large Language Models: Assessment and Enhancement
abstract
Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence.As Large Language Models (LLMs) become increasingly integrated into daily life, understanding their ability to interpret and respond to human mental states is crucial for enabling effective interactions.In this paper, we review LLMs' ToM capabilities by analyzing both evaluation benchmarks and enhancement strategies.For evaluation, we focus on recently proposed and widely used story-based benchmarks.For enhancement, we provide an indepth analysis of recent methods aimed at improving LLMs' ToM abilities.Furthermore, we outline promising directions for future research to further advance these capabilities and better adapt LLMs to more realistic and diverse scenarios.Our survey serves as a valuable resource for researchers interested in evaluating and advancing LLMs' ToM capabilities.
Ruirui Chen 0002, Weifeng Jiang, Chengwei Qin, Cheston Tan
ACL (1)4
2025 GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
abstract
Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual grounding (3DVG) methods treat text instructions with multiple steps as a whole, without extracting useful temporal information from each step. However, the instructions in SG3D often contain pronouns such as "it", "here" and "the same" to make language expressions concise. This requires grounding methods to understand the context and retrieve relevant information from previous steps to correctly locate object sequences. Due to the lack of an effective module for collecting related historical information, state-of-the-art 3DVG methods face significant challenges in adapting to the SG3D task. To fill this gap, we propose GroundFlow -- a plug-in module for temporal reasoning on 3D point cloud sequential grounding. Firstly, we demonstrate that integrating GroundFlow improves the task accuracy of 3DVG baseline methods by a large margin (+7.5\% and +10.2\%) in the SG3D benchmark, even outperforming a 3D large language model pre-trained on various datasets. Furthermore, we selectively extract both short-term and long-term step information based on its relevance to the current instruction, enabling GroundFlow to take a comprehensive view of historical information and maintain its temporal understanding advantage as step counts increase. Overall, our work introduces temporal reasoning capabilities to existing 3DVG models and achieves state-of-the-art performance in the SG3D benchmark across five datasets.
Shuting He, Cheston Tan, Bihan Wen
ICCV3
2025 Stencil: Subject-Driven Generation with Context Guidance
abstract
Recent text-to-image diffusion models can produce impressive visuals from textual prompts, but they struggle to reproduce the same subject consistently across multiple generations or contexts. Existing fine-tuning based methods for subject-driven generation face a trade-off between quality and efficiency. Fine-tuning larger models yield higher-quality images but is computationally expensive, while fine-tuning smaller models is more efficient but compromises image quality. To this end, we present Stencil. Stencil resolves this trade-off by leveraging the superior contextual priors of large models and efficient fine-tuning of small models. Stencil uses a small model for fine-tuning while a large pre-trained model provides contextual guidance during inference, injecting rich priors into the generation process with minimal overhead. Stencil excels at generating high-fidelity, novel renditions of the subject in less than a minute, delivering state-of-the-art performance and setting a new benchmark in subject-driven generation. Supplementary materials are available at IEEE SigPort.
Gordon Chen, Cheston Tan, Ziwei Liu 0002
ICIP3
2025 FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong, Roger Wattenhofer, Wei Tsang Ooi
AAMAS2
2025 FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation
Wenzheng Jiang, Ji Wang 0002, Xiongtao Zhang, Weidong Bao 0001, Cheston Tan, Flint Xiaofeng Fan
AAMAS5
2025 Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
abstract
The rise of reinforcement learning (RL) in critical real-world applications demands a fundamental rethinking of privacy in AI systems. Traditional privacy frameworks, designed to protect isolated data points, fall short for sequential decision-making systems where sensitive information emerges from temporal patterns, behavioral strategies, and collaborative dynamics. Modern RL paradigms, such as federated RL (FedRL) and RL with human feedback (RLHF) in large language models (LLMs), exacerbate these challenges by introducing complex, interactive, and context-dependent learning environments that traditional methods do not address. In this position paper, we argue for a new privacy paradigm built on four core principles: multi-scale protection, behavioral pattern protection, collaborative privacy preservation, and context-aware adaptation. These principles expose inherent tensions between privacy, utility, and interpretability that must be navigated as RL systems become more pervasive in high-stakes domains like healthcare, autonomous vehicles, and decision support systems powered by LLMs. To tackle these challenges, we call for the development of new theoretical frameworks, practical mechanisms, and rigorous evaluation methodologies that collectively enable effective privacy protection in sequential decision-making systems.
Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer, Yew-Soon Ong
IJCNN2
2025 Inferring Past Human Actions in Homes with Abductive Reasoning
abstract
Abductive reasoning aims to make the most likely inference for a given set of incomplete observations. In this paper, we introduce “Abductive Past Action Inference”, a novel research task aimed at identifying the past actions performed by individuals within homes to reach specific states captured in a single image, using abductive inference. The research explores three key abductive inference problems: past action set prediction, past action sequence prediction, and abductive past action verification. We introduce several models tailored for abductive past action inference, including a relational graph neural network, a relational bilinear pooling model, and a relational transformer model. Notably, the newly proposed object-relational bilinear graph encoder-decoder (BiGED) model emerges as the most effective among all methods evaluated, demonstrating good proficiency in handling the intricacies of the Action Genome dataset. The contributions of this research significantly advance the ability of deep learning models to reason about current scene evidence and make highly plausible inferences about past human actions. This advancement enables a deeper understanding of events and behaviors, which can enhance decision-making and improve system capabilities across various real-world applications such as Human-Robot Interaction and Elderly Care and Health Monitoring. Code and data available at https://github.com/LUNAProject22/AAR
Clement Tan, Chai Kiat Yeo, Cheston Tan, Basura Fernando
WACV3
2024 Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
abstract
While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting biases and spurious features? Hence, to provide insights, we design QUAG (QUadrant AveraGe), a lightweight and non-parametric probe, to conduct dataset-model combined representation analysis by impairing modality fusion. We find that the models achieve high performance on many datasets without leveraging multimodal representations. To validate QUAG further, we design QUAG-attention, a less-expressive replacement of self-attention with restricted token interactions. Models with QUAG-attention achieve similar performance with significantly fewer multiplication operations without any finetuning. Our findings raise doubts about the current models’ abilities to learn highly-coupled multimodal representations. Hence, we design the CLAVI (Complements in LAnguage and VIdeo) dataset, a stress-test dataset curated by augmenting real-world videos to have high modality coupling. Consistent with the findings of QUAG, we find that most of the models achieve near-trivial performance on CLAVI. This reasserts the limitations of current models for learning highly-coupled multimodal representations, that is not evaluated by the current datasets.
Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando, Cheston Tan
ICML5
2024 Can LLMs Perform Structured Graph Reasoning Tasks?
Palaash Agrawal, Shavak Vasania, Cheston Tan
ICPR (6)3
2024 Social Learning through Interactions with Other Agents: A Survey
Dylan Hillier, Cheston Tan
IJCAI2
2024 Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
abstract
Vision-language models (VLMs) have shown impressive zero-and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visual reasoning engines. However, the benchmarks being used conflate "pure" visual reasoning with world knowledge, and also have questions that involve a limited number of reasoning steps. Thus, it remains unclear whether a VLM’s apparent visual reasoning performance is due to its world knowledge, or due to actual visual reasoning capabilities.To clarify this ambiguity, we systematically benchmark and dissect the zero-shot visual reasoning capabilities of VLMs through synthetic datasets that require minimal world knowledge, and allow for analysis over a broad range of reasoning steps. We focus on two novel aspects of zero-shot visual reasoning: i) evaluating the impact of conveying scene information as either visual embeddings or purely textual scene descriptions to the underlying large language model (LLM) of the VLM, and ii) comparing the effectiveness of chain-of-thought prompting to standard prompting for zero-shot visual reasoning.We find that the underlying LLMs, when provided textual scene descriptions, consistently perform better compared to being provided visual embeddings. In particular, 18% higher accuracy is achieved on the PTR dataset. We also ∼find that CoT prompting performs marginally better than standard prompting only for the comparatively large GPT-3.5-Turbo (175B) model, and does worse for smaller-scale models. This suggests the emergence of CoT abilities for visual reasoning in LLMs at larger scales even when world knowledge is limited. Overall, we find limitations in the abilities of VLMs and LLMs for more complex visual reasoning, and highlight the important role that LLMs can play in visual reasoning.
Aishik Nagar, Shantanu Jaiswal, Cheston Tan
IJCNN3
2024 Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
abstract
Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the immediate recognition and localization of objects and events. Here, we introduce a fully neural Iterative and Parallel Reasoning Mechanism (IPRM) that combines two distinct forms of computation -- iterative and parallel -- to better address complex VQA scenarios. Specifically, IPRM's "iterative" computation facilitates compositional step-by-step reasoning for scenarios wherein individual operations need to be computed, stored, and recalled dynamically (e.g. when computing the query “determine the color of pen to the left of the child in red t-shirt sitting at the white table”). Meanwhile, its "parallel'' computation allows for the simultaneous exploration of different reasoning paths and benefits more robust and efficient execution of operations that are mutually independent (e.g. when counting individual colors for the query: "determine the maximum occurring color amongst all t-shirts'"). We design IPRM as a lightweight and fully-differentiable neural module that can be conveniently applied to both transformer and non-transformer vision-language backbones. It notably outperforms prior task-specific methods and transformer-based attention modules across various image and video VQA benchmarks testing distinct complex reasoning capabilities such as compositional spatiotemporal reasoning (AGQA), situational reasoning (STAR), multi-hop reasoning generalization (CLEVR-Humans) and causal event linking (CLEVRER-Humans). Further, IPRM's internal computations can be visualized across reasoning steps, aiding interpretability and diagnosis of its errors.
Shantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston Tan
NeurIPS4
2023 A Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event Categories
Arijit Dasgupta, Jiafei Duan, Su-Hua Wang, Renée Baillargeon, Cheston Tan
CogSci6
2023 DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners
abstract
State-of-the-art visual grounding models can achieve high detection accuracy, but they are not designed to distinguish between all objects versus only certain objects of interest. In natural language, in order to specify a particular object or set of objects of interest, humans use determiners such as "my", "either" and "those". Determiners, as an important word class, are a type of schema in natural language about the reference or quantity of the noun. Existing grounded referencing datasets place much less emphasis on determiners, compared to other word classes such as nouns, verbs and adjectives. This makes it difficult to develop models that understand the full variety and complexity of object referencing. Thus, we have developed and released the DetermiNet dataset1, which comprises 250,000 synthetically generated images and captions based on 25 determiners. The task is to predict bounding boxes to identify objects of interest, constrained by the semantics of the given determiner. We find that current state-of-the-art visual grounding models do not perform well on the dataset, highlighting the limitations of existing models on reference and quantification tasks.
Clarence Lee, M. Ganesh Kumar, Cheston Tan
ICCV3
2022 PIP: Physical Interaction Prediction via Mental Simulation with Span Selection
Jiafei Duan, Samson Yu Bai Jian, Soujanya Poria, Bihan Wen, Cheston Tan
ECCV (35)5
2022 TDAM: Top-Down Attention Module for Contextually Guided Feature Selection in CNNs
Shantanu Jaiswal, Basura Fernando, Cheston Tan
ECCV (25)3
2022 A Survey on Machine Learning Approaches for Modelling Intuitive Physics
abstract
Research in cognitive science has provided extensive evidence of human cognitive ability in performing physical reasoning of objects from noisy perceptual inputs. Such a cognitive ability is commonly known as intuitive physics. With advancements in deep learning, there is an increasing interest in building intelligent systems that are capable of performing physical reasoning from a given scene for the purpose of building better AI systems. As a result, many contemporary approaches in modelling intuitive physics for machine cognition have been inspired by literature from cognitive science. Despite the wide range of work in physical reasoning for machine cognition, there is a scarcity of reviews that organize and group these deep learning approaches. Especially at the intersection of intuitive physics and artificial intelligence, there is a need to make sense of the diverse range of ideas and approaches. Therefore, this paper presents a comprehensive survey of recent advances and techniques in intuitive physics-inspired deep learning approaches for physical reasoning. The survey will first categorize existing deep learning approaches into three facets of physical reasoning before organizing them into three general technical approaches and propose six categorical tasks of the field. Finally, we highlight the challenges of the current field and present some future research directions.
Jiafei Duan, Arijit Dasgupta, Jason Fischer, Cheston Tan
IJCAI4
2021 A Diagnostic Study Of Visual Question Answering With Analogical Reasoning
abstract
The deep learning community has made rapid progress in low-level visual perception tasks such as object localization, detection and segmentation. However, for tasks such as Visual Question Answering (VQA) and visual language grounding that require high-level reasoning abilities, huge gaps still exist between artificial systems and human intelligence. In this work, we perform a diagnostic study on recent popular VQA in terms of analogical reasoning. We term it as Analogical VQA, where a system needs to reason on a group of images to find analogical relations among them in order to correctly answer a natural language question. To study the task in depth, we propose an initial diagnostic synthetic dataset CLEVR-Analogy, which tests a range of analogical reasoning abilities (e.g. reasoning on object attributes, spatial relationships, existence, and arithmetic analogies). We benchmark various recent state-of-the-art methods on our dataset and compare the results against human performance, and discover that existing systems fall shorts when facing analogical reasoning involving spatial relationships. The dataset and code will be publicly available to facilitate future research.
Hongyuan Zhu 0002, Ying Sun 0001, Dongkyu Choi, Cheston Tan, Joo-Hwee Lim
ICIP5
2021 Fault-Tolerant Federated Reinforcement Learning with Theoretical Guarantee
abstract
The growing literature of Federated Learning (FL) has recently inspired Federated Reinforcement Learning (FRL) to encourage multiple agents to federatively build a better decision-making policy without sharing raw trajectories. Despite its promising applications, existing works on FRL fail to I) provide theoretical analysis on its convergence, and II) account for random system failures and adversarial attacks. Towards this end, we propose the first FRL framework the convergence of which is guaranteed and tolerant to less than half of the participating agents being random system failures or adversarial attackers. We prove that the sample efficiency of the proposed framework is guaranteed to improve with the number of agents and is able to account for such potential failures or attacks. All theoretical results are empirically verified on various RL benchmark tasks.
Flint Xiaofeng Fan, Yining Ma 0001, Zhongxiang Dai, Cheston Tan, Kian Hsiang Low
NeurIPS5
2021 A comprehensive survey of procedural video datasets
Hui Li Tan, Hongyuan Zhu 0002, Joo-Hwee Lim, Cheston Tan
Comput. Vis. Image Underst.4
2020 Actionet: An Interactive End-To-End Platform For Task-Based Data Collection And Augmentation In 3D Environment
abstract
The problem of task planning for artificial agents remains largely unsolved. While there has been increasing interest in data-driven approaches for the study of task planning for artificial agents, a significant remaining bottleneck is the dearth of large-scale comprehensive task-based datasets. In this paper, we present ActioNet, an interactive end-to-end platform for data collection and augmentation of task-based dataset in 3D environment. Using ActioNet, we collected a large-scale comprehensive task-based dataset, comprising over 3000 hierarchical task structures and videos. Using the hierarchical task structures, the videos are further augmented across 50 different scenes to give over 150,000 video. To our knowledge, ActioNet is the first interactive end-to-end platform for such task-based dataset generation and the accompanying dataset is the largest task-based dataset of such comprehensive nature. The ActioNet platform and dataset will be made available to facilitate research in hierarchical task planning.11https://github.com/SamsonYuBaiJian/actionet
Jiafei Duan, Samson Yu Bai Jian, Hui Li Tan, Cheston Tan
ICIP4
2020 6D Pose Estimation with Correlation Fusion
abstract
6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illumination, so it is important to complement them with depth information. However, existing methods using RGB-D data cannot adequately exploit consistent and complementary information between RGB and depth modalities. In this paper, we present a novel method to effectively consider the correlation within and across both modalities with attention mechanism to learn discriminative and compact multi-modal features. Then, effective fusion strategies for intra- and inter-correlation modules are explored to ensure efficient information flow between RGB and depth. To our best knowledge, this is the first work to explore effective intra- and inter-modality fusion in 6D pose estimation. The experimental results show that our method can achieve the state-of-the-art performance on LineMOD and YCB-Video dataset. We also demonstrate that the proposed method can benefit a real-world robot grasping task by providing accurate object pose estimation.
Hongyuan Zhu 0002, Ying Sun 0001, Cihan Acar, Yan Wu 0002, Liyuan Li, Cheston Tan, Joo-Hwee Lim
ICPR8
2020 Weakly Supervised Gaussian Networks for Action Detection
abstract
Detecting temporal extents of human actions in videos is a challenging computer vision problem that requires detailed manual supervision including frame-level labels. This expensive annotation process limits deploying action detectors to a limited number of categories. We propose a novel method, called WSGN, that learns to detect actions from weak supervision, using only video-level labels. WSGN learns to exploit both video-specific and dataset-wide statistics to predict relevance of each frame to an action category. This strategy leads to significant gains in action detection for two standard benchmarks THU-MOS14 and Charades. Our method obtains excellent results compared to state-of-the-art methods that uses similar features and loss functions on THUMOS14 dataset. Similarly, our weakly supervised method is only 0.3% mAP behind a state-of-the-art supervised method on challenging Charades dataset for action localization.
Basura Fernando, Cheston Tan, Hakan Bilen
WACV2
2019 An End-To-End Network for Generating Social Relationship Graphs
abstract
Socially-intelligent agents are of growing interest in artificial intelligence. To this end, we need systems that can understand social relationships in diverse social contexts. Inferring the social context in a given visual scene not only involves recognizing objects, but also demands a more in-depth understanding of the relationships and attributes of the people involved. To achieve this, one computational approach for representing human relationships and attributes is to use an explicit knowledge graph, which allows for high-level reasoning. We introduce a novel end-to-end-trainable neural network that is capable of generating a Social Relationship Graph - a structured, unified representation of social relationships and attributes - from a given input image. Our Social Relationship Graph Generation Network (SRG-GN) is the first to use memory cells like Gated Recurrent Units (GRUs) to iteratively update the social relationship states in a graph using scene and attribute context. The neural network exploits the recurrent connections among the GRUs to implement message passing between nodes and edges in the graph, and results in significant improvement over previous methods for social relationship recognition.
Arushi Goel, Keng Teck Ma, Cheston Tan
CVPR3
2018 A Probabilistic Model of Social Working Memory for Information Retrieval in Social Interactions
abstract
Social working memory (SWM) plays an important role in navigating social interactions. Inspired by studies in psychology, neuroscience, cognitive science, and machine learning, we propose a probabilistic model of SWM to mimic human social intelligence for personal information retrieval (IR) in social interactions. First, we establish a semantic hierarchy as social long-term memory to encode personal information. Next, we propose a semantic Bayesian network as the SWM, which integrates the cognitive functions of accessibility and self-regulation. One subgraphical model implements the accessibility function to learn the social consensus about IR-based on social information concept, clustering, social context, and similarity between persons. Beyond accessibility, one more layer is added to simulate the function of self-regulation to perform the personal adaptation to the consensus based on human personality. Two learning algorithms are proposed to train the probabilistic SWM model on a raw dataset of high uncertainty and incompleteness. One is an efficient learning algorithm of Newton's method, and the other is a genetic algorithm. Systematic evaluations show that the proposed SWM model is able to learn human social intelligence effectively and outperforms the baseline Bayesian cognitive model. Toward real-world applications, we implement our model on Google Glass as a wearable assistant for social interaction.
Liyuan Li, Qianli Xu, Tian Gan 0002, Cheston Tan, Joo-Hwee Lim
IEEE Trans. Cybern.4
2017 Object Detection Meets Knowledge Graphs
abstract
Object detection in images is a crucial task in computer vision, with important applications ranging from security surveillance to autonomous vehicles. Existing state-of-the-art algorithms, including deep neural networks, only focus on utilizing features within an image itself, largely neglecting the vast amount of background knowledge about the real world. In this paper, we propose a novel framework of knowledge-aware object detection, which enables the integration of external knowledge such as knowledge graphs into any object detection algorithm. The framework employs the notion of semantic consistency to quantify and generalize knowledge, which improves object detection through a re-optimization process to achieve better consistency with background knowledge. Finally, empirical evaluation on two benchmark datasets show that our approach can significantly increase recall by up to 6.3 points without compromising mean average precision, when compared to the state-of-the-art baseline.
Yuan Fang 0001, Kingsley Kuan, Jie Lin 0001, Cheston Tan, Vijay Chandrasekhar 0001
IJCAI4
2017 The effect of different types of navigation assistance on indoor scene memorability
abstract
With the rapid growing of wearable computing devices, indoor navigation guidance will become popular in the near future like the GPS-based navigation tools for drivers today. However, how the guided indoor navigation affects human’s memory of a novel environment has not been well studied. In this paper, we investigate route memory with three types of navigation assistance, that is, 2D map, wearable navigation assistant, and human usher. Twenty participants were asked to remember the route while being guided through a novel indoor environment. Our results show that the participants have similar patterns in remembering visual scenes, even using different types of assistance. These findings support previous work on scene memorability and provide the new insight that scene memorability is not affected by the type of navigation guidance. This may indicate that spatial working memory and visual memory are dissociated. We also show that scenes with navigation information are more memorable than scenes without such information. Finally, we provide some evidence that the location of a scene is linked to its memorability. In general, our findings provide valuable information about indoor scene memorability.
Michal Mukawa, Cheston Tan, Joo-Hwee Lim, Qianli Xu, Liyuan Li
Behav. Inf. Technol.2
2017 A Wearable Virtual Usher for Vision-Based Cognitive Indoor Navigation
abstract
Inspired by progresses in cognitive science, artificial intelligence, computer vision, and mobile computing technologies, we propose and implement a wearable virtual usher for cognitive indoor navigation based on egocentric visual perception. A novel computational framework of cognitive wayfinding in an indoor environment is proposed, which contains a context model, a route model, and a process model. A hierarchical structure is proposed to represent the cognitive context knowledge of indoor scenes. Given a start position and a destination, a Bayesian network model is proposed to represent the navigation route derived from the context model. A novel dynamic Bayesian network (DBN) model is proposed to accommodate the dynamic process of navigation based on real-time first-person-view visual input, which involves multiple asynchronous temporal dependencies. To adapt to large variations in travel time through trip segments, we propose an online adaptation algorithm for the DBN model, leading to a self-adaptive DBN. A prototype system is built and tested for technical performance and user experience. The quantitative evaluation shows that our method achieves over 13% improvement in accuracy as compared to baseline approaches based on hidden Markov model. In the user study, our system guides the participants to their destinations, emulating a human usher in multiple aspects.
Liyuan Li, Qianli Xu, Vijay Chandrasekhar 0001, Joo-Hwee Lim, Cheston Tan, Michal Mukawa
IEEE Trans. Cybern.5
2017 Summarization of Egocentric Videos: A Comprehensive Survey
abstract
The introduction of wearable video cameras (e.g., GoPro) in the consumer market has promoted video life-logging, motivating users to generate large amounts of video data. This increasing flow of first-person video has led to a growing need for automatic video summarization adapted to the characteristics and applications of egocentric video. With this paper, we provide the first comprehensive survey of the techniques used specifically to summarize egocentric videos. We present a framework for first-person view summarization and compare the segmentation methods and selection algorithms used by the related work in the literature. Next, we describe the existing egocentric video datasets suitable for summarization and, then, the various evaluation methods. Finally, we analyze the challenges and opportunities in the field and propose new lines of research.
Ana Garcia del Molino, Cheston Tan, Joo-Hwee Lim, Ah-Hwee Tan
IEEE Trans. Hum. Mach. Syst.2
2014 An Automated Estimator of Image Visual Realism Based on Human Cognition
abstract
Assessing the visual realism of images is increasingly becoming an essential aspect of fields ranging from computer graphics (CG) rendering to photo manipulation. In this paper we systematically evaluate factors underlying human perception of visual realism and use that information to create an automated assessment of visual realism. We make the following unique contributions. First, we established a benchmark dataset of images with empirically determined visual realism scores. Second, we identified attributes potentially related to image realism, and used correlational techniques to determine that realism was most related to image naturalness, familiarity, aesthetics, and semantics. Third, we created an attributes-motivated, automated computational model that estimated image visual realism quantitatively. Using human assessment as a benchmark, the model was below human performance, but outperformed other state-of-the-art algorithms.
Shaojing Fan, Tian-Tsong Ng, Jonathan S. Herberg, Bryan L. Koenig, Cheston Tan, Rangding Wang
CVPR5
2014 Incremental Graph Clustering for Efficient Retrieval from Streaming Egocentric Video Data
abstract
With wearable devices like Google Glass, it will soon become possible to record everything we see. We envision a system where one's entire visual memory is captured, stored and indexed. One of the biggest challenges is the scale of the retrieval problem. In this work, we focus on how to organize streaming egocentric video data. Egocentric video data is highly redundant, in that, we see several objects and scenes repeatedly as we go about our lives. To exploit this redundancy, we propose an evolving sparse-graph representation for egocentric video data. We propose an incremental local density clustering scheme, which learns salient objects and scenes for streaming egocentric video data. We use the density clustering scheme to prune redundant data in the database. For image-retrieval applications, by retaining only representative nodes from dense sub graphs in the streaming data source, we show we can achieve 90% of peak recall by retaining only 1% of data, with a significant 18% improvement in absolute recall over naive uniform sub sampling of the egocentric video data.
Vijay Chandrasekhar 0001, Cheston Tan, Wu Min, Liyuan Li, Xiaoli Li 0001, Joo-Hwee Lim
ICPR2
2014 A wearable virtual guide for context-aware cognitive indoor navigation
abstract
In this paper, we explore a new way to provide context-aware assistance for indoor navigation using a wearable vision system. We investigate how to represent the cognitive knowledge of wayfinding based on first-person-view videos in real-time and how to provide context-aware navigation instructions in a human-like manner. Inspired by the human cognitive process of wayfinding, we propose a novel cognitive model that represents visual concepts as a hierarchical structure. It facilitates efficient and robust localization based on cognitive visual concepts. Next, we design a prototype system that provides intelligent context-aware assistance based on the cognitive indoor navigation knowledge model. We conducted field tests and evaluated the system's efficacy by benchmarking it against traditional 2D maps and human guidance. The results show that context-awareness built on cognitive visual perception enables the system to emulate the efficacy of a human guide, leading to positive user experience.
Qianli Xu, Liyuan Li, Joo-Hwee Lim, Cheston Tan, Michal Mukawa, Gang S. Wang
Mobile HCI4
2014 Robust and Efficient Saliency Modeling from Image Co-Occurrence Histograms
abstract
This paper presents a visual saliency modeling technique that is efficient and tolerant to the image scale variation. Different from existing approaches that rely on a large number of filters or complicated learning processes, the proposed technique computes saliency from image histograms. Several two-dimensional image co-occurrence histograms are used, which encode not only "how many" (occurrence) but also "where and how" (co-occurrence) image pixels are composed into a visual image, hence capturing the "unusualness" of an object or image region that is often perceived by either global "uncommonness" (i.e., low occurrence frequency) or local "discontinuity" with respect to the surrounding (i.e., low co-occurrence frequency). The proposed technique has a number of advantageous characteristics. It is fast and very easy to implement. At the same time, it involves minimal parameter tuning, requires no training, and is robust to image scale variation. Experiments on the AIM dataset show that a superior shuffled AUC (sAUC) of 0.7221 is obtained, which is higher than the state-of-the-art sAUC of 0.7187.
Shijian Lu, Cheston Tan, Joo-Hwee Lim
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Human Perception of Visual Realism for Photo and Computer-Generated Face Images
abstract
Computer-generated (CG) face images are common in video games, advertisements, and other media. CG faces vary in their degree of realism, a factor that impacts viewer reactions. Therefore, efficient control of visual realism of face images is important. Efficient control is enabled by a deep understanding of visual realism perception: the extent to which viewers judge an image as a real photograph rather than a CG image. Across two experiments, we explored the processes involved in visual realism perception of face images. In Experiment 1, participants made visual realism judgments on original face images, inverted face images, and images of faces that had the top and bottom halves misaligned. In Experiment 2, participants made visual realism judgments on original face images, scrambled faces, and images that showed different parts of faces. Our findings indicate that both holistic and piecemeal processing are involved in visual realism perception of faces, with holistic processing becoming more dominant when resolution is lower. Our results also suggest that shading information is more important than color for holistic processing, and that inversion makes visual realism judgments harder for realistic images but not for unrealistic images. Furthermore, we found that eyes are the most influential face part for visual realism, and face context is critical for evaluating realism of face parts. To the best of our knowledge, this work is a first realism-centric study attempting to bridge the human perception of visual realism on face images with general face perception tasks.
Shaojing Fan, Rangding Wang, Tian-Tsong Ng, Cheston Tan, Jonathan S. Herberg, Bryan L. Koenig
ACM Trans. Appl. Percept.4
2013 Visual Recognition using a Combination of Shape and Color Features
Sepehr Jalali, Cheston Tan, Joo-Hwee Lim, Jo Yew Tham, Sim Heng Ong, Paul J. Seekings, Elizabeth A. Taylor
CogSci2
2013 Encoding Co-occurrence of Features in the HMAX Model
Sepehr Jalali, Cheston Tan, Joo-Hwee Lim, Jo Yew Tham, Sim Heng Ong, Paul J. Seekings, Elizabeth A. Taylor
CogSci2
2013 A Wearable Cognitive Vision System for Navigation Assistance in Indoor Environment
Liyuan Li, Gang S. Wang, Weixun Goh, Joo-Hwee Lim, Cheston Tan
ICONIP (3)5
2013 The use of optical and sonar images in the human and dolphin brain for image classification
abstract
In this paper we propose a new biologically inspired model which simulates the visual pathways in the human brain used for classification of matching optical and sonar derived images. Marine mammals, such as dolphins, that live in waters with poor optical clarity and low light levels such as littoral zones, use a combination of optical vision and biosonar to navigate and hunt for prey. Given that dolphins have evolved a synergistic combination of optical visual input and acoustic/sonar input, the primary focus of this paper is on reaching a similar level of synergy for a diver or Autonomous Underwater Vehicle (AUV) platform equipped with a system to extend the range and resolution of vision in poor ambient visibility. We propose a biologically inspired model that combines and processes visual images acquired via optical and acoustic pathways and show that the combined model enhances the accuracy of automatic classification of target objects in underwater images.
Sepehr Jalali, Paul J. Seekings, Cheston Tan, Aiswarya Ratheesh, Joo-Hwee Lim, Elizabeth A. Taylor
IJCNN3
2013 Classification of marine organisms in underwater images using CQ-HMAX biologically inspired color approach
abstract
In many coastal environments, particularly in tropical zones, coral reef ecosystems have exceptional biodiversity, contribute to coastal defense, provide unique and important habitats and valuable commercial resources. Assessment of environmental impacts on biodiversity in such areas are increasingly important to mitigate potential adverse effects on specific ecosystems. Visual classification of marine organisms is necessary for population estimates of individual species of corals or other benthic organisms. In this paper, we introduce a new image dataset of benthic organisms that are of different colors, shapes, scales, visibility and are taken from different viewpoints. We evaluate several different classification approaches on this dataset, and show that CQ-HMAX, our new biologically inspired approach to utilizing color information for object and scene recognition, that is inspired by the characteristics of color- and object-selective neurons in the high-level inferotemporal (IT) cortex of the primate visual system, results in better classification results in comparison with existing computational models such as support vectors machines, SIFT based approaches and the HMAX biologically inspired approach. We show that concatenating our model which encodes color information with the HMAX model which encodes grayscale shape information results in the highest classification accuracy.
Sepehr Jalali, Paul J. Seekings, Cheston Tan, Hazel Z. W. Tan, Joo-Hwee Lim, Elizabeth A. Taylor
IJCNN3
2013 Neural representation of action sequences: how far can a simple snippet-matching model take us?
abstract
The macaque Superior Temporal Sulcus (STS) is a brain area that receives and integrates inputs from both the ventral and dorsal visual processing streams (thought to specialize in form and motion processing respectively). For the processing of articulated actions, prior work has shown that even a small population of STS neurons contains sufficient information for the decoding of actor invariant to action, action invariant to actor, as well as the specific conjunction of actor and action. This paper addresses two questions. First, what are the invariance properties of individual neural representations (rather than the population representation) in STS? Second, what are the neural encoding mechanisms that can produce such individual neural representations from streams of pixel images? We find that a baseline model, one that simply computes a linear weighted sum of ventral and dorsal responses to short action “snippets”, produces surprisingly good fits to the neural data. Interestingly, even using inputs from a single stream, both actor-invariance and action-invariance can be produced simply by having different linear weights.
Cheston Tan, Jedediah M. Singer, Thomas Serre, David L. Sheinberg, Tomaso A. Poggio
NIPS1
2006 Systematic gene function prediction from gene expression data by using a fuzzy nearest-cluster method
abstract
BACKGROUND: Quantitative simultaneous monitoring of the expression levels of thousands of genes under various experimental conditions is now possible using microarray experiments. However, there are still gaps toward whole-genome functional annotation of genes using the gene expression data. RESULTS: In this paper, we propose a novel technique called Fuzzy Nearest Clusters for genome-wide functional annotation of unclassified genes. The technique consists of two steps: an initial hierarchical clustering step to detect homogeneous co-expressed gene subgroups or clusters in each possibly heterogeneous functional class; followed by a classification step to predict the functional roles of the unclassified genes based on their corresponding similarities to the detected functional clusters. CONCLUSION: Our experimental results with yeast gene expression data showed that the proposed method can accurately predict the genes' functions, even those with multiple functional roles, and the prediction performance is most independent of the underlying heterogeneity of the complex functional classes, as compared to the other conventional gene function prediction approaches.
Xiaoli Li 0001, Cheston Tan, See-Kiong Ng
BMC Bioinform.2