EDBT 2026 Demo / reviewers in the wild / expert
Qianli Xu
dblp:30/3276
· DBLP profile ↗
38ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0003-0105-5903ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting Social Engagement of Elderly From Lifelog Image-streams to Identify Effective Cues for Autobiographic RecallabstractLifelog images captured automatically by wearable cam-eras serve as effective cues that induce Autobiographic Memory Recall (AMR) of social interactions. This is very useful for personalized memory interventions. However, manual selection of images for such therapy imposes significant load on the caregivers who need to browse through a voluminous collection of images. To reduce this load, auto-mated tools that identify moments involving significant engagement of the camera wearer in social interactions are needed. To achieve this, we reannotate images extracted from public lifelog datasets for the presence of non-verbal social signals and the perceived engagement of the lifelogger during interactions. We use this data to develop models and explore how social signals and the detected intensity of social engagement are helpful for predicting AMR. We show that understanding visual social engagement can enhance AMR prediction, demonstrating the potential of the models in reducing caregivers’ effort. Vengateswaran Subramaniam, Vigneshwaran Subbaraju, Debaditya Roy, Pramath Krishna, Thivya Kandappu, Qianli Xu |
WACV | 6 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingabstractLarge Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we propose the Visually Grounded Text-Video Prompting (VG-TVP) method which is a novel LLM-empowered Multimodal Procedural Planning (MPP) framework. It generates cohesive text and video procedural plans given a specified high-level objective. The main challenges are achieving textual and visual informativeness, temporal coherence, and accuracy in procedural plans. VG-TVP leverages the zero-shot reasoning capability of LLMs, the video-to-text generation ability of the video captioning models, and the text-to-video generation ability of diffusion models. VG-TVP improves the interaction between modalities by proposing a novel Fusion of Captioning (FoC) method and using Text-to-Video Bridge (T2V-B) and Video-to-Text Bridge (V2T-B). They allow LLMs to guide the generation of visually-grounded text plans and textual-grounded video plans. To address the scarcity of datasets suitable for MPP, we have curated a new dataset called Daily-Life Task Procedural Plans (Daily-PP). We conduct comprehensive experiments and benchmarks to evaluate human preferences (regarding textual and visual informativeness, temporal coherence, and plan accuracy). Our VG-TVP method outperforms unimodal baselines on the Daily-PP dataset. Muhammet Furkan Ilaslan, Ali Koksal, Qinghong Lin, Burak Satar, Zheng Shou 0001, Qianli Xu |
AAAI | 6 |
| 2025 | SPASCA: Social Presence and Support with Conversational Agent for Persons Living with DementiaabstractWe present SPASCA - a conversational AI system that promotes psychological and cognitive well-being of persons living with dementia (PLWD). This system features an AI agent that provides social presence and support to PLWD through verbal communications, without physical presence of human caregivers. The system integrates (1) a novel dialogue model that generates dialogue items relevant to the user's experiences and lifestyle, (2) a digital avatar in the form of a talking head with the identity of a caregiver who is familiar to the demented user. We develop prototypes that adopt various interaction modalities and conversational styles and report the pros and cons of different system configurations through expert review. Our system shows the potential of conversational AI for personalized and affordable healthcare services. Ali Koksal, Jingjing Gu, Kotaro Hara, Joo-Hwee Lim, Qianli Xu |
AAAI | 6 |
| 2025 | DOTA: Distributional Test-time Adaptation of Vision-Language ModelsabstractVision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when significant distribution gaps exist between training and test data, while fine-tuning for diverse scenarios is often costly. Cache-based test-time adapters offer an efficient alternative by storing representative test samples to guide subsequent classifications. Yet, these methods typically employ naive cache management with limited capacity, leading to severe catastrophic forgetting when samples are inevitably dropped during updates. In this paper, we propose DOTA (DistributiOnal Test-time Adaptation), a simple yet effective method addressing this limitation. Crucially, instead of merely memorizing individual test samples, DOTA continuously estimates the underlying distribution of the test data stream. Test-time posterior probabilities are then computed using these dynamically estimated distributions via Bayes' theorem for adaptation. This distribution-centric approach enables the model to continually learn and adapt to the deployment environment. Extensive experiments validate that DOTA significantly mitigates forgetting and achieves state-of-the-art performance compared to existing methods. Zongbo Han, Jialong Yang, Junfan Li, Qianli Xu, Zheng Shou 0001, Changqing Zhang 0002 |
NeurIPS | 5 |
| 2024 | Exploring Conversations between a Practitioner and a Person with DementiaabstractIn social service centers, practitioners engage in conversations with clients with dementia to facilitate their daily activities and provide support when they are distressed. However, the nature of the care demands the practitioner’s active engagement, which becomes difficult to deliver as the number of people who need care expands. Researchers have been investigating the efficacy of developing agents that assume conversational tasks to alleviate this work. To contribute to the future design of agents for caregiving, we collected and analyzed ten conversations between clients with mild dementia and practitioners who provide care. Our analyses of turn-taking dynamics and dialogue acts with 15k utterances uncovered patterns such as noticeable differences in clients’ and practitioners’ conversational dynamics and the prevalence of neutral-toned, question-oriented utterances by practitioners. We then prototyped a large language model-based script that generates responses to client utterances. We found potential approaches and challenges for making its utterance pattern more similar to that of a practitioner. Kotaro Hara, Rosiana Natalie, Wei Soon Cheong, Jingjing Gu, Qianli Xu |
ASSETS | 5 |
| 2024 | See, Predict, Plan: Diffusion for Procedure Planning in Robotic Surgical Videos
Ziyuan Zhao, Fen Fang, Xulei Yang, Qianli Xu, Cuntai Guan, Shaohua Kevin Zhou |
MICCAI (6) | 4 |
| 2024 | VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationabstractA well-known dilemma in large vision-language models (e.g., GPT-4, LLaVA) is that while increasing the number of vision tokens generally enhances visual understanding, it also significantly raises memory and computational costs, especially in long-term, dense video frame streaming scenarios. Although learnable approaches like Q-Former and Perceiver Resampler have been developed to reduce the vision token burden, they overlook the context causally modeled by LLMs (i.e., key-value cache), potentially leading to missed visual cues when addressing user queries. In this paper, we introduce a novel approach to reduce vision compute by leveraging redundant vision tokens ``skipping layers'' rather than decreasing the number of vision tokens. Our method, VideoLLM-MoD, is inspired by mixture-of-depths LLMs and addresses the challenge of numerous vision tokens in long-term or streaming video. Specifically, for certain transformer layer, we learn to skip the computation for a high proportion (e.g., 80\%) of vision tokens, passing them directly to the next layer. This approach significantly enhances model efficiency, achieving approximately 42% time and 30% memory savings for the entire training. Moreover, our method reduces the computation in the context and avoid decreasing the vision tokens, thus preserving or even improving performance compared to the vanilla model. We conduct extensive experiments to demonstrate the effectiveness of VideoLLM-MoD, showing its state-of-the-art results on multiple benchmarks, including narration, forecasting, and summarization tasks in COIN, Ego4D, and Ego-Exo4D datasets. The code and checkpoints will be made available at github.com/showlab/VideoLLM-online. Joya Chen, Qinghong Lin, Qimeng Wang, Yan Gao 0017, Qianli Xu, Tong Xu 0001, Yao Hu 0002, Enhong Chen, Zheng Shou 0001 |
NeurIPS | 6 |
| 2024 | Localizing discriminative regions for fine-grained visual recognition: One could be better than many
Fen Fang, Yun Liu 0011, Qianli Xu |
Neurocomputing | 3 |
| 2024 | Enhancing Representation Learning With Spatial Transformation and Early Convolution for Reinforcement Learning-Based Small Object DetectionabstractAlthough object detection has achieved significant progress in the past decade, detecting small objects is still far from satisfactory due to the high variability of object scales and complex backgrounds. The common way to enhance small object detection is to use high-resolution (HR) images. However, this method incurs huge computational resources which grow squarely with the resolution of images. To achieve both accuracy and efficiency, we propose a novel reinforcement learning framework that employs an efficient policy network consisting of a Spatial Transformation Network to enhance the state representation learning and a Transformer model with early convolution to improve feature extraction. Our method has two main steps: (1) coarse location query (CLQ), where an RL agent is trained to predict the locations of small objects on low-resolution (LR) (down-sampled version of HR) images; (2) context-sensitive object detection where HR image patches are used to detect objects on the selected coarse locations and LR image patches on background areas (containing no small objects). In this way, we can obtain high detection performance on small objects while avoiding unnecessary computation on background areas. The proposed method has been tested and benchmarked on various datasets. On the Caltech Pedestrians Detection and Web Pedestrians datasets, the proposed method improves the detection accuracy by 2%, while reducing the number of processed pixels. On the Vision meets Drone object detection dataset and the Oil and Gas Storage Tank dataset, the proposed method outperforms the state-of-the-art (SotA) methods. On MS COCO mini-val set, our method outperforms SotA methods on small object detection, while also achieving comparable performance on medium and large objects. Fen Fang, Wenyu Liang, Qianli Xu, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented CollaborationsabstractThe usage of exocentric and egocentric videos in Video Question Answering (VQA) is a new endeavor in human-robot interaction and collaboration studies.Particularly for egocentric videos, one may leverage eye-gaze information to understand human intentions during the task.In this paper, we build a novel task-oriented VQA dataset, called GazeVQA, for collaborative tasks where gaze information is captured during the task process.GazeVQA is designed with a novel QA format that covers thirteen different reasoning types to capture multiple aspects of task information and user intent.For each participant, GazeVQA consists of more than 1,100 textual questions and more than 500 labeled images that were annotated with the assistance of the Segment Anything Model.In total, 2,967 video clips, 12,491 labeled images, and 25,040 questions from 22 participants were included in the dataset.Additionally, inspired by the assisting models and common ground theory for industrial task collaboration, we propose a new AI model called AssistGaze that is designed to answer the questions with three different answer types, namely textual, image, and video.AssistGaze can effectively ground the perceptual input into semantic information while reducing ambiguities.We conduct comprehensive experiments to demonstrate the challenges of GazeVQA 1 and the effectiveness of AssistGaze 2 . Muhammet Furkan Ilaslan, Chenan Song, Joya Chen, Difei Gao, Weixian Lei, Qianli Xu, Joo Lim, Zheng Shou 0001 |
EMNLP | 6 |
| 2023 | Data Augmentation Using Corner CutMix and an Auxiliary Self-Supervised LossabstractDeep convolutional neural networks (CNNs) have achieved remarkable success in computer vision tasks, but their training is susceptible to overfitting when the training sample size is insufficient. In this paper, we introduce Corner CutMix, a novel data augmentation technique for CNN training. During training, Corner CutMix randomly selects a region from one of four corner areas in an image and replaces it with a randomly chosen region from a distractor image. Additionally, we design an auxiliary self-supervised loss function to learn the position of the selected corner region, thereby improving the transferability and generalizability of the learned representation. Corner CutMix is easy to implement, adding little computational overhead, and can be combined with other augmentation methods such as random cropping, color distortion, and flipping. Our extensive classification task experiments in self-supervised learning on public datasets (e.g., CIFAR10, CIFAR100, and STL10) demonstrate the effectiveness of Corner CutMix, which consistently outperforms strong baselines such as CutOut and CutMix. Fen Fang, Nhat M. Hoang, Qianli Xu, Joo-Hwee Lim |
ICIP | 3 |
| 2022 | Improving Generalization of Reinforcement Learning Using a Bilinear Policy NetworkabstractIn deep reinforcement learning (DRL), the agent is usually trained on seen environments by optimizing a policy network. However, it is difficult to be generalized to unseen environments properly, even when the environmental variations are insignificant. This is partly because the policy network cannot effectively learn the representation of visual difference that is subtle among highly similar states in the environments. Because a bilinear structured model containing two feature extractors allows pairwise feature interactions in a translation-ally invariant manner which makes it particularly useful for subtle difference recognition among highly similar states, in this work, a bilinear policy network is employed to enhance representation learning, and thus to improve generalization of the DRL. The proposed bilinear policy network is tested on various DRL task, including a control task on path planning for active object detection, and Grid World, an AI game task. The test results show that the generalization of DRL can be improved by the proposed network. Fen Fang, Wenyu Liang, Yan Wu 0002, Qianli Xu, Joo-Hwee Lim |
ICIP | 4 |
| 2022 | Hierarchical Defect Detection Based On Reinforcement LearningabstractIn this paper, we propose a novel reinforcement learning (RL) based method for defects detection in high-resolution (HR) images. e.g. cracks and scratches on the surfaces of buildings, constructions, and products. Our innovation leverages RL to explore challenging images in progressive manner, using pre-trained deep learning (DL) detection as feedback mechanism. First, The DL model is pre-trained on low resolution (LR) images with relatively high defect background ratio (DBR). The RL agent is trained by optimizing a policy network according to feedback of DL model on selected regions of HR images with fairly low DBR to coarsely predict defective region by executing two actions: defective region selection and region refinement. Then, the selected defective regions are evaluated using the DL model to generate final defect region which will be mapped back to the HR images. Experimental results on HR crack and scratch images indicate that our method is able to achieve state-of-the-art performance with 0.976 and 0.965 F1-score respectively. Fen Fang, Qianli Xu, Joo-Hwee Lim |
ICIP | 2 |
| 2022 | Image Understanding With Reinforcement Learning: Auto-Tuning Image Attributes and Model Parameters for Object Detection and SegmentationabstractModels for image semantics understanding, such as deep learning (DL) models and mathematical models, are often trained on specific dataset or configured with specific parameters. When deploying such models on new tasks in a different test environment, it requires considerable effort to re-train the model or extensive expertise to tune the parameters. In this paper, we propose a smart reinforcement learning (RL) agent that could learn to tune parameters automatically to enhance model performance. The learning process is formulated as a generic control task for parameter adjustment, and applied to two use scenarios: (1) image attributes tuning to improve object detection performance on fixed DL model, and (2) parameter tuning of the mathematical model (Level Set) for image segmentation. We design a novel dynamic threshold mechanism in a multi-branch RL agent to effectively tune parameters of image qualities (for object detection) and Level Set models (for object segmentation). We conduct experiments on Pascal-VOC testing set, MS COCO validation set and a proprietary dataset of industrial components, where we achieve substantial improvement on object detection accuracy. We also perform experiments on the automatic parameter tuning of Level Set models. Results show that our method facilitates considerable performance improvement on public datasets compared with baseline method. Fen Fang, Qianli Xu, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | TAILOR: Teaching with Active and Incremental Learning for Object RegistrationabstractWhen deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions. Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
AAAI | 1 |
| 2021 | Enhancing Multi-Step Action Prediction for Active Object DetectionabstractActive vision for robots is one promising solution to open world visual detection problems. A fundamental issue is view planning, i.e., predicting next best views to capture images of interest to reduce uncertainty. While multi-step action in a reinforcement learning (RL) setup can boost the efficiency of view planning, existing methods suffer from unstable detection outcome when the Q-values of multiple branches of action advantages (i.e., action range and action type) are combined naively. To tackle this issue, we propose a novel mechanism to disentangle action range from action type through a two-stage training strategy on a deep Q-network. It combines well-crafted loss functions with respect to action range and action type to enforce separated training of these two branches. We evaluate our method on two public datasets and show that it facilitates substantial gain in view planning efficiency, while enhancing detection accuracy. Fen Fang, Qianli Xu, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 2 |
| 2021 | Towards Efficient Multiview Object Detection with Adaptive Action PredictionabstractActive vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection. We formulate the multi-object detection task as an active multiview object detection problem given the initial location of the objects. Next, we propose a novel adaptive action prediction (A2P) method built on a deep Q-learning network with a dueling architecture. The A2P method is able to perform view planning based on visual information of multiple objects; and adjust action ranges according to the task status. Evaluated on the AVD dataset, A2P leads to 21.9% increase in detection accuracy in unfamiliar environments, while improving efficiency by 22.7%. On the T-LESS dataset, multi-object detection boosts efficiency by more than 30% while achieving equivalent detection accuracy. Qianli Xu, Fen Fang, Nicolas Gauthier, Wenyu Liang, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
ICRA | 1 |
| 2021 | Predicting Event Memorability from Contextual Visual SemanticsabstractEpisodic event memory is a key component of human cognition. Predicting event memorability,i.e., to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorability according to a cued recall process. Specifically, we explore whether event memorability is contingent on the event context, as well as the intrinsic visual attributes of image cues. We design a novel experiment protocol and conduct a large-scale experiment with 47 elder subjects over 3 months. Subjects’ memory of life events is tested in a cued recall process. Using advanced visual analytics methods, we build a first-of-its-kind event memorability dataset (called R3) with rich information about event context and visual semantic features. Furthermore, we propose a contextual event memory network (CEMNet) that tackles multi-modal input to predict item-wise event memorability, which outperforms competitive benchmarks. The findings inform deeper understanding of episodic event memory, and open up a new avenue for prediction of human episodic memory. Source code is available at https://github.com/ffzzy840304/Predicting-Event-Memorability. Qianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju, Joo-Hwee Lim |
NeurIPS | 1 |
| 2021 | PrivacyPrimer: Towards Privacy-Preserving Episodic Memory Support For Older AdultsabstractBuilt-in pervasive cameras have become an integral part of mobile/wearable devices and enabled a wide range of ubiquitous applications with their ability to be "always-on". In particular, life-logging has been identified as a means to enhance the quality of life of older adults by allowing them to reminisce about their own life experiences. However, the sensitive images captured by the cameras threaten individuals' right to have private social lives and raise concerns about privacy and security in the physical world. This threat gets worse when image recognition technologies can link images to people, scenes, and objects, hence, implicitly and unexpectedly reveal more sensitive information such as social connections. In this paper, we first examine life-log images obtained from 54 older adults to extract (a) the artifacts or visual cues, and (b) the context of the image that influences an older life-logger's ability to recall the life events associated with a life-log image. We call these artifacts and contextual cues "stimuli". Using the set of stimuli extracted, we then propose a set of obfuscation strategies that naturally balances the trade-off between reminiscability and privacy (revealing social ties) while selectively obfuscating parts of the images. More specifically, our platform yields privacy-utility tradeoff by compromising, on average, modest 13.4% reminiscability scores while significantly improving privacy guarantees -- around 40% error in cloud estimation. Thivya Kandappu, Vigneshwaran Subbaraju, Qianli Xu |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Lifelog Image Retrieval Based on Semantic Relevance MappingabstractLifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method. Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Task-Oriented Multi-Modal Question Answering For Collaborative ApplicationsabstractCobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task. Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 3 |
| 2020 | Active Image Sampling on Canonical Views for Novel Object DetectionabstractTo alleviate the costly data annotation problem in deep learning-based object detection, we leverage the canonical view model for active sample selection to improve the effectiveness of learning. Inspired by the view-approximation model, we hypothesize that visual features learned from canonical views denote better representations of objects, thus boosting the effectiveness of object learning. We validate the hypothesis empirically in the context of robot learning for novel object detection. Based on this, we propose a novel on-line viewpoint exploration (OLIVE) method that (1) defines goodness-of-view by combining informativeness of visual features and consistency of model-based object detection, and (2) systematically explores and selects viewpoints to boost learning efficiency. Furthermore, we train a legacy Faster R-CNN model with a data augmentation method while leveraging data samples generated by the OLIVE pipeline. We test our method on the T-LESS dataset and show that the proposed method outperforms competitive benchmarking methods, especially when the samples are few. Qianli Xu, Fen Fang, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 1 |
| 2020 | Gesture Enhanced Comprehension of Ambiguous Human-to-Robot InstructionsabstractThis work demonstrates the feasibility and benefits of using pointing gestures, a naturally-generated additional input modality, to improve the multi-modal comprehension accuracy of human instructions to robotic agents for collaborative tasks.We present M2Gestic, a system that combines neural-based text parsing with a novel knowledge-graph traversal mechanism, over a multi-modal input of vision, natural language text and pointing. Via multiple studies related to a benchmark table top manipulation task, we show that (a) M2Gestic can achieve close-to-human performance in reasoning over unambiguous verbal instructions, and (b) incorporating pointing input (even with its inherent location uncertainty) in M2Gestic results in a significant (30%) accuracy improvement when verbal instructions are ambiguous. Dulanga Weerakoon, Vigneshwaran Subbaraju, Nipuni Karumpulli, Qianli Xu, U-Xuan Tan, Joo-Hwee Lim, Archan Misra |
ICMI | 5 |
| 2020 | Detecting Objects with High Object Region PercentageabstractObject shape is a subtle but important factor for object detection. It has been observed that the object-region-percentage (ORP) can be utilized to improve detection accuracy for elongated objects, which have much lower ORPs than other types of objects. In this paper, we propose an approach to improve the detection performance for objects with high ORPs. Our method consists of three steps. First, we adjust the ground truth bounding boxes of high-ORP objects to an optimal range. Second, we train an object detector, Faster R-CNN, based on adjusted bounding boxes to achieve high recall. Finally, we train a DCNN to learn the adjustment ratios towards four directions and adjust detected bounding boxes of objects to get better localization for higher precision. We evaluate the effectiveness of our method on 12 high-ORP objects in COCO and 8 objects in a proprietary gearbox dataset. The experimental results show that our method can achieve state-of-the-art performance on these objects while costing less resources in training and inference stages. Fen Fang, Qianli Xu, Liyuan Li, Joo-Hwee Lim |
ICPR | 2 |
| 2018 | Image-based Parking Place Identification for Regulating Shared Bicycle ParkingabstractWe propose a novel method and system to prevent indiscriminate parking of dockless shared bicycles using location-based geo-fencing and image-based parking place identification. The geo-fencing is used to define the approximate regions for different types of bicycle parking regulations. The parking place identification uses a method based on deep Convolutional Neural Network (DCNN) to automatically identify designated bicycle parking places from photos captured by the cyclist using a mobile phone. Combining these two modalities, the parking of shared bicycles can be restricted in designated zones in various environments. Experiments are conducted using photos taken from the designated parking places with different parking indications at various locations. We evaluate the performance of the image-based parking place identification and use heatmaps to analyze potential features that are exploit by the DCNN models. The method achieves high performance on the testing dataset; and the features used for parking place identification are largely consistent with human perceptions. Shudong Xie, Qianli Xu, Fen Fang, Liyuan Li |
ICARCV | 3 |
| 2018 | Personalized Serious Games for Cognitive Intervention with Lifelog Visual AnalyticsabstractThis paper presents a novel serious game app and a method to cre- ate and integrate personalized game content based on lifelog visual analytics. The main objective is to extract personalized content from visual lifelogs, integrate it into mobile games, and evaluate the effect of personalization on user experience. First, a suite of visual analysis methods is proposed to extract semantic informa- tion from visual lifelogs and discover the association among the lifelog entities. The outcome is dataset that contains augmented and personal lifelog images. Next, a mobile game app is developed that makes use of the dataset as game content. Finally, an experiment is conducted to evaluate user gameplay behaviors in the wild over three months, where a mixture of generic and personalized game content is deployed. It is observed that user adherence is heightened by personalized game content as compared to generic content. Also observed is a higher enjoyment level in personalized than generic game content. The result provides the first empirical evidence of the effect of personalized games on user adherence and preference for cognitive intervention. This work paves the way for effective cognitive training with user-generated content. Qianli Xu, Vigneshwaran Subbaraju, Chee How Cheong, Aijing Wang, Kathleen Kang, Munirah Bashir, Yanhong Dong, Liyuan Li, Joo-Hwee Lim |
ACM Multimedia | 1 |
| 2018 | A Probabilistic Model of Social Working Memory for Information Retrieval in Social InteractionsabstractSocial working memory (SWM) plays an important role in navigating social interactions. Inspired by studies in psychology, neuroscience, cognitive science, and machine learning, we propose a probabilistic model of SWM to mimic human social intelligence for personal information retrieval (IR) in social interactions. First, we establish a semantic hierarchy as social long-term memory to encode personal information. Next, we propose a semantic Bayesian network as the SWM, which integrates the cognitive functions of accessibility and self-regulation. One subgraphical model implements the accessibility function to learn the social consensus about IR-based on social information concept, clustering, social context, and similarity between persons. Beyond accessibility, one more layer is added to simulate the function of self-regulation to perform the personal adaptation to the consensus based on human personality. Two learning algorithms are proposed to train the probabilistic SWM model on a raw dataset of high uncertainty and incompleteness. One is an efficient learning algorithm of Newton's method, and the other is a genetic algorithm. Systematic evaluations show that the proposed SWM model is able to learn human social intelligence effectively and outperforms the baseline Bayesian cognitive model. Toward real-world applications, we implement our model on Google Glass as a wearable assistant for social interaction. Liyuan Li, Qianli Xu, Tian Gan 0002, Cheston Tan, Joo-Hwee Lim |
IEEE Trans. Cybern. | 2 |
| 2017 | The effect of different types of navigation assistance on indoor scene memorabilityabstractWith the rapid growing of wearable computing devices, indoor navigation guidance will become popular in the near future like the GPS-based navigation tools for drivers today. However, how the guided indoor navigation affects human’s memory of a novel environment has not been well studied. In this paper, we investigate route memory with three types of navigation assistance, that is, 2D map, wearable navigation assistant, and human usher. Twenty participants were asked to remember the route while being guided through a novel indoor environment. Our results show that the participants have similar patterns in remembering visual scenes, even using different types of assistance. These findings support previous work on scene memorability and provide the new insight that scene memorability is not affected by the type of navigation guidance. This may indicate that spatial working memory and visual memory are dissociated. We also show that scenes with navigation information are more memorable than scenes without such information. Finally, we provide some evidence that the location of a scene is linked to its memorability. In general, our findings provide valuable information about indoor scene memorability. Michal Mukawa, Cheston Tan, Joo-Hwee Lim, Qianli Xu, Liyuan Li |
Behav. Inf. Technol. | 4 |
| 2017 | A Wearable Virtual Usher for Vision-Based Cognitive Indoor NavigationabstractInspired by progresses in cognitive science, artificial intelligence, computer vision, and mobile computing technologies, we propose and implement a wearable virtual usher for cognitive indoor navigation based on egocentric visual perception. A novel computational framework of cognitive wayfinding in an indoor environment is proposed, which contains a context model, a route model, and a process model. A hierarchical structure is proposed to represent the cognitive context knowledge of indoor scenes. Given a start position and a destination, a Bayesian network model is proposed to represent the navigation route derived from the context model. A novel dynamic Bayesian network (DBN) model is proposed to accommodate the dynamic process of navigation based on real-time first-person-view visual input, which involves multiple asynchronous temporal dependencies. To adapt to large variations in travel time through trip segments, we propose an online adaptation algorithm for the DBN model, leading to a self-adaptive DBN. A prototype system is built and tested for technical performance and user experience. The quantitative evaluation shows that our method achieves over 13% improvement in accuracy as compared to baseline approaches based on hidden Markov model. In the user study, our system guides the participants to their destinations, emulating a human usher in multiple aspects. Liyuan Li, Qianli Xu, Vijay Chandrasekhar 0001, Joo-Hwee Lim, Cheston Tan, Michal Mukawa |
IEEE Trans. Cybern. | 2 |
| 2015 | Stress level detection using double-layer subband filter
Tin Lay Nwe, Qianli Xu, Cuntai Guan, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2015 | Cluster-Based Analysis for Personalized Stress Evaluation Using Physiological SignalsabstractTechnology development in wearable sensors and biosignal processing has made it possible to detect human stress from the physiological features. However, the intersubject difference in stress responses presents a major challenge for reliable and accurate stress estimation. This research proposes a novel cluster-based analysis method to measure perceived stress using physiological signals, which accounts for the intersubject differences. The physiological data are collected when human subjects undergo a series of task-rest cycles, incurring varying levels of stress that is indicated by an index of the State Trait Anxiety Inventory. Next, a quantitative measurement of stress is developed by analyzing the physiological features in two steps: 1) a k -means clustering process to divide subjects into different categories (clusters), and 2) cluster-wise stress evaluation using the general regression neural network. Experimental results show a significant improvement in evaluation accuracy as compared to traditional methods without clustering. The proposed method is useful in developing intelligent, personalized products for human stress management. Qianli Xu, Tin Lay Nwe, Cuntai Guan |
IEEE J. Biomed. Health Informatics | 1 |
| 2014 | A wearable virtual guide for context-aware cognitive indoor navigationabstractIn this paper, we explore a new way to provide context-aware assistance for indoor navigation using a wearable vision system. We investigate how to represent the cognitive knowledge of wayfinding based on first-person-view videos in real-time and how to provide context-aware navigation instructions in a human-like manner. Inspired by the human cognitive process of wayfinding, we propose a novel cognitive model that represents visual concepts as a hierarchical structure. It facilitates efficient and robust localization based on cognitive visual concepts. Next, we design a prototype system that provides intelligent context-aware assistance based on the cognitive indoor navigation knowledge model. We conducted field tests and evaluated the system's efficacy by benchmarking it against traditional 2D maps and human guidance. The results show that context-awareness built on cognitive visual perception enables the system to emulate the efficacy of a human guide, leading to positive user experience. Qianli Xu, Liyuan Li, Joo-Hwee Lim, Cheston Tan, Michal Mukawa, Gang S. Wang |
Mobile HCI | 1 |
| 2013 | Designing engagement-aware agents for multiparty conversationsabstractRecognizing users' engagement state and intentions is a pressing task for computational agents to facilitate fluid conversations in situated interactions. We investigate how to quantitatively evaluate high-level user engagement and intentions based on low-level visual cues, and how to design engagement-aware behaviors for the conversational agents to behave in a sociable manner. Drawing on machine learning techniques, we propose two computational models to quantify users' attention saliency and engagement intentions. Their performances are validated by a close match between the predicted values and the ground truth annotation data. Next, we design a novel engagement-aware behavior model for the agent to adjust its direction of attention and manage the conversational floor based on the estimated users' engagement. In a user study, we evaluated the agent's behaviors in a multiparty dialog scenario. The results show that the agent's engagement-aware behaviors significantly improved the effectiveness of communication and positively affected users' experience. Qianli Xu, Liyuan Li, Gang S. Wang |
CHI | 1 |
| 2012 | Effect of scenario media on human-robot interaction evaluationabstractDifferent media used to present the human-robot interaction (HRI) scenarios may affect users' perception of a robot in the user studies. We investigated how different scenario media (text, video, and live interaction) might influence user evaluation of social robots based on a controlled experiment. We found that multiple aspects of user acceptance were influenced by the scenario media. Moreover, more design problems and redesign proposals were elicited when users were exposed to media with higher fidelity. The results led to useful insights into choosing scenario media in HRI evaluation. Qianli Xu, Jamie Ng, Yian Ling Cheong, Odelia Yiling Tan, Ji Bin Wong, Benedict Tay Tiong Chee, Taezoon Park |
HRI | 1 |
| 2012 | Effect of scenario media on elder adults' evaluation of human-robot interactionabstractWhen evaluating user attitudes toward social robots in human-robot interactions (HRIs), one should exploit the rich contextual information in the HRI. Such information is usually represented as HRI scenarios, which can be conveyed using different media. We investigated how different media of scenarios might influence elder adults' evaluation of social robots. Three media (text, video, and live interaction) were used to elicit user acceptance and feedback, where the content of scenarios was kept similar. We found that multiple aspects of user acceptance were influenced by the scenario media. Moreover, more design problems were elicited when users were exposed to media with higher fidelity. The results shed light on the selection of scenario media in HRI evaluation. Qianli Xu, Jamie Ng, Yian Ling Cheong, Odelia Yiling Tan, Ji Bin Wong, Benedict Tay Tiong Chee, Taezoon Park |
RO-MAN | 1 |
| 2012 | User Experience Modeling and Simulation for Product Ecosystem Design Based on Fuzzy Reasoning Petri NetsabstractProduct ecosystem design entails complex user experience (UX) that involves interactions among multiple users, products, and the ambience. This paper aims to capture causal relationships between UX and design elements and in turn to provide decision support to product ecosystem analysis. A fuzzy reasoning Petri net is developed to deal with the uncertainty, complexity, and dynamics associated with UX modeling. Reasoning of diverse constructs of UX is embedded in the fuzzy production rules that are derived from self-report UX data based on rough set mining. A fuzzy reasoning algorithm is implemented to perform parallel inference by multicriteria rules and to simulate most likely UX under different ambient factors. A case study of subway station UX design demonstrates the potential of product ecosystem FRPN formulation. Feng Zhou 0003, Roger Jianxin Jiao, Qianli Xu, Koji Takahashi |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2009 | Coordinating product, process, and supply chain decisions: A constraint satisfaction approach
Roger Jianxin Jiao, Qianli Xu, Zhang Wu, Ngai-Kheong Ng |
Eng. Appl. Artif. Intell. | 2 |