EDBT 2026 Demo / reviewers in the wild / expert
Liyuan Li
dblp:14/333
· DBLP profile ↗
72ranked-venue papers
18as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 29 · 10 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-authorSystems, architecture and hardware · 4 · 3 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Multi-Camera Inspection System for AircraftabstractIn this paper, we present the development of an automated visual inspection system designed to detect defects on the upper surface of an aircraft airframe. Specifically, the system employs a multi-camera PTZ (Pan-Tilt-Zoom) set-up to capture and process images from designated regions. Custom developed software manages path planning and camera localization, while a hybrid-AI framework is integrated to identify various defect types, including missing and damaged components. The demonstration highlights the system’s detection capabilities and prototype functionalities using a large aircraft model, supported by a user interface to monitor progress and visualize results. To help validate this work, performance evaluations were conducted using selected multimodal and object detection models. Mark D. Rice, Kelvin Wei Lim, Qing Yu Hoo, Liyuan Li, Lee Jue Ying, Jacky Jie Wei Tan, Lai Xing Ng, Jamie Ng |
AAAI | 5 |
| 2026 | KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Gang Liao, Hongsen Qin, Alicia Golden, Michael Kuchnik, Yavuz Yetim, Ruichao Xiao, Jia Jiunn Ang, Chunli Fu, Yihan He, Samuel Hsia, Zewei Jiang, Roman Levenstein, Dianshi Li, Liyuan Li, Ajit Mathews, Varna Puvvada, Feng Shi 0001, Nathan Yan, Xiayu Yu, Uladzimir Pashkevich, Matt Steiner, Carole-Jean Wu, Gaoxiang Liu |
ISCA | 15 |
| 2025 | Digital Low-Cost FPGA Implementation of Two-Coupled and Grid-Based Network of 2D Artificial Cochlea Using the Hopf Resonator ApproachabstractThe Cochlea, a spiral-shaped structure in the inner ear, plays a crucial role in the process of hearing by converting sound waves into electrical signals that the brain can interpret. This study introduces a cost-effective adaptation of 2D artificial Cochlea mathematical modeling using a planar approximation technique. The main novelty and contribution of our work is a method employs surface-based functions and is known as the Surface-Based Approximation Model of Cochlea (SBAMoC). By simplifying complex multiplication processes in nonlinear components, the SBAMoC reduces costs and enhances efficiency, making it suitable for FPGA implementation with minimal hardware requirements. The proposed model is evaluated in scenarios involving two-coupled oscillations and grid-based cochlear networks to better understand its performance. Through hardware synthesis on a Virtex-II board, the SBAMoC demonstrates improved efficiency and reduced computational expenses compared to the original model, achieving faster speeds and greater cost-effectiveness. In practical tests, the SBAMoC exhibits higher operational speeds and increased scalability, outperforming the original model by replicating accurate cochlear behaviors with minimal deviations. Specifically, the single SBAMoC implementation in our model achieves a speed boost of approximately 1.333 times compared to the original model (381.292 MHz vs. 286.029 MHz) and supports a greater number of fitted SBAMoCs (75 vs. 35), showcasing its superior efficiency and performance enhancements. In case of real-world applications, it can be considered for the development of more efficient and cost-effective cochlear implants, leading to improved hearing restoration solutions for individuals with hearing impairments. Also, the findings from this study could also be leveraged to enhance the design and implementation of signal processing systems in various audio and communication devices, paving the way for advanced audio processing technologies with increased efficiency and reduced hardware costs. Songjie Xiang, Liyuan Li, Yisu Ge, Xiaoyun Gao, Mohammad Sharif Daoud, Abdulilah M. Mayet, Yanling Chu, Yideng Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | RD-Diff: RLTransformer -Based Diffusion Model with Diversity-Inducing Modulator for Human Motion Prediction
Haosong Zhang 0001, Mei Chee Leong, Liyuan Li, Weisi Lin |
ACCV (1) | 3 |
| 2024 | PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action RecognitionabstractRecent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However, due to a large gap between vision and text, they might not be able to sufficiently utilize the benefits of cross-modality information. In the field of human action recognition, the additional pose modality may bridge the gap between vision and text to improve the effective-ness of cross-modality learning. In this paper, we propose a novel framework, called Pose-enhanced Vision-Language (Pe VL) model, to adapt the VL model with pose modality to learn effective knowledge offine-grained human actions. Our PeVL model includes two novel components: an Un-symmetrical Cross-Modality Refinement (UCMR) block and a Semantic-Guided Multi-level Contrastive (SGMC) mod-ule. The UCMR block includes Pose-guided Visual Refine-ment (P2V-R) and Visual-enriched Pose Refinement (V2P-R) for effective cross-modality learning. The SGMC module includes Multi-level Contrastive Associations of vision-text and pose-text at both action and sub-action levels, and a Semantic-Guided Loss, enabling effective contrastive learning with text. Built upon a pre-trained VLfoundation model, our model integrates trainable adapters and can be trained end-to-end. Our novel PeVL design over VL foundation model yields remarkable performance gains on four fine-grained human action recognition datasets, achieving a new SOTA with a significantly small number of FLOPs for low-cost re-training.1 Haosong Zhang 0001, Mei Chee Leong, Liyuan Li, Weisi Lin |
CVPR | 3 |
| 2024 | PGVT: Pose-Guided Video Transformer for Fine-Grained Action RecognitionabstractBased on recent advancements in transformer-based video models and multi-modal joint learning, we propose a novel model, named Pose-Guided Video Transformer (PGVT), to incorporate sparse high-level body joints locations and dense low-level visual pixels for effective learning and accurate recognition of human actions. PGVT leverages the pre-trained image models by freezing their parameters and introducing trainable adapters to effectively integrate two input modalities, i.e., human poses and video frames, to learn a pose-focused spatiotemporal representation of human actions. We design two novel core modules, i.e., Pose Temporal Attention and Pose-Video Spatial Attention, to facilitate interaction between body joint locations and uniform video tokens, enriching each modality with contextualized information from the other. We evaluate PGVT model on four action recognition datasets: Diving48, Gym99, and Gym288 for fine-grained action recognition, and Kinetics400 for coarse-grained action recognition. Our model achieves new SOTA performance on the three fine-grained human action recognition datasets and comparable performance on Kinetics400 with a small number of tunable parameters compared with SOTA methods. Various ablation studies are performed which verify the benefits of our new designs. Haosong Zhang 0001, Mei Chee Leong, Liyuan Li, Weisi Lin |
WACV | 3 |
| 2024 | Fast Thermal Infrared Image Restoration Method Based on On-Orbit Invariant Modulation Transfer FunctionabstractAlthough the thermal infrared remote sensing camera plays a pivotal role in Earth observation, and impacts the target detection, surface temperature inversion, and subsequent space missions significantly, the imaging quality of the camera is constrained by its optics, image sensors, and electronics during on-orbit operation. At the same time, the traditional blind recovery algorithms, which require extensive time for estimating intricate blur kernels, encounter challenges due to varying atmospheric conditions and other factors leading to dissimilar blur kernels across different observation scenes. In this context, this article introduces a rapid image recovery algorithm rooted in the concept of the invariant modulation transfer function (IMTF) specific to on-orbit cameras. The IMTF model remains stable and impervious to influences stemming from ground targets, atmospheric conditions, and orbital or environmental fluctuations, contingent upon the camera’s inherent characteristics. The extraction of the IMTF involves subjecting the transfer function’s region to a modified edge methodology, followed by image recovery through a hyper-Laplacian prior inverse convolution approach. The resolution of the inverse problem is achieved by employing an alternating minimization scheme. This method addresses the mitigation of imaging artifacts originating from the camera’s limitations. Comparative analysis against the state-of-the-art image recovery techniques establishes the competitiveness of the method proposed in this article, both in terms of recovery efficacy and operational efficiency. Substantiating this, experimental validation using in-orbit thermal infrared remote sensing images reveals a notable improvement in the average gradient (AG) (by a factor of 3.2), edge intensity (EI) (by a factor of 2.5), and modulation transfer function (by a factor of 1.3) of the restored images. Consequently, this approach introduces a novel perspective for enhancing the restoration of in-orbit remote sensing images. Lintong Qi, Rongguo Zhang, Zhuoyue Hu, Liyuan Li, Qiyao Wang, Xinyue Ni |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Slow and Weak Attractor Computation Embedded in Fast and Strong E-I Balanced Neural DynamicsabstractAttractor networks require neuronal connections to be highly structured in order to maintain attractor states that represent information, while excitation and inhibition balanced networks (E-INNs) require neuronal connections to be random and sparse to generate irregular neuronal firings. Despite being regarded as canonical models of neural circuits, both types of networks are usually studied in isolation, and it remains unclear how they coexist in the brain, given their very different structural demands. In this study, we investigate the compatibility of continuous attractor neural networks (CANNs) and E-INNs. In line with recent experimental data, we find that a neural circuit can exhibit both the traits of CANNs and E-INNs if the neuronal synapses consist of two sets: one set is strong and fast for irregular firing, and the other set is weak and slow for attractor dynamics. Our results from simulations and theoretical analysis reveal that the network also exhibits enhanced performance compared to the case of using only one set of synapses, with accelerated convergence of attractor states and retained E-I balanced condition for localized input. We also apply the network model to solve a real-world tracking problem and demonstrate that it can track fast-moving objects well. We hope that this study provides insight into how structured neural computations are realized by irregular firings of neurons. Xiaohan Lin, Liyuan Li, Boxin Shi, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001 |
NeurIPS | 2 |
| 2023 | In-Orbit Geometric Calibration for Long-Linear-Array and Wide-Swath Whisk-Broom TIS of SDGSAT-1abstractBecause of the imaging mechanism complexity of long-linear-array and wide-swath whisk-broom thermal infrared spectrometer (TIS) of the first Sustainable Development Goals Satellite (SDGSAT-1), how to achieve a high geometric positioning accuracy (GPA) becomes the core factor in subsequent geometric quantitative applications. Here, in this article, a three-step in-orbit geometric calibration (GC) strategy comprising the estimations of exterior orientation parameters (EOPs), interior orientation parameters (IOPs), and scanning compensation parameters (SCPs) is proposed to correct the geo-location displacements for whisk-broom TIS. First, in accordance with the optical-mechanical structure and pinhole imaging theory, we establish the rigorous geometric positioning model (RGPM) of TIS and analyze the error resources term-by-term along the error propagation link elaborately. Second, the corresponding rigorous geometric calibration model (RGCM) is constructed in detail based on the 2-D look-angle model and the generalized bias correction matrix. Especially for eliminating the systematic nonlinear errors in the scanning direction, a fifth-degree polynomial is put forward to be employed to fit and compensate for the angular measurement errors of the scanning mirror. Finally, a three-step estimation method is presented to estimate the calibration parameters with ground control points (GCPs). Experimental results based on the spatial references of Landsat 8 panchromatic images and version 2 of advanced spaceborne thermal emission and reflection radiometer (ASTER) global digital elevation model (GDEM2) show that the GPA of the proposed method in along-track and cross-track directions can be better than 1.0 pixels for all three bands, which makes a great sense for associated geometric measurements. Liyuan Li, Lixing Zhao, Jingjie Jiao, Linyi Jiang, Lan Yang 0001, Shengli Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Infrared Small Target Detection Based on Weighted Three-Layer Window Local ContrastabstractThe performance of small target detection restricts the development of the infrared search and track (IRST) system. Against the complicated background clutter of the infrared (IR) image, the small targets are difficult to separate from a noisy background. Aiming at solving the problem of residual background clutter in the local contrast method, a weighted three-layer window local contrast method (WTLLCM) is proposed in this letter. First, the images are filtered by a layered gradient kernel to enhance the contrast between targets and background. Then, a three-layer window is utilized to calculate the local contrast of the filtered images. Next, it is worth mentioning that a simple target aggregation strategy is considered to preserve the integrity of the target. Especially, a new region intensity level (NRIL) algorithm is proposed to weigh the local contrast map to further suppress the background. Finally, the targets are detected by adaptive threshold segmentation. Compared with state-of-the-art small targets detection baseline algorithms based on local contrast, extensive experimental results demonstrate the superiority of the proposed method, especially in complex backgrounds. And instead of utilizing multi-scale windows, multi-scale targets detection is accomplished by using a single-scale window to reduce the calculation of the method. Huixin Cui, Liyuan Li, Xin Liu 0084, Xiaofeng Su |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Moving Dim and Small Target Detection in Multiframe Infrared Sequence With Low SCR Based on Temporal Profile SimilarityabstractResearch about infrared dim and small target detection (DSTD) is concentrated on single-frame algorithms, which are limited by the contrast between target and background and face with the problems of low detection probability, high false alarm rate, and lack of robustness in low signal-to-clutter ratio (SCR) and strong noise environment. Studying the use of multiframe sequences adequately is necessary. In order to effectively utilize the temporal and local spatial information of infrared sequences, we propose a similarity model for pixel temporal profile (TP). Different from current TP detection methods, we study waveform similarity calculation for detection, and local time-shift characteristics to eliminate false alarms. First, fast Fourier transform (FFT) and KL divergence are applied to calculate the similarity of TP and reference waveform, and errors due to time offsets can be avoided through the frequency domain; second, the peak ratio of the FFT is applied to calculate the time shift of the neighboring pixels relative to the center pixel; third, maximum suppression strategy is used to reduce false alarms. Experiments show that the model and algorithms proposed in this letter have excellent performance in target enhancement and background suppression, and have higher performance than other methods in receiver operator characteristic curve (ROC). Xin Liu 0084, Liyuan Li, Liqi Liu, Xiaofeng Su |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | On-Orbit Spatial Quality Evaluation of SDGSAT-1 Thermal Infrared SpectrometerabstractThe Sustainable Development Science Satellite (SDGSAT-1) was successfully launched on 5th November 2021. It is the world’s first scientific satellite dedicated to serving the United Nations 2030 Agenda for Sustainable Development. To evaluate the on-orbit image quality of its thermal infrared spectrometer (TIS), we propose an improved edge slope calculation method. Learning from past academic research, we found that the traditional edge slop (ES) calculation methods are limited by various factors. For example, it demands highly on geographical aspects and surface temperature uniformity. In this paper, we interpolate the narrow sea-land boundary region. The edge signal-to-noise ratio and dip angle are calculated by the selected region to determine whether it satisfies the conditions of ES calculation. The Otsu algorithm is used to find the optimal threshold for the sea-land division, and the sub-pixel edge positions are extracted by combining with the canny operator. In order to achieve higher precision, we use the modified Fermi function to fit the extracted edge spread function (ESF). During the on-orbit testing of the SDGSAT-1/TIS, we performed ES calculations for a large number of test areas with a flight direction of 0.5 and a scan direction of 0.45. Lintong Qi, Liyuan Li, Xinyue Ni, Xiaoxuan Zhou |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A Multi-Task Framework for Infrared Small Target Detection and SegmentationabstractDue to the complicated background and noise of infrared images, infrared small target detection is one of the most difficult problems in the field of computer vision. In most existing studies, semantic segmentation methods are typically used to achieve better results. The centroid of each target is calculated from the segmentation map as the detection result. In contrast, we propose a novel end-to-end framework for infrared small target detection and segmentation in this paper. First, with the use of UNet as the backbone to maintain resolution and semantic information, our model can achieve a higher detection accuracy than other state-of-the-art methods by attaching a simple anchor-free head. Then, a pyramid pool module is used to further extract features and improve the precision of target segmentation. Next, we use semantic segmentation tasks that pay more attention to pixel-level features to assist in the training process of object detection, which increases the average precision and allows the model to detect some targets that were previously not detectable. Furthermore, we develop a multi-task framework for infrared small target detection and segmentation. Our multi-task learning model reduces complexity by nearly half and speeds up inference by nearly twice compared to the composite single-task model, while maintaining accuracy. The code and models are publicly available at https://github.com/Chenastron/MTUNet. Liyuan Li, Xin Liu 0084, Xiaofeng Su |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Combined Stripe Noise Removal and Deblurring Recovering Method for Thermal Infrared Remote Sensing ImagesabstractRemote sensing images (RSIs) have been applied to many fields such as environmental monitoring, urban planning, and military defense. All-day thermal infrared imaging observation system can finely portray the trajectory of human activities and provide data support for United Nations sustainable development planning. However, in the process of RSIs, image quality degradation containing image blurring is caused by optical system aberration, satellite platform vibration, imaging system out-of-focus, and atmospheric turbulence. Moreover, stripe noise in the image is produced by the non-uniformity of infrared sensors. Image deblurring and destriping are classical tasks in RSIs, but the above two problems are discussed separately in almost all research. i.e., denoising after adding stripe noise on clear images or deblurring under the assumption that only random noise exists. In this paper, stripe noise and blur are jointly removed from on-orbit RSIs acquired by the thermal infrared spectrometer on the SDGSAT-1 satellite, using a method based on stripe component residuals and gradient property. According to the experimental results, the performance of the proposed method is greatly improved compared with processing these two tasks separately, which can provide a valuable reference for the study of high-resolution thermal infrared RSIs recovery. Xiaoxuan Zhou, Liyuan Li, Tingliang Hu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | TAILOR: Teaching with Active and Incremental Learning for Object RegistrationabstractWhen deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions. Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
AAAI | 8 |
| 2021 | Enhancing Multi-Step Action Prediction for Active Object DetectionabstractActive vision for robots is one promising solution to open world visual detection problems. A fundamental issue is view planning, i.e., predicting next best views to capture images of interest to reduce uncertainty. While multi-step action in a reinforcement learning (RL) setup can boost the efficiency of view planning, existing methods suffer from unstable detection outcome when the Q-values of multiple branches of action advantages (i.e., action range and action type) are combined naively. To tackle this issue, we propose a novel mechanism to disentangle action range from action type through a two-stage training strategy on a deep Q-network. It combines well-crafted loss functions with respect to action range and action type to enforce separated training of these two branches. We evaluate our method on two public datasets and show that it facilitates substantial gain in view planning efficiency, while enhancing detection accuracy. Fen Fang, Qianli Xu, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 4 |
| 2021 | Joint Learning on the Hierarchy Representation for Fine-Grained Human Action RecognitionabstractFine-grained human action recognition is a core research topic in computer vision. Inspired by the recently proposed hierarchy representation of fine-grained actions in FineGym and SlowFast network for action recognition, we propose a novel multi-task network which exploits the FineGym hierarchy representation to achieve effective joint learning and prediction for fine-grained human action recognition. The multi-task network consists of three pathways of SlowOnly networks with gradually increased frame rates for events, sets and elements of fine-grained actions, followed by our proposed integration layers for joint learning and prediction. It is a two-stage approach, where it first learns deep feature representation at each hierarchical level, and is followed by feature encoding and fusion for multi-task learning. Our empirical results on the FineGym dataset achieve a new state-of-the-art performance, with 91.80% Top-1 accuracy and 88.46% mean accuracy for element actions, which are 3.40% and 7.26% higher than the previous best results. Mei Chee Leong, Hui Li Tan, Haosong Zhang 0001, Liyuan Li, Feng Lin 0002, Joo-Hwee Lim |
ICIP | 4 |
| 2021 | Towards Efficient Multiview Object Detection with Adaptive Action PredictionabstractActive vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection. We formulate the multi-object detection task as an active multiview object detection problem given the initial location of the objects. Next, we propose a novel adaptive action prediction (A2P) method built on a deep Q-learning network with a dueling architecture. The A2P method is able to perform view planning based on visual information of multiple objects; and adjust action ranges according to the task status. Evaluated on the AVD dataset, A2P leads to 21.9% increase in detection accuracy in unfamiliar environments, while improving efficiency by 22.7%. On the T-LESS dataset, multi-object detection boosts efficiency by more than 30% while achieving equivalent detection accuracy. Qianli Xu, Fen Fang, Nicolas Gauthier, Wenyu Liang, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
ICRA | 6 |
| 2021 | Modeling-Assisted InSAR Phase-Unwrapping Method for Mapping Mine SubsidenceabstractCompared with traditional measurement technologies, synthetic aperture radar interferometry (InSAR) has unique advantages in monitoring ground subsidence due to underground mining. However, when the subsidence gradient of the subsidence trough exceeds the maximum measurable gradient of InSAR technology, the interference fringes will be too dense, causing phase aliasing. As a result, it is impossible to obtain correct phase-unwrapping result. The main objectives of this letter are two folded. First is to develop an unwrapping strategy to deal with the unwrapping problem caused by large subsidence gradient at the mine subsidence trough. The main idea of this strategy is to estimate most of the subsidence phase by multiple model inversions based on iterative approach. Then, the model phases from multiple models are combined with the final unwrapped residual phase. Another objective of this letter is to evaluate the feasibility of the three common deformation models, i.e., Mogi, probability integral method (PIM), and Okada, in solving the phase-unwrapping problem. Their advantages and disadvantages are outlined. Both the simulated data and real data are used for this experiment. The result shows that the problem of large subsidence gradient in the differential interferometric synthetic aperture radar (DInSAR) results can be solved by multiple model inversion. Among the three models, the use of Okada model seems to provide slightly more accurate result for solving the large-scale subsidence in the mining area than the other two models with the proposed strategy. Yiwei Dai, Alex Hayman Ng, Liyuan Li, Linlin Ge, Tingye Tao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Single-Image Dehazing via Compositional Adversarial NetworkabstractSingle-image dehazing has been an important topic given the commonly occurred image degradation caused by adverse atmosphere aerosols. The key to haze removal relies on an accurate estimation of global air-light and the transmission map. Most existing methods estimate these two parameters using separate pipelines which reduces the efficiency and accumulates errors, thus leading to a suboptimal approximation, hurting the model interpretability, and degrading the performance. To address these issues, this article introduces a novel generative adversarial network (GAN) for single-image dehazing. The network consists of a novel compositional generator and a novel deeply supervised discriminator. The compositional generator is a densely connected network, which combines fine-scale and coarse-scale information. Benefiting from the new generator, our method can directly learn the physical parameters from data and recover clean images from hazy ones in an end-to-end manner. The proposed discriminator is deeply supervised, which enforces that the output of the generator to look similar to the clean images from low-level details to high-level structures. To the best of our knowledge, this is the first end-to-end generative adversarial model for image dehazing, which simultaneously outputs clean images, transmission maps, and air-lights. Extensive experiments show that our method remarkably outperforms the state-of-the-art methods. Furthermore, to facilitate future research, we create the HazeCOCO dataset which is currently the largest dataset for single-image dehazing. Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Zhao Kang 0001, Shijian Lu, Zhiwen Fang, Liyuan Li, Joo-Hwee Lim |
IEEE Trans. Cybern. | 8 |
| 2021 | Lifelog Image Retrieval Based on Semantic Relevance MappingabstractLifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method. Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Task-Oriented Multi-Modal Question Answering For Collaborative ApplicationsabstractCobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task. Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 4 |
| 2020 | Active Image Sampling on Canonical Views for Novel Object DetectionabstractTo alleviate the costly data annotation problem in deep learning-based object detection, we leverage the canonical view model for active sample selection to improve the effectiveness of learning. Inspired by the view-approximation model, we hypothesize that visual features learned from canonical views denote better representations of objects, thus boosting the effectiveness of object learning. We validate the hypothesis empirically in the context of robot learning for novel object detection. Based on this, we propose a novel on-line viewpoint exploration (OLIVE) method that (1) defines goodness-of-view by combining informativeness of visual features and consistency of model-based object detection, and (2) systematically explores and selects viewpoints to boost learning efficiency. Furthermore, we train a legacy Faster R-CNN model with a data augmentation method while leveraging data samples generated by the OLIVE pipeline. We test our method on the T-LESS dataset and show that the proposed method outperforms competitive benchmarking methods, especially when the samples are few. Qianli Xu, Fen Fang, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 4 |
| 2020 | 6D Pose Estimation with Correlation Fusionabstract6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illumination, so it is important to complement them with depth information. However, existing methods using RGB-D data cannot adequately exploit consistent and complementary information between RGB and depth modalities. In this paper, we present a novel method to effectively consider the correlation within and across both modalities with attention mechanism to learn discriminative and compact multi-modal features. Then, effective fusion strategies for intra- and inter-correlation modules are explored to ensure efficient information flow between RGB and depth. To our best knowledge, this is the first work to explore effective intra- and inter-modality fusion in 6D pose estimation. The experimental results show that our method can achieve the state-of-the-art performance on LineMOD and YCB-Video dataset. We also demonstrate that the proposed method can benefit a real-world robot grasping task by providing accurate object pose estimation. Hongyuan Zhu 0002, Ying Sun 0001, Cihan Acar, Yan Wu 0002, Liyuan Li, Cheston Tan, Joo-Hwee Lim |
ICPR | 7 |
| 2020 | Detecting Objects with High Object Region PercentageabstractObject shape is a subtle but important factor for object detection. It has been observed that the object-region-percentage (ORP) can be utilized to improve detection accuracy for elongated objects, which have much lower ORPs than other types of objects. In this paper, we propose an approach to improve the detection performance for objects with high ORPs. Our method consists of three steps. First, we adjust the ground truth bounding boxes of high-ORP objects to an optimal range. Second, we train an object detector, Faster R-CNN, based on adjusted bounding boxes to achieve high recall. Finally, we train a DCNN to learn the adjustment ratios towards four directions and adjust detected bounding boxes of objects to get better localization for higher precision. We evaluate the effectiveness of our method on 12 high-ORP objects in COCO and 8 objects in a proprietary gearbox dataset. The experimental results show that our method can achieve state-of-the-art performance on these objects while costing less resources in training and inference stages. Fen Fang, Qianli Xu, Liyuan Li, Joo-Hwee Lim |
ICPR | 3 |
| 2020 | The 5G communication technology-oriented intelligent building system planning and design
Liyuan Li |
Comput. Commun. | 2 |
| 2020 | A novel hybrid approach for crack detection
Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
Pattern Recognit. | 2 |
| 2020 | Combining Faster R-CNN and Model-Driven Clustering for Elongated Object DetectionabstractWhile analyzing the performance of state-of-the-art R-CNN based generic object detectors, we find that the detection performance for objects with low object-region-percentages (ORPs) of the bounding boxes are much lower than the overall average. Elongated objects are examples. To address the problem of low ORPs for elongated object detection, we propose a hybrid approach which employs a Faster R-CNN to achieve robust detections of object parts, and a novel model-driven clustering algorithm to group the related partial detections and suppress false detections. First, we train a Faster R-CNN with partial region proposals of suitable and stable ORPs. Next, we introduce a deep CNN (DCNN) for orientation classification on the partial detections. Then, on the outputs of the Faster R-CNN and DCNN, the algorithm of adaptive model-driven clustering first initializes a model of an elongated object with a data-driven process on local partial detections, and refines the model iteratively by model-driven clustering and data-driven model updating. By exploiting Faster R-CNN to produce robust partial detections and model-driven clustering to form a global representation, our method is able to generate a tight oriented bounding box for elongated object detection. We evaluate the effectiveness of our approach on two typical elongated objects in the COCO dataset, and other typical elongated objects, including rigid objects (pens, screwdrivers and wrenches) and non-rigid objects (cracks). Experimental results show that, compared with the state-of-the-art approaches, our method achieves a large margin of improvements for both detection and localization of elongated objects in images. Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
IEEE Trans. Image Process. | 2 |
| 2019 | Singe Image Rain Removal with Unpaired Information: A Differentiable Programming PerspectiveabstractSingle image rain-streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. Previous works solve this problem using various hand-designed priors or by explicitly mapping synthetic rain to paired clean image in a supervised way. In practice, however, the pre-defined priors are easily violated and the paired training data are hard to collect. To overcome these limitations, in this work, we propose RainRemoval-GAN (RRGAN), the first end-to-end adversarial model that generates realistic rain-free images using only unpaired supervision. Our approach alleviates the paired training constraints by introducing a physical-model which explicitly learns a recovered images and corresponding rain-streaks from the differentiable programming perspective. The proposed network consists of a novel multiscale attention memory generator and a novel multiscale deeply supervised discriminator. The multiscale attention memory generator uses a memory with attention mechanism to capture the latent rain streaks context at different stages to recover the clean images. The deeply supervised multiscale discriminator imposes constraints at the recovered output in terms of local details and global appearance to the clean image set. Together with the learned rainstreaks, a reconstruction constraint is employed to ensure the appearance consistent with the input image. Experimental results on public benchmark demonstrates our promising performance compared with nine state-of-the-art methods in terms of PSNR, SSIM, visual qualities and running time. Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Songfan Yang, Vijay Chanderasekh, Liyuan Li, Joo-Hwee Lim |
AAAI | 6 |
| 2019 | Towards Real-Time Crack Detection Using a Deep Neural Network With a Bayesian Fusion AlgorithmabstractSurface cracks can represent very small and thin objects in images. With irregular shapes and sizes, and non-fixed texture patterns, the detection of cracks can be a challenging problem in computer vision. Prior work has been undertaken on detecting cracks for images using a sliding window mode. However, such methods can be time consuming, and result in high false alarms. To help address this problem, a new crack detection and segmentation method is proposed in this paper. Specifically, our method includes three main features: (1) a Faster R-CNN model to detect crack patches in images; (2) the use of a Bayesian fusion algorithm to suppress false alarms based on detected patch orientation; and (3) image processing functions to obtain final segmentation masks, such as for Gaussian blur, erosion, etc. Experimental results show that our method can achieve high detection accuracy on sampled images in real-time. Fen Fang, Liyuan Li, Mark D. Rice, Joo-Hwee Lim |
ICIP | 2 |
| 2019 | An Adaptive Fitting Approach for the Visual Detection and Counting of Small Circular Objects in Manufacturing ApplicationsabstractDetecting, localizing and counting small circular objects in machine parts is an important task in many applications for manufacturing. Existing methods of circle detection face difficulties due to the high-curvature and limited edge points of circles. As a result, in this paper we propose a novel two-stage circle detection method, which integrates bottom-up coarse detection and top-down circle fitting. First, a circle detector combining low-level feature descriptors and a linear SVM is developed. This is used to scan an input image in a sliding window mode to detect small circles with coarse estimates of locations and scales. Next, a hierarchical Bayesian model performs a top-down adaptive circle fitting, with the ability to achieve a maximum a posteriori probability to fit circles to local image features. The evaluation of our approach with manufacturing images has demonstrated to be efficient in detecting small circles in machine parts. Liyuan Li, Fen Fang, Mark D. Rice, Jamie Ng, Wei Xiong 0001, Joo-Hwee Lim |
ICIP | 2 |
| 2018 | Image-based Parking Place Identification for Regulating Shared Bicycle ParkingabstractWe propose a novel method and system to prevent indiscriminate parking of dockless shared bicycles using location-based geo-fencing and image-based parking place identification. The geo-fencing is used to define the approximate regions for different types of bicycle parking regulations. The parking place identification uses a method based on deep Convolutional Neural Network (DCNN) to automatically identify designated bicycle parking places from photos captured by the cyclist using a mobile phone. Combining these two modalities, the parking of shared bicycles can be restricted in designated zones in various environments. Experiments are conducted using photos taken from the designated parking places with different parking indications at various locations. We evaluate the performance of the image-based parking place identification and use heatmaps to analyze potential features that are exploit by the DCNN models. The method achieves high performance on the testing dataset; and the features used for parking place identification are largely consistent with human perceptions. Shudong Xie, Qianli Xu, Fen Fang, Liyuan Li |
ICARCV | 5 |
| 2018 | DehazeGAN: When Image Dehazing Meets Differential ProgrammingabstractSingle image dehazing has been a classic topic in computer vision for years. Motivated by the atmospheric scattering model, the key to satisfactory single image dehazing relies on an estimation of two physical parameters, i.e., the global atmospheric light and the transmission coefficient. Most existing methods employ a two-step pipeline to estimate these two parameters with heuristics which accumulate errors and compromise dehazing quality. Inspired by differentiable programming, we re-formulate the atmospheric scattering model into a novel generative adversarial network (DehazeGAN). Such a reformulation and adversarial learning allow the two parameters to be learned simultaneously and automatically from data by optimizing the final dehazing performance so that clean images with faithful color and structures are directly produced. Moreover, our reformulation also greatly improves the GAN’s interpretability and quality for single image dehazing. To the best of our knowledge, our method is one of the first works to explore the connection among generative adversarial models, image dehazing, and differentiable programming, which advance the theories and application of these areas. Extensive experiments on synthetic and realistic data show that our method outperforms state-of-the-art methods in terms of PSNR, SSIM, and subjective visual quality. Hongyuan Zhu 0002, Xi Peng 0001, Vijay Chandrasekhar 0001, Liyuan Li, Joo-Hwee Lim |
IJCAI | 4 |
| 2018 | Personalized Serious Games for Cognitive Intervention with Lifelog Visual AnalyticsabstractThis paper presents a novel serious game app and a method to cre- ate and integrate personalized game content based on lifelog visual analytics. The main objective is to extract personalized content from visual lifelogs, integrate it into mobile games, and evaluate the effect of personalization on user experience. First, a suite of visual analysis methods is proposed to extract semantic informa- tion from visual lifelogs and discover the association among the lifelog entities. The outcome is dataset that contains augmented and personal lifelog images. Next, a mobile game app is developed that makes use of the dataset as game content. Finally, an experiment is conducted to evaluate user gameplay behaviors in the wild over three months, where a mixture of generic and personalized game content is deployed. It is observed that user adherence is heightened by personalized game content as compared to generic content. Also observed is a higher enjoyment level in personalized than generic game content. The result provides the first empirical evidence of the effect of personalized games on user adherence and preference for cognitive intervention. This work paves the way for effective cognitive training with user-generated content. Qianli Xu, Vigneshwaran Subbaraju, Chee How Cheong, Aijing Wang, Kathleen Kang, Munirah Bashir, Yanhong Dong, Liyuan Li, Joo-Hwee Lim |
ACM Multimedia | 8 |
| 2018 | A Probabilistic Model of Social Working Memory for Information Retrieval in Social InteractionsabstractSocial working memory (SWM) plays an important role in navigating social interactions. Inspired by studies in psychology, neuroscience, cognitive science, and machine learning, we propose a probabilistic model of SWM to mimic human social intelligence for personal information retrieval (IR) in social interactions. First, we establish a semantic hierarchy as social long-term memory to encode personal information. Next, we propose a semantic Bayesian network as the SWM, which integrates the cognitive functions of accessibility and self-regulation. One subgraphical model implements the accessibility function to learn the social consensus about IR-based on social information concept, clustering, social context, and similarity between persons. Beyond accessibility, one more layer is added to simulate the function of self-regulation to perform the personal adaptation to the consensus based on human personality. Two learning algorithms are proposed to train the probabilistic SWM model on a raw dataset of high uncertainty and incompleteness. One is an efficient learning algorithm of Newton's method, and the other is a genetic algorithm. Systematic evaluations show that the proposed SWM model is able to learn human social intelligence effectively and outperforms the baseline Bayesian cognitive model. Toward real-world applications, we implement our model on Google Glass as a wearable assistant for social interaction. Liyuan Li, Qianli Xu, Tian Gan 0002, Cheston Tan, Joo-Hwee Lim |
IEEE Trans. Cybern. | 1 |
| 2017 | Multi-layer linear model for top-down modulation of visual attention in natural egocentric visionabstractTop-down attention plays an important role in guidance of human attention in real-world scenarios, but less efforts in computational modeing of visual attention has been put on it. Inspired by the mechanisms of top-down attention in human visual perception, we propose a multi-layer linear model of top-down attention to modulate bottom-up saliency maps actively. The first layer is a linear regression model which combines the bottom-up saliency maps on various visual features and objects. A contextual dependent upper layer is introduced to tune the parameters of the lower layer model adaptively. Finally, a mask of selection history is applied to the fused attention map to bias the attention selection towards the task related regions. Efficient learning algorithm with single-pass polynomial complexity is derived. We evaluate our model on a set of natural egocentric videos captured from a wearable glass in real-world environments. Our model outperforms the baseline and state-of-the-art bottom-up saliency models. Keng Teck Ma, Liyuan Li, Peilun Dai, Joo-Hwee Lim, Chengyao Shen, Qi Zhao 0001 |
ICIP | 2 |
| 2017 | The effect of different types of navigation assistance on indoor scene memorabilityabstractWith the rapid growing of wearable computing devices, indoor navigation guidance will become popular in the near future like the GPS-based navigation tools for drivers today. However, how the guided indoor navigation affects human’s memory of a novel environment has not been well studied. In this paper, we investigate route memory with three types of navigation assistance, that is, 2D map, wearable navigation assistant, and human usher. Twenty participants were asked to remember the route while being guided through a novel indoor environment. Our results show that the participants have similar patterns in remembering visual scenes, even using different types of assistance. These findings support previous work on scene memorability and provide the new insight that scene memorability is not affected by the type of navigation guidance. This may indicate that spatial working memory and visual memory are dissociated. We also show that scenes with navigation information are more memorable than scenes without such information. Finally, we provide some evidence that the location of a scene is linked to its memorability. In general, our findings provide valuable information about indoor scene memorability. Michal Mukawa, Cheston Tan, Joo-Hwee Lim, Qianli Xu, Liyuan Li |
Behav. Inf. Technol. | 5 |
| 2017 | A Wearable Virtual Usher for Vision-Based Cognitive Indoor NavigationabstractInspired by progresses in cognitive science, artificial intelligence, computer vision, and mobile computing technologies, we propose and implement a wearable virtual usher for cognitive indoor navigation based on egocentric visual perception. A novel computational framework of cognitive wayfinding in an indoor environment is proposed, which contains a context model, a route model, and a process model. A hierarchical structure is proposed to represent the cognitive context knowledge of indoor scenes. Given a start position and a destination, a Bayesian network model is proposed to represent the navigation route derived from the context model. A novel dynamic Bayesian network (DBN) model is proposed to accommodate the dynamic process of navigation based on real-time first-person-view visual input, which involves multiple asynchronous temporal dependencies. To adapt to large variations in travel time through trip segments, we propose an online adaptation algorithm for the DBN model, leading to a self-adaptive DBN. A prototype system is built and tested for technical performance and user experience. The quantitative evaluation shows that our method achieves over 13% improvement in accuracy as compared to baseline approaches based on hidden Markov model. In the user study, our system guides the participants to their destinations, emulating a human usher in multiple aspects. Liyuan Li, Qianli Xu, Vijay Chandrasekhar 0001, Joo-Hwee Lim, Cheston Tan, Michal Mukawa |
IEEE Trans. Cybern. | 1 |
| 2017 | Towards Detection of Bus Driver Fatigue Based on Robust Visual Analysis of Eye StateabstractDriver's fatigue is one of the major causes of traffic accidents, particularly for drivers of large vehicles (such as buses and heavy trucks) due to prolonged driving periods and boredom in working conditions. In this paper, we propose a vision-based fatigue detection system for bus driver monitoring, which is easy and flexible for deployment in buses and large vehicles. The system consists of modules of head-shoulder detection, face detection, eye detection, eye openness estimation, fusion, drowsiness measure percentage of eyelid closure (PERCLOS) estimation, and fatigue level classification. The core innovative techniques are as follows: 1) an approach to estimate the continuous level of eye openness based on spectral regression; and 2) a fusion algorithm to estimate the eye state based on adaptive integration on the multimodel detections of both eyes. A robust measure of PERCLOS on the continuous level of eye openness is defined, and the driver states are classified on it. In experiments, systematic evaluations and analysis of proposed algorithms, as well as comparison with ground truth on PERCLOS measurements, are performed. The experimental results show the advantages of the system on accuracy and robustness for the challenging situations when a camera of an oblique viewing angle to the driver's face is used for driving state monitoring. Bappaditya Mandal, Liyuan Li, Gang S. Wang, Jie Lin 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Performance evaluation of local descriptors and distance measures on benchmarks and first-person-view videos for face identification
Bappaditya Mandal, Liyuan Li, Ashraf A. Kassim |
Neurocomputing | 3 |
| 2015 | Whole space subclass discriminant analysis for face recognitionabstractIn this work, we propose to divide each class (a person) into subclasses using spatial partition trees which helps in better capturing the intra-personal variances arising from the appearances of the same individual. We perform a comprehensive analysis on within-class and within-subclass eigen-spectrums of face images and propose a novel method of eigen-spectrum modeling which extracts discriminative features of faces from both within-subclass and total or between-subclass scatter matrices. Effective low-dimensional face discriminative features are extracted for face recognition (FR) after performing discriminant evaluation in the entire eigenspace. Experimental results on popular face databases (AR, FERET) and the challenging unconstrained YouTube Face database show the superiority of our proposed approach on all three databases. Bappaditya Mandal, Liyuan Li, Vijay Chandrasekhar 0001, Joo-Hwee Lim |
ICIP | 2 |
| 2014 | Incremental Graph Clustering for Efficient Retrieval from Streaming Egocentric Video DataabstractWith wearable devices like Google Glass, it will soon become possible to record everything we see. We envision a system where one's entire visual memory is captured, stored and indexed. One of the biggest challenges is the scale of the retrieval problem. In this work, we focus on how to organize streaming egocentric video data. Egocentric video data is highly redundant, in that, we see several objects and scenes repeatedly as we go about our lives. To exploit this redundancy, we propose an evolving sparse-graph representation for egocentric video data. We propose an incremental local density clustering scheme, which learns salient objects and scenes for streaming egocentric video data. We use the density clustering scheme to prune redundant data in the database. For image-retrieval applications, by retaining only representative nodes from dense sub graphs in the streaming data source, we show we can achieve 90% of peak recall by retaining only 1% of data, with a significant 18% improvement in absolute recall over naive uniform sub sampling of the egocentric video data. Vijay Chandrasekhar 0001, Cheston Tan, Wu Min, Liyuan Li, Xiaoli Li 0001, Joo-Hwee Lim |
ICPR | 4 |
| 2014 | A wearable virtual guide for context-aware cognitive indoor navigationabstractIn this paper, we explore a new way to provide context-aware assistance for indoor navigation using a wearable vision system. We investigate how to represent the cognitive knowledge of wayfinding based on first-person-view videos in real-time and how to provide context-aware navigation instructions in a human-like manner. Inspired by the human cognitive process of wayfinding, we propose a novel cognitive model that represents visual concepts as a hierarchical structure. It facilitates efficient and robust localization based on cognitive visual concepts. Next, we design a prototype system that provides intelligent context-aware assistance based on the cognitive indoor navigation knowledge model. We conducted field tests and evaluated the system's efficacy by benchmarking it against traditional 2D maps and human guidance. The results show that context-awareness built on cognitive visual perception enables the system to emulate the efficacy of a human guide, leading to positive user experience. Qianli Xu, Liyuan Li, Joo-Hwee Lim, Cheston Tan, Michal Mukawa, Gang S. Wang |
Mobile HCI | 2 |
| 2014 | Extended Spectral Regression for efficient scene recognition
Liyuan Li, Weixun Goh, Joo-Hwee Lim, Sinno Jialin Pan |
Pattern Recognit. | 1 |
| 2013 | Designing engagement-aware agents for multiparty conversationsabstractRecognizing users' engagement state and intentions is a pressing task for computational agents to facilitate fluid conversations in situated interactions. We investigate how to quantitatively evaluate high-level user engagement and intentions based on low-level visual cues, and how to design engagement-aware behaviors for the conversational agents to behave in a sociable manner. Drawing on machine learning techniques, we propose two computational models to quantify users' attention saliency and engagement intentions. Their performances are validated by a close match between the predicted values and the ground truth annotation data. Next, we design a novel engagement-aware behavior model for the agent to adjust its direction of attention and manage the conversational floor based on the estimated users' engagement. In a user study, we evaluated the agent's behaviors in a multiparty dialog scenario. The results show that the agent's engagement-aware behaviors significantly improved the effectiveness of communication and positively affected users' experience. Qianli Xu, Liyuan Li, Gang S. Wang |
CHI | 2 |
| 2013 | A Wearable Cognitive Vision System for Navigation Assistance in Indoor Environment
Liyuan Li, Gang S. Wang, Weixun Goh, Joo-Hwee Lim, Cheston Tan |
ICONIP (3) | 1 |
| 2012 | Vision-based attention estimation and selection for social robot to perform natural interaction in the open worldabstractIn this paper, a novel vision system is proposed to estimate attention of people from rich visual clues for social robot to perform natural interactions with multiple participants in public environments. The vision detection and recognition modules include multi-person detection and tracking, upper-body pose recognition, face and gaze detection, lip motion analysis for speaking recognition, and facial expression recognition. A computational approach is proposed to generate a quantitative estimation of human attention. The vision system is implemented on a robotic receptionist "EVE" and encouraging results have been obtained. Liyuan Li, Xinguo Yu, Jun Li 0005, Gang S. Wang, Ji Yu Shi, Yeow Kee Tan, Haizhou Li 0001 |
HRI | 1 |
| 2012 | Robust Multiperson Detection and Tracking for Mobile Service and Social RobotsabstractThis paper proposes an efficient system which integrates multiple vision models for robust multiperson detection and tracking for mobile service and social robots in public environments. The core technique is a novel maximum likelihood (ML)-based algorithm which combines the multimodel detections in mean-shift tracking. First, a likelihood probability which integrates detections and similarity to local appearance is defined. Then, an expectation-maximization (EM)-like mean-shift algorithm is derived under the ML framework. In each iteration, the E-step estimates the associations to the detections, and the M-step locates the new position according to the ML criterion. To be robust to the complex crowded scenarios for multiperson tracking, an improved sequential strategy to perform the mean-shift tracking is proposed. Under this strategy, human objects are tracked sequentially according to their priority order. To balance the efficiency and robustness for real-time performance, at each stage, the first two objects from the list of the priority order are tested, and the one with the higher score is selected. The proposed method has been successfully implemented on real-world service and social robots. The vision system integrates stereo-based and histograms-of-oriented-gradients-based human detections, occlusion reasoning, and sequential mean-shift tracking. Various examples to show the advantages and robustness of the proposed system for multiperson tracking from mobile robots are presented. Quantitative evaluations on the performance of multiperson tracking are also performed. Experimental results indicate that significant improvements have been achieved by using the proposed method. Liyuan Li, Shuicheng Yan, Xinguo Yu, Yeow Kee Tan, Haizhou Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Spatialized epitome and its applicationsabstractDue to the lack of explicit spatial consideration, existing epitome model may fail for image recognition and target detection, which directly motivates us to propose the so-called spatialized epitome in this paper. Extended from the original graphical model of epitome, the spatialized epitome provides a general framework to integrate both appearance and spatial arrangement of patches in the image to achieve a more precise likelihood representation for image(s) and eliminate ambiguities in image reconstruction and recognition. From the extended graphical model of epitome, an EM learning procedure is derived under the framework of variational approximation. The learning procedure can generate an optimized summary of the image appearance with spatial distribution of the similar patches. From the spatialized epitome, we present a principled way of inferring the probability of a new input image under the learnt model and thereby enabling image recognition and target detection. We show how the incorporation of spatial information enhances the epitome's ability for discrimination on several vision tasks, e.g., misalignment/cross-pose face recognition and vehicle detection with a few training samples. Xinqi Chu, Shuicheng Yan, Liyuan Li, Kap Luk Chan, Thomas S. Huang |
CVPR | 3 |
| 2010 | Human Action Recognition by Negative Space AnalysisabstractWe propose a novel region-based method to recognize human actions by analyzing regions surrounding the human body, termed as negative space according to art theory, whereas other region-based approaches work with silhouette of the human body. We find that negative space provides sufficient information to describe each pose. It can also overcome some limitations of silhouette based methods such as leaks or holes in the silhouette. Each negative space can be approximately represented by simple shapes, resulting in computationally inexpensive feature description that supports fast and accurate action recognition. The proposed system has obtained 100% accuracy on the Weizmann human action dataset and is found more robust with respect to partial occlusion, shadow, noisy segmentation and non-rigid deformation of actions than other methods. Shah Atiqur Rahman, Liyuan Li, Maylor K. H. Leung |
CW | 2 |
| 2010 | HOG based multi-stage object detection and pose recognition for service robotabstractThis paper develops a HOG-based multistage approach for object detection and object pose recognition for service robots. This approach makes use of the merits of both multi-class and bi-class HOG-based detectors to form a three-stage algorithm at low computing cost. In the first stage, the multi-class classifier with coarse features is employed to estimate the orientation of a potential target object in the image; in the second stage, a bi-class detector corresponding to the detected orientation with intermediate level features is used to filter out most of false positives; and in the third stage, a bi-class detector corresponding to the detected orientation using fine features is used to achieve accurate detection with low rate of false positives. The training of multi-class and bi-class SVMs with their respective features in different levels is described. Experiments in real-world environments have shown that the proposed method is much more accurate than the detection method as it uses only multi-class detector. The proposed method is also much more efficient than the detection method as it uses a bi-class detector for each possible orientation. The approach works well on the scenarios where the SIFT-based detector may fail. The method can achieve real-time object detection, localization, and pose recognition on a P4 2.4GHz PC. Xinguo Yu, Liyuan Li, Kah Eng Hoe |
ICARCV | 3 |
| 2010 | Epitomized Summarization of Wireless Capsule Endoscopic Videos for Efficient Visualization
Xinqi Chu, Chee Khun Poh, Liyuan Li, Kap Luk Chan, Shuicheng Yan, Weijia Shen, That Mon Htwe, Jiang Liu 0001, Joo-Hwee Lim, Eng Hui Ong |
MICCAI (2) | 3 |
| 2009 | Lift-button detection and recognition for service robot in buildingsabstractLift operation is one of critical functions for mobile service robot to move across levels in buildings. Lift operation poses the problem of lift-button detection and recognition for computer vision. This paper presents a framework for vision-based lift-button detection and recognition. This problem is challenging due to reflection and complex background. To achieve the robustness in lift operation, we adopt multiple techniques to combat the challenges. First, we propose a multiple partial models method to increase the robustness of button panel detection. Second, we do button recognition by combining structural inference, Hough transform, and multi-symbol recognition techniques. Lastly, we use ultrasonic distance measure devices to aid the vision of the robot. The experimental results show that our framework can achieve the promising results in recognizing both the internal and external buttons of lift. Xinguo Yu, Liyuan Li, Kah Eng Hoe |
ICIP | 3 |
| 2009 | ML-fusion based multi-model human detection and tracking for robust human-robot interfacesabstractA novel stereo vision system for real-time human detection and tracking on a mobile service robot is presented in this paper. The system integrates the individually enhanced stereo-based human detection, HOG-based human detection, color-based tracking, and motion estimation for the robust detection and tracking of humans with large appearance and scale variations in real-world environments. A new framework of maximum likelihood based multi-model fusion is proposed to fuse these four human detection and tracking models according to the detection-track associations in 3D space, which is robust to the possible missed detections, false detections, and duplicated responses from the individual models. Multi-person tracking is implemented in a sequential near-to-far way, which well alleviates the difficulties caused by human-over-human occlusions. Extensive experimental results demonstrate the robustness of the proposed system under real-world scenarios with large variations in lighting conditions, cluttered backgrounds, human clothes and postures, and complex occlusion situations. Significant improvements in human detection and tracking have been achieved. The system has been deployed on six robot butlers to serve drinks, and showed encouraging performance in open ceremony events. Liyuan Li, Kah Eng Hoe, Shuicheng Yan, Xinguo Yu |
WACV | 1 |
| 2009 | Interactive broadcast services for live soccer video based on instant semantics acquisition
Xinguo Yu, Liyuan Li, Hon Wai Leong |
J. Vis. Commun. Image Represent. | 2 |
| 2008 | Unsupervised learning of human perspective context using ME-DT for efficient human detection in surveillanceabstractA novel and automated technique for learning human perspective context (HPC) from a scene is proposed in this paper. It is found that two models are required to describe HPC for camera tilt angle ranging from 0° to 50°. From a scene, the tilt angle can be inferred from the observed human shapes and head/foot positions. Afterward, a novel ME-DT (Model Estimation - Data Tuning) algorithm is proposed to learn human perspective context from live data of various degrees of uncertainties. The uncertainties may come from the variations of human individual heights and poses, and segmentation/recognition errors. ME-DT not only estimates the model parameters from the training data but also tunes the data to achieve a better head-foot correlation. The human perspective context provides a feasible constraint on the scales, positions, and orientations of humans in the scene. Applying this constraint to the HOG human detection, great reduction of the detection windows and improved performances have been obtained compared to conventional methods. Liyuan Li, Maylor K. H. Leung |
CVPR | 1 |
| 2008 | Multi-strategy object tracking in complex situation for video surveillanceabstractIn this paper, a novel method of multi-strategy object tracking for video surveillance is proposed. Under this framework, the moving and stationary objects are tracked separately, so that different reliable features for different types of objects can be exploited efficiently. For a moving object, the global color features called Dominant Color Histogram (DCH) are reliable for object tracking. An efficient sequential approach is employed, which first estimates the depth order of the objects using DCH, then tracks each individual one-by-one with mean-shift and exclusion operations. For a stationary object, the image template is accurate for object representation and matching. A layer model is employed for stationary object tracking. Stationary objects are classified as “visible”, “occluded”, and “removed”. For people stop moving in scene, they seldom stay completely motionless. Therefore, two more states, “changing pose”, and “start moving” are added. Once the stationary person is detected as “start moving”, he will be switched to moving object tracking seamlessly. The proposed method has been successfully applied in a real-time intelligent CCTV surveillance system for unusual event detection and tested in both real-world public sites and public datasets from PETS2006. Very encouraging results have been obtained. Ruijiang Luo, Liyuan Li, Qibin Sun |
ISCAS | 2 |
| 2008 | An Efficient Sequential Approach to Tracking Multiple Objects Through Crowds for Real-Time Intelligent CCTV SystemsabstractEfficiency and robustness are the two most important issues for multiobject tracking algorithms in real-time intelligent video surveillance systems. We propose a novel 2.5-D approach to real-time multiobject tracking in crowds, which is formulated as a maximum a posteriori estimation problem and is approximated through an assignment step and a location step. Observing that the occluding object is usually less affected by the occluded objects, sequential solutions for the assignment and the location are derived. A novel dominant color histogram (DCH) is proposed as an efficient object model. The DCH can be regarded as a generalized color histogram, where dominant colors are selected based on a given distance measure. Comparing with conventional color histograms, the DCH only requires a few color components (31 on average). Furthermore, our theoretical analysis and evaluation on real data have shown that DCHs are robust to illumination changes. Using the DCH, efficient implementations of sequential solutions for the assignment and location steps are proposed. The assignment step includes the estimation of the depth order for the objects in a dispersing group, one-by-one assignment, and feature exclusion from the group representation. The location step includes the depth-order estimation for the objects in a new group, the two-phase mean-shift location, and the exclusion of tracked objects from the new position in the group. Multiobject tracking results and evaluation from public data sets are presented. Experiments on image sequences captured from crowded public environments have shown good tracking results, where about 90% of the objects have been successfully tracked with the correct identification numbers by the proposed method. Our results and evaluation have indicated that the method is efficient and robust for tracking multiple objects (>or= 3) in complex occlusion for real-world surveillance scenarios. Liyuan Li, Weimin Huang 0002, Irene Y. H. Gu, Ruijiang Luo, Qi Tian 0002 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2006 | Region-Based Statistical Background Modeling for Foreground Object SegmentationabstractThis paper proposes a novel region-based scheme for dynamically modeling time-evolving statistics of video background, leading to an effective segmentation of foreground moving objects for a video surveillance system. In (L. Li et al., 2004) statistical-based video surveillance systems employ a Bayes decision rule for classifying foreground and background changes in individual pixels. Although principal feature representations significantly reduce the size of tables of statistics, pixel-wise maintenance remains a challenge due to the computations and memory requirement. The proposed region-based scheme, which is an extension of the above method, replaces pixel-based statistics by region-based statistics through introducing dynamic background region (or pixel) merging and splitting. Simulations have been performed to several outdoor and indoor image sequences, and results have shown a significant reduction of memory requirements for tables of statistics while maintaining relatively good quality in foreground segmented video objects. Kristof Op De Beeck, Irene Y. H. Gu, Liyuan Li, Mats Viberg, Bart De Moor |
ICIP | 3 |
| 2005 | Affine Object Tracking with Kernel-Based Spatial-Color RepresentationabstractThis paper presents a new visual tracking method that can achieve accurate estimation of affine transformation and precise spatial-color representation. The estimation of transformation provides more information than translation for better motion understanding and also helps maintain the precise representation; the precise representation enables tracking objects in highly-cluttered environment. The basis of the method is a kernel-based similarity measure called affine matching that describes the relationship between image regions with respect to affine transformation parameters. Based on the similarity measure, a mathematical solution is derived for estimating the transformation parameters for moving objects in videos. Various experiments have yielded positive results. Haihong Zhang, Liyuan Li |
CVPR (1) | 4 |
| 2004 | Stereo-based human detection for mobile service robotsabstractWithout knowledge of background or motion feature, detecting humans from a 2D image is still a tough task. In this paper, a novel stereo-based method to detect human objects for mobile service robots is proposed. Human objects are detected from the stereo spatial space through three distinct steps: (i) human oriented scale-adaptive filtering to aggregate and enhance the evidence of human presence, (ii) human like object segmentation, and (iii) human object identification based on the matching of a deformable head shoulder template to the evidence from both stereo and edge information. Systematic evaluation of the experimental results show that the high accuracy rates for human detection have been achieved with fewer constraints on the human operator, robot and the environment which they are in. Liyuan Li, Ying Ting Koh, Shuzhi Sam Ge |
ICARCV | 1 |
| 2004 | Stereo-based human read detection from crowd scenesabstractIn this paper, a novel stereo-based head detection method is proposed for human detection in crowd scene. It contains three steps: (1) scale-adaptive filtering, (2) spurious clue suppression and (3) human head location. With the depth information, the sizes of human heads could be estimated. From this, 3D scale-adaptive filtering is proposed. It is applied for extracting the likelihood evidence of heads from the stereo image. In the second step, the extracted points whose positions in the real space are much higher or lower than the average human height above the ground surface are further suppressed. Finally, human heads are located by applying a mean-shift algorithm to the likelihood map. Good results of detecting human heads in crowds have been obtained from the experiments on real scene. Liyuan Li, Terence Sim |
ICIP | 2 |
| 2004 | Statistical modeling of complex backgrounds for foreground object detectionabstractThis paper addresses the problem of background modeling for foreground object detection in complex environments. A Bayesian framework that incorporates spectral, spatial, and temporal features to characterize the background appearance is proposed. Under this framework, the background is represented by the most significant and frequent features, i.e., the principal features, at each pixel. A Bayes decision rule is derived for background and foreground classification based on the statistics of principal features. Principal feature representation for both the static and dynamic background pixels is investigated. A novel learning method is proposed to adapt to both gradual and sudden "once-off" background changes. The convergence of the learning process is analyzed and a formula to select a proper learning rate is derived. Under the proposed framework, a novel algorithm for detecting foreground objects from complex environments is then established. It consists of change detection, change classification, foreground segmentation, and background maintenance. Experiments were conducted on image sequences containing targets of interest in a variety of environments, e.g., offices, public buildings, subway stations, campuses, parking lots, airports, and sidewalks. Good results of foreground detection were obtained. Quantitative evaluation and comparison with the existing method show that the proposed method provides much improved results. Liyuan Li, Weimin Huang 0002, Irene Y. H. Gu, Qi Tian 0002 |
IEEE Trans. Image Process. | 1 |
| 2003 | Foreground object detection from videos containing complex backgroundabstractThis paper proposes a novel method for detection and segmentation of foreground objects from a video which contains both stationary and moving background objects and undergoes both gradual and sudden "once-off" changes. A Bayes decision rule for classification of background and foreground from selected feature vectors is formulated. Under this rule, different types of background objects will be classified from foreground objects by choosing a proper feature vector. The stationary background object is described by the color feature, and the moving background object is represented by the color co-occurrence feature. Foreground objects are extracted by fusing the classification results from both stationary and moving pixels. Learning strategies for the gradual and sudden "once-off" background changes are proposed to adapt to various changes in background through the video. The convergence of the learning process is proved and a formula to select a proper learning rate is also derived. Experiments have shown promising results in extracting foreground objects from many complex backgrounds including wavering tree branches, flickering screens and water surfaces, moving escalators, opening and closing doors, switching lights and shadows of moving objects. Liyuan Li, Weimin Huang 0002, Irene Y. H. Gu, Qi Tian 0002 |
ACM Multimedia | 1 |
| 2003 | Face recognition by incremental learningabstractOne of the important features for human machine interaction is its ability to recognize human faces. This paper presents a novel architecture suitable for real time robotic face recognition by learning a person's face incrementally, where the Gabor features at respective feature locations of a face are used to derive a similarity measurement. A face tracking followed by a clustering technique is used to learn a person's face appearance variance when the system interacts with the person. The recognition by learning proposed in this paper is similar to the partial memory incremental learning method, where we proposed a novel approach to the learning and updating process. Experiment shows significant improvement in the face recognition performance after learning over the time and with more interaction between a person and the system. Weimin Huang 0002, Benghai Lee, Liyuan Li, Karianto Leman |
SMC | 3 |
| 2003 | Principal color representation for tracking personsabstractThis paper proposes a novel method for tracking persons based on the principal colors of human objects. First, an efficient human object representation method, principal color representation (PCR), is proposed. Asymmetric similarity measures are then proposed based on the principal color representation. These asymmetric similarity measures could be used to evaluate the matching between two individuals as well as visual evident of an individual in a group. An efficient algorithm for tracking persons as individuals or in groups is then described. The method has been tested using image sequences containing multiple moving persons frequently gathering and separating. Our test results have shown that proposed method has successfully tracked both persons as individuals or in groups, and is robust to illumination changes. Liyuan Li, Weimin Huang 0002, Irene Y. H. Gu, Karianto Leman, Qi Tian 0002 |
SMC | 1 |
| 2002 | Foreground Object Detection in Changing Background Based on Color Co-Occurrence StatisticsabstractThis paper proposes a novel method for detecting foreground objects in nonstationary complex environments containing moving background objects. We derive a Bayes decision rule for classification of background and foreground changes based on inter-frame color co-occurrence statistics. An approach to store and fast retrieve color co-occurrence statistics is also established In the proposed method, foreground objects are detected in two steps. First, both foreground and background changes are extracted using background subtraction and temporal differencing. The frequent background changes are then recognized using the Bayes decision rule based on the learned color co-occurrence statistics. Both short-term and longterm strategies to learn the frequent background changes are proposed Experiments have shown promising results in detecting foreground objects from video containing wavering tree branches and flickering screens/water surface. The proposed method has shown better performance as compared with two existing methods. Liyuan Li, Weimin Huang 0002, Irene Y. H. Gu, Qi Tian 0002 |
WACV | 1 |
| 2002 | Integrating intensity and texture differences for robust change detectionabstractWe propose a novel technique for robust change detection based upon the integration of intensity and texture differences between two frames. A new accurate texture difference measure based on the relations between gradient vectors is proposed. The mathematical analysis shows that the measure is robust with respect to noise and illumination changes. Two ways to integrate the intensity and texture differences have been developed. The first combines the two measures adaptively according to the weightage of texture evidence, while the second does it optimally with additional constraint of smoothness. The parameters of the algorithm are selected automatically based on a statistic analysis. An algorithm is developed for fast implementation. The computational complexity analysis indicates that the proposed technique can run in real-time. The experiment results are evaluated both visually and quantitatively. They show that by exploiting both intensity and texture differences for change detection, one can obtain much better segmentation results than using the intensity or structure difference alone. Liyuan Li, Maylor K. H. Leung |
IEEE Trans. Image Process. | 1 |
| 2001 | Robust Change Detection by Fusing Intensity and Texture DifferencesabstractThe paper proposes a novel technique for robust change detection based upon the integration of intensity and texture differences between two frames. A new texture difference measure based on the relations between gradient vectors is described. The robustness of the measure with respect to noise and illumination changes has been analyzed. Two ways to integrate the intensity and texture differences are proposed. The first combines two measures according to the weightage of texture evidence, while the second takes into additional constraint of smoothness. The parameters of the algorithm are selected automatically. The computational complexity analysis indicates that the proposed technique can run in real-time. Experimental results show that by exploiting both intensity and texture differences for change detection, one can obtain much better segmentation results than using the intensity or structure difference alone. Liyuan Li, Maylor K. H. Leung |
CVPR (1) | 1 |
| 1999 | Corner Detection and Interpretation on Planar Curves Using Fuzzy ReasoningabstractThe problem of corner detection on planar curves is examined based on human perception of local graphic features. First, a set of fuzzy patterns of contour points are established. Then, corner detection is characterized as a fuzzy classification problem that contains three stages: evaluation, classification, and location. Compared with existing methods, the proposed approach is superior in that it explains the curve, instead of simple labeling, and it performs based on human perception. Experimental results on shapes of various complexities are presented. The performance with respect to noise is also addressed. Liyuan Li, Weinan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Fast recursive algorithms for two-dimensional thresholding
Liyuan Li, Weinan Chen |
Pattern Recognit. | 2 |
| 1997 | Gray level image thresholding based on fisher linear projection of two-dimensional histogram
Liyuan Li, Weinan Chen |
Pattern Recognit. | 1 |