Frank P.-W. Lo

dblp:246/7640 · also Frank Po Wen Lo · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-0358-6567ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 An AI-Driven Bionic Whisker System Assisting for Clinical Gastrointestinal Disease Screening
abstract
Effective early screenings for gastrointestinal diseases are crucial for reducing mortality through timely interventions and improving life expectancy. In this paper, a strain effect-based biomimetic artificial whisker system is proposed to extract the structural and textural information of the tissues in the lumen based on interactive tactile perception data and an end-to-end screening algorithm. Benchmark experiment of the proposed method and an ex-vivo pilot study are conducted to characterize the baseline performance and feasibility of detecting several common tissue structures in surgical application scenarios. Our method shows promising results, with a test accuracy of up to 97.27% and a kappa value of 0.9590. This integrated hardware-software, end-to-end solution is promising to become an emerging human-machine interaction paradigm, empowering traditional healthcare applications.
Frank P.-W. Lo, James Calo, Benny P. L. Lo, Alex J. Thompson 0001, Eric M. Yeatman
IJCNN2
2024 Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring
abstract
Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eating behaviors of the subject. However, there currently is no method that is able to incorporate these visual clues and provide a comprehensive context of dietary intake from passive recording (e.g., is the subject sharing food with others, what food the subject is eating, and how much food is left in the bowl). On the other hand, privacy is a major concern while egocentric wearable cameras are used for capturing. In this article, we propose a privacy-preserved secure solution (i.e., egocentric image captioning) for dietary assessment with passive monitoring, which unifies food recognition, volume estimation, and scene understanding. By converting images into rich text descriptions, nutritionists can assess individual dietary intake based on the captions instead of the original images, reducing the risk of privacy leakage from images. To this end, an egocentric dietary image captioning dataset has been built, which consists of in-the-wild images captured by head-worn and chest-worn cameras in field studies in Ghana. A novel transformer-based architecture is designed to caption egocentric dietary images. Comprehensive experiments have been conducted to evaluate the effectiveness and to justify the design of the proposed architecture for egocentric dietary image captioning. To the best of our knowledge, this is the first work that applies image captioning for dietary intake assessment in real-life settings.
Jianing Qiu, Frank P.-W. Lo, Xiao Gu 0003, Modou L. Jobarteh, Wenyan Jia, Thomas Baranowski, Matilda Steiner-Asiedu, Alex K. Anderson, Megan A. McCrory, Edward Sazonov, Mingui Sun, Gary S. Frost, Benny P. L. Lo
IEEE Trans. Cybern.2
2024 Dietary Assessment With Multimodal ChatGPT: A Systematic Analysis
abstract
Conventional approaches to dietary assessment are primarily grounded in self-reporting methods or structured interviews conducted under the supervision of dietitians. These methods, however, are often subjective, inaccurate, and time-intensive. Although artificial intelligence (AI)-based solutions have been devised to automate the dietary assessment process, prior AI methodologies tackle dietary assessment in a fragmented landscape (e.g., merely recognizing food types or estimating portion size) and encounter challenges in their ability to generalize across a diverse range of food categories, dietary behaviors, and cultural contexts. Recently, the emergence of multimodal foundation models, such as GPT-4V, has exhibited transformative potential across a wide range of tasks in various research domains. These models have demonstrated remarkable generalist intelligence and accuracy, owing to their large-scale pre-training on broad datasets and substantially scaled model size. In this study, we explore the application of GPT-4V powering multimodal ChatGPT for dietary assessment, along with prompt engineering and passive monitoring techniques. We evaluated the proposed pipeline using a self-collected, semi free-living dietary intake dataset, captured through wearable cameras. Our findings reveal that GPT-4V excels in food detection under challenging conditions without any fine-tuning or adaptation using food-specific datasets. By guiding the model with specific language prompts (e.g., African cuisine), it shifts from recognizing common staples like rice and bread to accurately identifying regional dishes like banku and ugali. Another standout feature of GPT-4V is its contextual awareness. GPT-4V can leverage surrounding objects as scale references to deduce the portion sizes of food items, further facilitating the process of dietary assessment.
Frank P.-W. Lo, Jianing Qiu, Bo Xiao 0002, Wu Yuan 0001, Stamatia Giannarou, Gary S. Frost, Benny P. L. Lo
IEEE J. Biomed. Health Informatics1
2023 Learning-Based Inverse Kinematics Identification of the Tendon-Driven Robotic Manipulator for Minimally Invasive Surgery
abstract
It is well-known that the tendon-driven robotic manipulator plays an important role in robotic-assisted minimally invasive surgery (MIS). However, due to the intrinsic nonlinearities, uncertainties, slack and hysteresis introduced by the tendon-driven actuation, the tendon-driven robotic manipulator is difficult to model and control when compared with the traditional actuation styles. To serve the modeling purpose, in this paper, the deep-learning-based intelligent modeling of inverse kinematics in the snake-like tendon-driven surgical instrument is presented. In the proposed approach the Deep Recurrent Neural Network (DRNN) with Long Short-Term Memory (LSTM) architecture is adopted to memorize and identify the nonlinear inverse kinematics of the tendon-driven surgical instrument through the history of the motor and tip positions. To collect highly reliable data to train the DRNN, the experiment to generate training data is carefully designed with the consideration of the stainless tendon characters and motor limitations. During the designed controller movements, the kinematics data is obtained by recording the motor positions and the tip positions. Besides, it is noticed that there are correlations of the sequential data samples, which could significantly reduce the modeling accuracy. To remove the correlations and improve the modeling performance, the correlations of the sequential data samples are removed by modifying the training processes. Modeling results and detailed discussions verified the effectiveness of the proposed approach.
Bo Xiao 0002, Wuzhou Hong, Ziwei Wang 0001, Frank P.-W. Lo, Zhenhua Yu 0004, Ravi Vaidyanathan, Eric M. Yeatman
IECON4
2023 Large AI Models in Health Informatics: Applications, Challenges, and the Future
abstract
Large AI models, or foundation models, are models recently emerging with massive scales both parameter-wise and data-wise, the magnitudes of which can reach beyond billions. Once pretrained, large AI models demonstrate impressive performance in various downstream tasks. A prime example is ChatGPT, whose capability has compelled people's imagination about the far-reaching influence that large AI models can have and their potential to transform different domains of our lives. In health informatics, the advent of large AI models has brought new paradigms for the design of methodologies. The scale of multi-modal data in the biomedical and health domain has been ever-expanding especially since the community embraced the era of deep learning, which provides the ground to develop, validate, and advance large AI models for breakthroughs in health-related areas. This article presents a comprehensive review of large AI models, from background to their applications. We identify seven key sectors in which large AI models are applicable and might have substantial influence, including: 1) bioinformatics; 2) medical diagnosis; 3) medical imaging; 4) medical informatics; 5) medical education; 6) public health; and 7) medical robotics. We examine their challenges, followed by a critical discussion about potential future directions and pitfalls of large AI models in transforming the field of health informatics.
Jianing Qiu, Lin Li 0070, Jiankai Sun, Jiachuan Peng, Peilun Shi, Ruiyang Zhang, Yinzhao Dong, Kyle Lam, Frank P.-W. Lo, Bo Xiao 0002, Wu Yuan 0001, Ningli Wang, Dong Xu 0002, Benny P. L. Lo
IEEE J. Biomed. Health Informatics9
2023 An Intelligent Vision-Based Nutritional Assessment Method for Handheld Food Items
abstract
Dietary assessment has proven to be effective to evaluate the dietary intake of patients with diabetes and obesity. The traditional approach of accessing the dietary intake is to conduct a 24-hour dietary recall, a structured interview designed to obtain information on food categories and volume consumed by the participants. Due to unconscious biases in this kind of self-reporting approaches, many research studies have explored the use of vision-based approaches to provide accurate and objective assessments. Despite the promising results of food recognition by deep neural networks, there still exist several hurdles in deep learning-based food volume estimation ranging from domain shift between synthetic and raw 3D models, shape completion ambiguity and lack of large-scale paired training dataset. Therefore, this paper proposed an intelligent nutritional assessment approach via weakly-supervised point cloud completion, which aims to close the reality gap in 3D point cloud completion tasks and address the targeted challenges. Then the volume can be easily estimated from the completed representation of the food. Another major merit of our system is that it can be used to estimate the volume of handheld food items without requiring the constraints including placing the food items on a table or next to fiducial markers, which facilitates the implementation on both wearable and handheld cameras. Comprehensive experiments have been carried out on major benchmark datasets and self-constructed volume-annotated dataset respectively, in which the proposed method demonstrates comparable results with several strong fully-supervised baseline methods and shows superior completion ability in handling food volume estimation.
Frank P.-W. Lo, Yao Guo 0002, Yingnan Sun, Jianing Qiu, Benny P. L. Lo
IEEE Trans. Multim.1
2022 Lightweight Internet of Things Device Authentication, Encryption, and Key Distribution Using End-to-End Neural Cryptosystems
abstract
Device authentication, encryption, and key distribution are of vital importance to any Internet of Things (IoT) systems, such as the new smart city infrastructures. This is due to the concern that attackers could easily exploit the lack of strong security in IoT devices to gain unauthorized access to the system or to hijack IoT devices to perform denial-of-service attacks on other networks. With the rise of fog and edge computing in IoT systems, increasing numbers of IoT devices have been equipped with computing capabilities to perform data analysis with deep learning technologies. Deep learning on edge devices can be deployed in numerous applications, such as local cardiac arrhythmia detection on a smart sensing patch, but it is rarely applied to device authentication and wireless communication encryption. In this article, we propose a novel lightweight IoT device authentication, encryption, and key distribution approach using neural cryptosystems and binary latent space. The neural cryptosystems adopt three types of end-to-end encryption schemes: 1) symmetric; 2) public-key; and 3) without keys. A series of experiments was conducted to test the performance and security strength of the proposed neural cryptosystems. The experimental results demonstrate the potential of this novel approach as a promising security and privacy solution for the next-generation of IoT systems.
Yingnan Sun, Frank P.-W. Lo, Benny P. L. Lo
IEEE Internet Things J.2
2021 Deep3DRanker: A Novel Framework for Learning to Rank 3D Models with Self-Attention in Robotic Vision
abstract
Research on generating or processing point clouds has become an increasingly popular domain in robotic research due to its extensive applications, such as robotic grasping, augmented reality and autonomous vehicle navigation. In this paper, we explore a new research area on point clouds - Learning to rank 3D models captured from a single depth image. In the Learning To Rank (LTR) task, we aim at optimizing the order of a list of 3D models according to the given query. Inspired by the recent advances in Natural Language Processing (NLP), we propose a novel framework, namely Deep3DRanker, for ranking 3D models by leveraging graph-based encoding and self-attention mechanisms. Comprehensive experiments are conducted to validate our methods on publicly available YCB synthetic and YCB video datasets. The promising results have shown that our proposed framework is generic enough to be applicable with any combinations of randomly positioned, oriented, and unseen object items with accuracy ranging from 59.2% to 94.9%, which shows great potentials of the proposed framework for robotic applications, in particular, for making decisions under different circumstances.
Frank P.-W. Lo, Yao Guo 0002, Yingnan Sun, Jianing Qiu, Benny P. L. Lo
ICRA1
2021 Indoor Future Person Localization from an Egocentric Wearable Camera
abstract
Accurate prediction of future person location and movement trajectory from an egocentric wearable camera can benefit a wide range of applications, such as assisting visually impaired people in navigation, and the development of mobility assistance for people with disability. In this work, a new egocentric dataset was constructed using a wearable camera, with 8,250 short clips of a targeted person either walking 1) toward, 2) away, or 3) across the camera wearer in indoor environments, or 4) staying still in the scene, and 13,817 person bounding boxes were manually labelled. Apart from the bounding boxes, the dataset also contains the estimated pose of the targeted person as well as the IMU signal of the wearable camera at each time point. An LSTM-based encoder-decoder framework was designed to predict the future location and movement trajectory of the targeted person in this egocentric setting. Extensive experiments have been conducted on the new dataset, and have shown that the proposed method is able to reliably and better predict future person location and trajectory in egocentric videos captured by the wearable camera compared to three baselines.
Jianing Qiu, Frank P.-W. Lo, Xiao Gu 0003, Yingnan Sun, Benny P. L. Lo
IROS2
2021 Counting Bites and Recognizing Consumed Food from Videos for Passive Dietary Monitoring
abstract
Assessing dietary intake in epidemiological studies are predominantly based on self-reports, which are subjective, inefficient, and also prone to error. Technological approaches are therefore emerging to provide objective dietary assessments. Using only egocentric dietary intake videos, this work aims to provide accurate estimation on individual dietary intake through recognizing consumed food items and counting the number of bites taken. This is different from previous studies that rely on inertial sensing to count bites, and also previous studies that only recognize visible food items but not consumed ones. As a subject may not consume all food items visible in a meal, recognizing those consumed food items is more valuable. A new dataset that has 1,022 dietary intake video clips was constructed to validate our concept of bite counting and consumed food item recognition from egocentric videos. 12 subjects participated and 52 meals were captured. A total of 66 unique food items, including food ingredients and drinks, were labelled in the dataset along with a total of 2,039 labelled bites. Deep neural networks were used to perform bite counting and food item recognition in an end-to-end manner. Experiments have shown that counting bites directly from video clips can reach 74.15% top-1 accuracy (classifying between 0-4 bites in 20-second clips), and a MSE value of 0.312 (when using regression). Our experiments on video-based food recognition also show that recognizing consumed food items is indeed harder than recognizing visible ones, with a drop of 25% in F1 score.
Jianing Qiu, Frank P.-W. Lo, Ya-Yen Tsai, Yingnan Sun, Benny P. L. Lo
IEEE J. Biomed. Health Informatics2
2020 Point2Volume: A Vision-Based Dietary Assessment Approach Using View Synthesis
abstract
Dietary assessment is an important tool for nutritional epidemiology studies. To assess the dietary intake, the common approach is to carry out 24-h dietary recall (24HR), a structured interview conducted by experienced dietitians. Due to the unconscious biases in such self-reporting methods, many research works have proposed the use of vision-based approaches to provide accurate and objective assessments. In this article, a novel vision-based method based on real-time three-dimensional (3-D) reconstruction and deep learning view synthesis is proposed to enable accurate portion size estimation of food items consumed. A point completion neural network is developed to complete partial point cloud of food items based on a single depth image or video captured from any convenient viewing position. Once 3-D models of food items are reconstructed, the food volume can be estimated through meshing. Compared to previous methods, our method has addressed several major challenges in vision-based dietary assessment, such as view occlusion and scale ambiguity, and it outperforms previous approaches in accurate portion size estimation.
Frank P.-W. Lo, Yingnan Sun, Jianing Qiu, Benny P. L. Lo
IEEE Trans. Ind. Informatics1
2020 Image-Based Food Classification and Volume Estimation for Dietary Assessment: A Review
abstract
A daily dietary assessment method named 24-hour dietary recall has commonly been used in nutritional epidemiology studies to capture detailed information of the food eaten by the participants to help understand their dietary behaviour. However, in this self-reporting technique, the food types and the portion size reported highly depends on users' subjective judgement which may lead to a biased and inaccurate dietary analysis result. As a result, a variety of visual-based dietary assessment approaches have been proposed recently. While these methods show promises in tackling issues in nutritional epidemiology studies, several challenges and forthcoming opportunities, as detailed in this study, still exist. This study provides an overview of computing algorithms, mathematical models and methodologies used in the field of image-based dietary assessment. It also provides a comprehensive comparison of the state of the art approaches in food recognition and volume/weight estimation in terms of their processing speed, model accuracy, efficiency and constraints. It will be followed by a discussion on deep learning method and its efficacy in dietary assessment. After a comprehensive exploration, we found that integrated dietary assessment systems combining with different approaches could be the potential solution to tackling the challenges in accurate dietary intake assessment.
Frank P.-W. Lo, Yingnan Sun, Jianing Qiu, Benny P. L. Lo
IEEE J. Biomed. Health Informatics1
2019 Mining Discriminative Food Regions for Accurate Food Recognition
Jianing Qiu, Frank P.-W. Lo, Yingnan Sun, Siyao Wang, Benny P. L. Lo
BMVC2
2019 A Novel Vision-based Approach for Dietary Assessment using Deep Learning View Synthesis
abstract
Dietary assessment system has proven as an effective tool to evaluate the eating behavior of patients suffering from diabetes and obesity. To assess the dietary intake, the traditional method is to carry out a 24-hour dietary recall (24HR), a structured interview aimed at capturing information on food items and portion size consumed by participants. However, unconscious biases are developed easily due to individual's subjective perception in this self-reporting technique which may lead to inaccuracy. Thus, this paper proposed a novel vision-based approach for estimating the volume of food items based on deep learning view synthesis and depth sensing techniques. In this paper, a point completion network is applied to perform 3D reconstruction of food items using a single depth image captured from any convenient viewing angle. Compared to previous approaches, the proposed method has addressed several key challenges in vision-based dietary assessment, such as view occlusion and scale ambiguity. Experiments have been carried out to examine this approach and showed the feasibility of the algorithm in accurate estimation of food volume.
Frank P.-W. Lo, Yingnan Sun, Jianing Qiu, Benny P. L. Lo
BSN1
2019 Assessing Individual Dietary Intake in Food Sharing Scenarios with a 360 Camera and Deep Learning
abstract
A novel vision-based approach for estimating individual dietary intake in food sharing scenarios is proposed in this paper, which incorporates food detection, face recognition and hand tracking techniques. The method is validated using panoramic videos which capture subjects' eating episodes. The results demonstrate that the proposed approach is able to reliably estimate food intake of each individual as well as the food eating sequence. To identify the food items ingested by the subject, a transfer learning approach is designed. 4, 200 food images with segmentation masks, among which 1,500 are newly annotated, are used to fine-tune the deep neural network for the targeted food intake application. In addition, a method for associating detected hands with subjects is developed and the outcomes of face recognition are refined to enable the quantification of individual dietary intake in communal eating settings.
Jianing Qiu, Frank P.-W. Lo, Benny P. L. Lo
BSN2
2019 A Deep Learning Approach on Gender and Age Recognition using a Single Inertial Sensor
abstract
Extracting human attributes, such as gender and age, from biometrics have received much attention in recent years. Gender and age recognition can provide crucial information for applications such as security, healthcare, and gaming. In this paper, a novel deep learning approach on gender and age recognition using a single inertial sensors is proposed. The proposed approach is tested using the largest available inertial sensor-based gait database with data collected from more than 700 subjects. To demonstrate the robustness and effectiveness of the proposed approach, 10 trials of inter-subject Monte-Carlo cross validation were conducted, and the results show that the proposed approach can achieve an averaged accuracy of 86.6%±2.4% for distinguishing two age groups: teen and adult, and recognizing gender with averaged accuracies of 88.6%±2.5% and 73.9%±2.8% for adults and teens respectively.
Yingnan Sun, Frank P.-W. Lo, Benny P. L. Lo
BSN2
2019 Visual Guidance and Automatic Control for Robotic Personalized Stent Graft Manufacturing
abstract
Personalized stent graft is designed to treat Abdominal Aortic Aneurysms (AAA). Due to the individual difference in arterial structures, stent graft has to be custom made for each AAA patient. Robotic platforms for autonomous personalized stent graft manufacturing have been proposed in recently which rely upon stereo vision systems for coordinating multiple robots for fabricating customized stent grafts. This paper proposes a novel hybrid vision system for real-time visual-sevoing for personalized stent-graft manufacturing. To coordinate the robotic arms, this system is based on projecting a dynamic stereo microscope coordinate system onto a static wide angle view stereo webcam coordinate system. The multiple stereo camera configuration enables accurate localization of the needle in 3D during the sewing process. The scale-invariant feature transform (SIFT) method and color filtering are implemented for stereo matching and feature identifications for object localization. To maintain the clear view of the sewing process, a visual-servoing system is developed for guiding the stereo microscopes for tracking the needle movements. The deep deterministic policy gradient (DDPG) reinforcement learning algorithm is developed for real-time intelligent robotic control. Experimental results have shown that the robotic arm can learn to reach the desired targets autonomously.
Frank P.-W. Lo, Benny P. L. Lo
ICRA3
2019 EEG-based user identification system using 1D-convolutional long short-term memory neural networks
Yingnan Sun, Frank P.-W. Lo, Benny P. L. Lo
Expert Syst. Appl.2
2018 Food volume estimation for quantifying dietary intake with a wearable camera
abstract
A novel food volume measurement technique is proposed in this paper for accurate quantification of the daily dietary intake of the user. The technique is based on simultaneous localisation and mapping (SLAM), a modified version of convex hull algorithm, and a 3D mesh object reconstruction technique. This paper explores the feasibility of applying SLAM techniques for continuous food volume measurement with a monocular wearable camera. A sparse map will be generated by SLAM after capturing the images of the food item with the camera and the multiple convex hull algorithm is applied to form a 3D mesh object. The volume of the target object can then be computed based on the mesh object. Compared to previous volume measurement techniques, the proposed method can measure the food volume continuously with no prior information such as pre-defined food shape model. Experiments have been carried out to evaluate this new technique and showed the feasibility and accuracy of the proposed algorithm in measuring food volume.
Anqi Gao, Frank P.-W. Lo, Benny P. L. Lo
BSN2