Yu Gan 0003

dblp:89/8500-3 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-3409-3412ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 57% Vision and language · 33% Image recognition and object detection · 10%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 50% Rendering · 50%
Human-computer interaction and pervasive computing
2 papers
Wearable and physiological sensing · 79% Health and well-being technologies · 21%
Computer networks
2 papers
Wireless sensing and localization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.012026
Screen Detection From Egocentric Image Streams Leveraging Multi-View Vision Language Model · IEEE Trans. Multim. 2026
Computer vision › 3D vision
3d reconstruction
0.912025
Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation · IEEE Trans. Vis. Comput. Graph. 2025
Computer vision › 3D vision › motion capture
human performance capture
0.912025
Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation · IEEE Trans. Vis. Comput. Graph. 2025
Virtual and augmented reality
augmented reality
0.912025
Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation · IEEE Trans. Vis. Comput. Graph. 2025
Rendering › physically based rendering › wave optics rendering › computer-generated holography
holographic rendering
0.912025
Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation · IEEE Trans. Vis. Comput. Graph. 2025
Computer vision › Image recognition and object detection
object detection
0.312026
Screen Detection From Egocentric Image Streams Leveraging Multi-View Vision Language Model · IEEE Trans. Multim. 2026
Health and well-being technologies › rehabilitation technology
telerehabilitation
0.312025
Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation · IEEE Trans. Vis. Comput. Graph. 2025
Wireless sensing and localization › device-free sensing
device-free localization
0.212013
Adaptive device-free passive localization coping with dynamic target speed · INFOCOM 2013
Wireless sensing and localization › ranging
acoustic ranging
0.112012
Push the limit of WiFi based localization for smartphones · MobiCom 2012
Wireless sensing and localization
indoor localization
0.112012
Push the limit of WiFi based localization for smartphones · MobiCom 2012
Wireless sensing and localization › indoor localization
wifi fingerprinting
0.112012
Push the limit of WiFi based localization for smartphones · MobiCom 2012
Wireless sensing and localization › smartphone sensing
smartphone-based localization
0.012012
Push the limit of WiFi based localization for smartphones · MobiCom 2012

Methods — techniques the papers use, named apart from their topics

gaussian rasterization · 2.6SMPL · 2.6HumanNeRF · 2.63d gaussian splatting · 2.6vision-language model · 2.0egocentric image analysis · 2.0RSS-based localization · 0.2joint mapping with ranging constraints · 0.1acoustic ranging · 0.1
YearPublicationVenuePosition
2026 Screen Detection From Egocentric Image Streams Leveraging Multi-View Vision Language Model
abstract
Accurately monitoring the screen exposure of young children is important for research related to screen use, such as childhood obesity, physical activity, and social interaction. Most existing studies rely upon self-report or manual measures from bulky wearable sensors, thus lacking efficiency and accuracy in capturing quantitative screen exposure data. In this work, we developed a novel screen detection framework that utilizes egocentric images from a wearable sensor, named the screen time tracker (STT), and a vision language model (VLM). In particular, we devised a multi-view VLM that takes multiple views from egocentric image streams and interprets screen exposure dynamically. We validated our approach by using a dataset of children's free-living activities, demonstrating significant improvement over existing methods in conventional vision language models and object detection models. The combination of a vision language model and a lightweight hardware design provides a novel solution in screen detection for children. The proposed framework has great potential to benefit children's behavioral study. The code is available at https://github.com/YGanLab/MV-VLM.
Xueshen Li, Sen Shen, Xinlong Hou, Xinran Gao, Steven Holiday, Matthew Cribbet, Susan W. White, Edward Sazonov, Yu Gan 0003
IEEE Trans. Multim.10
2025 Detection of Screen Usage During Eating Events Among Preschool-Aged Children
abstract
Detection of screen use during eating events in young children is a challenging task, as preschool-aged children have limited ability to self-report. The relationship between screen use and dietary intake is not yet well understood. In this study, we utilized a wearable camera to capture egocentric images and identify screen usage during eating events. The study involved 25 children (ages 3–5) who wore the device for two full days. The Recognize Anything Model (RAM) was used to analyze the images, generating object tags along with their associated tag scores. Using ANOVA F-values, the top 20 image tags were selected, and their confidence scores were employed as features in a Random Forest classifier to distinguish food and beverage images from non-food images. Screen use was estimated using the YOLO-5 object detection model. Our proposed framework achieved an accuracy of 87.7% for eating event detection, 95.7% for detecting eating events involving screen use, 66.1% for food and beverage image detection, and 83.5% for detecting screen use with food and beverage images. These findings underscore the potential of wearable cameras for monitoring screen time during eating events in preschool-aged children.
Tonmoy Ghosh, Md Billlal Hossain, Steven Holiday, Matthew Cribbet, Susan W. White, Yu Gan 0003, Edward Sazonov
ICIP6
2025 A two-stage proactive dialogue generator for efficient clinical information collection using large language model
Xueshen Li, Xinlong Hou, Nirupama Ravi, Yu Gan 0003
Expert Syst. Appl.5
2025 Real-Time, Free-Viewpoint Holographic Patient Rendering for Telerehabilitation via a Single Camera: A Data-Driven Approach With 3D Gaussian Splatting for Real-World Adaptation
abstract
Telerehabilitation is a cost-effective alternative to in-clinic rehabilitation. Although convenient, it lacks immersive and free-viewpoint patient visualization. Current research explores two solutions to this issue. Mesh-based methods use 3D models and motion capture for AR visualization. However, they are labor-intensive and less photorealistic than 2D images. Microsoft's Holoportation generates photorealistic 3D models with eight RGBD cameras in real time. However, it requires complex setups, high GPU power, and high-speed communication infrastructure, making deployment challenging. This article presents a Real-Time Free-Viewpoint Holographic Patient Rendering (RT-FVHP) system for telerehabilitation. Unlike traditional methods that require manually crafted assets such as 3D meshes, texture maps, and skeletal rigging, our data-driven approach eliminates the need for explicit asset definitions. Inspired by the HumanNeRF framework, we retarget dynamic human poses to a canonical pose and leverage 3D Gaussian Splatting to train a neural network in canonical space for patient representation. The trained model generates 2D RGB$\sigma$σ outputs via Gaussian Splatting rasterization, guided by camera parameters and human pose inputs. Compatible with HoloLens 2 and web-based platforms, RT-FVHP operates effectively under real-world conditions, including handling occlusions caused by treadmills. Occlusion handling is accomplished using our Shape-Enforced Gaussian Density Control (SGDC), which initializes and densifies 3D Gaussians in occluded regions using estimated SMPL human body priors. This approach minimizes manual intervention while ensuring complete body reconstruction. With efficient Gaussian rasterization, the model delivers real-time performance of up to 400 FPS at 1080p resolution on a dedicated RTX6000 GPU.
Shengting Cao, Jiamiao Zhao, Fei Hu 0001, Yu Gan 0003
IEEE Trans. Vis. Comput. Graph.4
2023 Single-Belt Versus Split-Belt: Intelligent Treadmill Control via Microphase Gait Capture for Poststroke Rehabilitation
abstract
Stroke is the leading long-term disability and causes a significant financial burden associated with rehabilitation. In poststroke rehabilitation, individuals with hemiparesis have a specialized demand for coordinated movement between the paretic and the nonparetic legs. The split-belt treadmill can effectively facilitate the paretic leg by slowing down the belt speed for that leg while the patient is walking on a split-belt treadmill. Although studies have found that split-belt treadmills can produce better gait recovery outcomes than traditional single-belt treadmills, the high cost of split-belt treadmills is a significant barrier to stroke rehabilitation in clinics. In this article, we design an AI-based system for the single-belt treadmill to make it act like a split-belt by adjusting the belt speed instantaneously according to the patient's microgait phases. This system only requires a low-cost RGB camera to capture human gait patterns. A novel microgait classification pipeline model is used to detect gait phases in real time. The pipeline is based on self-supervised learning that can calibrate the anchor video with the real-time video. We then use a ResNet-LSTM module to handle temporal information and increase accuracy. A real-time filtering algorithm is used to smoothen the treadmill control. We have tested the developed system with 34 healthy individuals and four stroke patients. The results show that our system is able to detect the gait microphase accurately and requires less human annotation in training, compared to the ResNet50 classifier. Our system "Splicer" is boosted by AI modules and performs comparably as a split-belt system, in terms of timely varying left/right foot speed, creating a hemiparetic gait in healthy individuals, and promoting paretic side symmetry in force exertion for stroke patients. This innovative design can potentially provide cost-effective rehabilitation treatment for hemiparetic patients.
Shengting Cao, Mansoo Ko, Chih-Ying Li, Fei Hu 0001, Yu Gan 0003
IEEE Trans. Hum. Mach. Syst.7
2023 Cardiac Adipose Tissue Segmentation via Image-Level Annotations
abstract
Automatically identifying the structural substrates underlying cardiac abnormalities can potentially provide real-time guidance for interventional procedures. With the knowledge of cardiac tissue substrates, the treatment of complex arrhythmias such as atrial fibrillation and ventricular tachycardia can be further optimized by detecting arrhythmia substrates to target for treatment (i.e., adipose) and identifying critical structures to avoid. Optical coherence tomography (OCT) is a real-time imaging modality that aids in addressing this need. Existing approaches for cardiac image analysis mainly rely on fully supervised learning techniques, which suffer from the drawback of workload on labor-intensive annotation process of pixel-wise labeling. To lessen the need for pixel-wise labeling, we develop a two-stage deep learning framework for cardiac adipose tissue segmentation using image-level annotations on OCT images of human cardiac substrates. In particular, we integrate class activation mapping with superpixel segmentation to solve the sparse tissue seed challenge raised in cardiac tissue segmentation. Our study bridges the gap between the demand on automatic tissue analysis and the lack of high-quality pixel-wise annotations. To the best of our knowledge, this is the first study that attempts to address cardiac tissue segmentation on OCT images via weakly supervised learning techniques. Within an in-vitro human cardiac OCT dataset, we demonstrate that our weakly supervised approach on image-level annotations achieves comparable performance as fully supervised methods trained on pixel-wise annotations.
Yu Gan 0003, Theresa Lye, Haofeng Zhang 0002, Andrew F. Laine, Elsa D. Angelini, Christine P. Hendon
IEEE J. Biomed. Health Informatics2
2020 Heterogeneity Measurement of Cardiac Tissues Leveraging Uncertainty Information from Image Segmentation
Yu Gan 0003, Theresa Lye, Haofeng Zhang 0002, Andrew F. Laine, Elsa D. Angelini, Christine P. Hendon
MICCAI (1)2
2013 Adaptive device-free passive localization coping with dynamic target speed
abstract
Device-free passive localization enables locating targets (e.g., intruders or victims) that do not carry any radio devices nor do they actively participate in the wireless localization process. This is because the wireless environments will get affected when people move into the area, which result in the changes of Received Signal Strength (RSS) of the wireless links. In this paper, we first show that the localization performance degrades significantly when people are moving in dynamic speeds. This is because existing studies in device-free passive localization system have an implicit assumption that the target is moving at a constant speed, which is not always true in practical scenarios. To cope with targets moving with dynamic speeds, we propose an adaptive speed change detection framework including three components: speed change detection, determination of time-window size and adaptive localization. Two speed change detection schemes have been developed to capture the changes of moving speed and adjust the time-window size adaptively to facilitate effective localization. We demonstrate that our framework is flexible to work with any device-free localization method using signal strength. Results from the real experiments confirm that our approach has over 30% improvement on both median and max localization error, under dynamically changing speed of the target.
Xiuyuan Zheng, Jie Yang 0003, Yingying Chen 0001, Yu Gan 0003
INFOCOM4
2012 Push the limit of WiFi based localization for smartphones
abstract
Highly accurate indoor localization of smartphones is critical to enable novel location based features for users and businesses. In this paper, we first conduct an empirical investigation of the suitability of WiFi localization for this purpose. We find that although reasonable accuracy can be achieved, significant errors (e.g., $6\sim8m$) always exist. The root cause is the existence of distinct locations with similar signatures, which is a fundamental limit of pure WiFi-based methods. Inspired by high densities of smartphones in public spaces, we propose a peer assisted localization approach to eliminate such large errors. It obtains accurate acoustic ranging estimates among peer phones, then maps their locations jointly against WiFi signature map subjecting to ranging constraints. We devise techniques for fast acoustic ranging among multiple phones and build a prototype. Experiments show that it can reduce the maximum and 80-percentile errors to as small as $2m$ and $1m$, in time no longer than the original WiFi scanning, with negligible impact on battery lifetime.
Hongbo Liu 0002, Yu Gan 0003, Jie Yang 0003, Simon Sidhom, Yan Wang 0003, Yingying Chen 0001, Fan Ye 0003
MobiCom2