Sizhe An

dblp:253/1603 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-9211-4886ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 PHD: Personalized 3D Human Body Fitting with Point Diffusion
abstract
We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/
Hsuan-I Ho, Po-Chen Wu, Ivan Shugurov, Chengcheng Tang, Abhay Mittal, Sizhe An, Manuel Kaufmann, Linguang Zhang
ICCV7
2024 SphereHead: Stable 3D Full-Head Synthesis with Spherical Tri-Plane Representation
Heyuan Li, Ce Chen, Tianhao Shi, Yuda Qiu, Sizhe An, Guanying Chen, Xiaoguang Han 0001
ECCV (75)5
2023 PanoHead: Geometry-Aware 3D Full-Head Synthesis in 360°
abstract
Synthesis and reconstruction of 3D human head has gained increasing interests in computer vision and computer graphics recently. Existing state-of-the-art 3D generative adversarial networks (GANs) for 3D human head synthesis are either limited to near-frontal views or hard to preserve 3D consistency in large view angles. We propose PanoHead, the first 3D-aware generative model that enables high-quality view-consistent image synthesis of full heads in 360° with diverse appearance and detailed geometry using only in-the-wild unstructured images for training. At its core, we lift up the representation power of recent 3D GANs and bridge the data alignment gap when training from in-the-wild images with widely distributed views. Specifically, we propose a novel two-stage self-adaptive image alignment for robust 3D GAN training. We further introduce a tri-grid neural volume representation that effectively addresses front-face and back-head feature entanglement rooted in the widely-adopted tri-plane formulation. Our method instills prior knowledge of 2D image segmentation in adversarial learning of 3D neural scene structures, enabling compositable head synthesis in diverse backgrounds. Benefiting from these designs, our method significantly outperforms previous 3D GANs, generating high-quality 3D heads with accurate geometry and diverse appearances, even with long wavy and afro hairstyles, renderable from arbitrary poses. Furthermore, we show that our system can reconstruct full 3D heads from single input images for personalized realistic 3D avatars.
Sizhe An, Yichun Shi, Guoxian Song, Ümit Y. Ogras, Linjie Luo
CVPR1
2023 PAniC-3D: Stylized Single-view 3D Reconstruction from Portraits of Anime Characters
abstract
We propose PAniC-3D, a system to reconstruct stylized 3D character heads directly from illustrated (p)ortraits of (ani)me (c)haracters. Our anime-style domain poses unique challenges to single-view reconstruction; compared to natural images of human heads, character portrait illustrations have hair and accessories with more complex and diverse geometry, and are shaded with non-photorealistic contour lines. In addition, there is a lack of both 3D model and portrait illustration data suitable to train and evaluate this ambiguous stylized reconstruction task. Facing these challenges, our proposed PAniC-3D architecture crosses the illustration-to-3D domain gap with a line-filling model, and represents sophisticated geometries with a volumetric radiance field. We train our system with two large new datasets (11.2k Vroid 3D models, 1k Vtuber portrait illustrations), and evaluate on a novel AnimeRecon benchmark of illustration-to-3D pairs. PAniC-3D significantly outper-forms baseline methods, and provides data to establish the task of stylized reconstruction from portrait illustrations.
Shuhong Chen, Kevin Zhang 0003, Yichun Shi, Yiheng Zhu 0003, Guoxian Song, Sizhe An, Janus Kristjansson, Matthias Zwicker
CVPR7
2023 Energy-Efficient On-Chip Training for Customized Home-based Rehabilitation Systems
abstract
Rehabilitation is an essential process for patients suffering from motor disorders. It is generally performed by experts in a clinical environment. Home-based rehabilitation systems allow patients to perform rehabilitation without going to clinics, thus, reducing the commute and healthcare costs. Human joint estimation allows visualization of body movements required for rehabilitation. However, the estimations can be inaccurate if they are not customized for the specific patient. Therefore, we propose a personalized rehabilitation system customized to new patients utilizing energy-efficient on-chip training with in-memory acceleration. Experiments show that it customizes to new patients successfully with an average 28.01% lower error and provides energy-efficient estimates of human joint coordinates with 611.1× lower inference energy and 14.0× faster training.
A. Alper Goksoy, Sizhe An, Ümit Y. Ogras
DAC2
2023 Transfer Learning for Human Activity Recognition Using Representational Analysis of Neural Networks
abstract
Human activity recognition (HAR) has increased in recent years due to its applications in mobile health monitoring, activity recognition, and patient rehabilitation. The typical approach is training a HAR classifier offline with known users and then using the same classifier for new users. However, the accuracy for new users can be low with this approach if their activity patterns are different than those in the training data. At the same time, training from scratch for new users is not feasible for mobile applications due to the high computational cost and training time. To address this issue, we propose a HAR transfer learning framework with two components. First, a representational analysis reveals common features that can transfer across users and user-specific features that need to be customized. Using this insight, we transfer the reusable portion of the offline classifier to new users and fine-tune only the rest. Our experiments with five datasets show up to 43% accuracy improvement and 66% training time reduction when compared to the baseline without using transfer learning. Furthermore, measurements on the hardware platform reveal that the power and energy consumption decreased by 43% and 68%, respectively, while achieving the same or higher accuracy as training from scratch. Our code is released for reproducibility. 1
Sizhe An, Ganapati Bhat, Suat Gumussoy, Ümit Y. Ogras
ACM Trans. Comput. Heal.1
2022 Fast and scalable human pose estimation using mmWave point cloud
abstract
Millimeter-Wave (mmWave) radar can enable high-resolution human pose estimation with low cost and computational requirements. However, mmWave data point cloud, the primary input to processing algorithms, is highly sparse and carries significantly less information than other alternatives such as video frames. Furthermore, the scarce labeled mmWave data impedes the development of machine learning (ML) models that can generalize to unseen scenarios. We propose a fast and scalable human pose estimation (FUSE) framework that combines multi-frame representation and meta-learning to address these challenges. Experimental evaluations show that FUSE adapts to the unseen scenarios 4× faster than current supervised learning approaches and estimates human joint coordinates with about 7 cm mean absolute error.
Sizhe An, Ümit Y. Ogras
DAC1
2022 mRI: Multi-modal 3D Human Pose Estimation Dataset using mmWave, RGB-D, and Inertial Sensors
abstract
The ability to estimate 3D human body pose and movement, also known as human pose estimation (HPE), enables many applications for home-based health monitoring, such as remote rehabilitation training. Several possible solutions have emerged using sensors ranging from RGB cameras, depth sensors, millimeter-Wave (mmWave) radars, and wearable inertial sensors. Despite previous efforts on datasets and benchmarks for HPE, few dataset exploits multiple modalities and focuses on home-based health monitoring. To bridge the gap, we present mRI, a multi-modal 3D human pose estimation dataset with mmWave, RGB-D, and Inertial Sensors. Our dataset consists of over 160k synchronized frames from 20 subjects performing rehabilitation exercises and supports the benchmarks of HPE and action detection. We perform extensive experiments using our dataset and delineate the strength of each modality. We hope that the release of mRI can catalyze the research in pose estimation, multi-modal learning, and action understanding, and more importantly facilitate the applications of home-based health monitoring.
Sizhe An, Ümit Y. Ogras
NeurIPS1
2022 MGait: Model-Based Gait Analysis Using Wearable Bend and Inertial Sensors
abstract
Movement disorders, such as Parkinson’s disease, affect more than 10 million people worldwide. Gait analysis is a critical step in the diagnosis and rehabilitation of these disorders. Specifically, step and stride lengths provide valuable insights into the gait quality and rehabilitation process. However, traditional approaches for estimating step length are not suitable for continuous daily monitoring since they rely on special mats and clinical environments. To address this limitation, this article presents a novel and practical step-length estimation technique using low-power wearable bend and inertial sensors. Experimental results show that the proposed model estimates step length with 5.49% mean absolute percentage error and provides accurate real-time feedback to the user.
Sizhe An, Yigit Tuncel, Toygun Basaklar, Gokul K. Krishnakumar, Ganapati Bhat, Ümit Y. Ogras
ACM Trans. Internet Things1
2021 Wearable Devices and Low-Power Design for Smart Health Applications: Challenges and Opportunities
abstract
Wearable devices can enable affordable and accessible smart health care services with the help of innovative low-power design and edge computing technologies. Indeed, novel wearable devices are already fueling a shift from a hospital-centric setting to more personalized home-based solutions [1]. A wide variety of miniature, flexible, and stretchable sensors enable collecting real-time data without impeding users’ daily routines. For example, inertial measurement units (IMUs) based on MEMS technology integrate a 9-axis accelerometer/gyroscope/magnetometer into a small package. Similarly, bend and stretch sensors embedded into clothes measure knee and hip angles, while biosensors track biopotentials, such as electrocardiogram (ECG) and electromyography (EMG). Then, novel edge-AI algorithms process the real-time data using low-power processors to build smart health applications ranging from health and activity monitoring to early diagnosis and prognosis [2] (Section B). One of the most critical challenges in wearable smart health applications is the stringent energy capacity imposed by size and weight constraints [2], [3]. All the required sensing, processing, and communications tasks must be performed without any manual charging or battery maintenance effort to maximize the user experience (Section C). The rest of this extended abstract discusses the challenges and potential solutions for driver applications and energy management techniques.
Toygun Basaklar, Yigit Tuncel, Sizhe An, Ümit Y. Ogras
ISLPED3
2021 MARS: mmWave-based Assistive Rehabilitation System for Smart Healthcare
abstract
Rehabilitation is a crucial process for patients suffering from motor disorders. The current practice is performing rehabilitation exercises under clinical expert supervision. New approaches are needed to allow patients to perform prescribed exercises at their homes and alleviate commuting requirements, expert shortages, and healthcare costs. Human joint estimation is a substantial component of these programs since it offers valuable visualization and feedback based on body movements. Camera-based systems have been popular for capturing joint motion. However, they have high-cost, raise serious privacy concerns, and require strict lighting and placement settings. We propose a millimeter-wave (mmWave)-based assistive rehabilitation system (MARS) for motor disorders to address these challenges. MARS provides a low-cost solution with a competitive object localization and detection accuracy. It first maps the 5D time-series point cloud from mmWave to a lower dimension. Then, it uses a convolution neural network (CNN) to estimate the accurate location of human joints. MARS can reconstruct 19 human joints and their skeleton from the point cloud generated by mmWave radar. We evaluate MARS using ten specific rehabilitation movements performed by four human subjects involving all body parts and obtain an average mean absolute error of 5.87 cm for all joint positions. To the best of our knowledge, this is the first rehabilitation movements dataset using mmWave point cloud. MARS is evaluated on the Nvidia Jetson Xavier-NX board. Model inference takes only 64 s and consumes 442 J energy. These results demonstrate the practicality of MARS on low-power edge devices.
Sizhe An, Ümit Y. Ogras
ACM Trans. Embed. Comput. Syst.1
2019 An Ultra-Low Energy Human Activity Recognition Accelerator for Wearable Health Applications
abstract
Human activity recognition (HAR) has recently received significant attention due to its wide range of applications in health and activity monitoring. The nature of these applications requires mobile or wearable devices with limited battery capacity. User surveys show that charging requirement is one of the leading reasons for abandoning these devices. Hence, practical solutions must offer ultra-low power capabilities that enable operation on harvested energy. To address this need, we present the first fully integrated custom hardware accelerator (HAR engine) that consumes 22.4 μJ per operation using a commercial 65 nm technology. We present a complete solution that integrates all steps of HAR , i.e., reading the raw sensor data, generating features, and activity classification using a deep neural network (DNN). It achieves 95% accuracy in recognizing 8 common human activities while providing three orders of magnitude higher energy efficiency compared to existing solutions.
Ganapati Bhat, Yigit Tuncel, Sizhe An, Hyung Gyu Lee, Ümit Y. Ogras
ACM Trans. Embed. Comput. Syst.3