Shihui Guo

dblp:122/6395 · DBLP profile ↗
← Back
68ranked-venue papers
6as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 24 · 17 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 9 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
abstract
Sparse Inertial Measurement Units (IMUs) based human motion capture has gained significant momentum, driven by the adaptation of fundamental AI tools such as recurrent neural networks (RNNs) and transformers that are tailored for temporal and spatial modeling. Despite these achievements, current research predominantly focuses on pipeline and architectural designs, with comparatively little attention given to regularization methods, highlighting a critical gap in developing a comprehensive AI toolkit for this task. To bridge this gap, we propose motion label smoothing, a novel method that adapts the classic label smoothing strategy from classification to the sparse IMU-based motion capture task. Specifically, we first demonstrate that a naive adaptation of label smoothing, including simply blending a uniform vector or a "uniform" motion representation (e.g., dataset-average motion or a canonical T-pose), is suboptimal; and argue that a proper adaptation requires increasing the entropy of the smoothed labels. Second, we conduct a thorough analysis of human motion labels, identifying three critical properties: 1) Temporal Smoothness, 2) Joint Correlation, and 3) Low-Frequency Dominance, and show that conventional approaches to entropy enhancement (e.g., blending Gaussian noise) are ineffective as they disrupt these properties. Finally, we propose the blend of a novel skeleton-based Perlin noise for motion label smoothing, designed to raise label entropy while satisfying motion properties. Extensive experiments applying our motion label smoothing to three state-of-the-art methods across four real-world IMU datasets demonstrate its effectiveness and robust generalization (plug-and-play) capability.
Zhaorui Meng, Yangqing Hou, Anjun Chen, Shihui Guo, Yipeng Qin
AAAI5
2026 From Performers to Creators: Understanding Retired Women's Perceptions of Technology-Enhanced Dance Performance
abstract
Over 100 million retired women in China engage in dance, but their performances are constrained by limited resources and age-related decline. While interactive dance technologies can enhance artistic expression, existing systems are largely inaccessible to non-professional older dancers. This paper explores how interactive dance technologies can be designed with an age-sensitive approach to support retired women in enhancing their stage performance. We conducted two workshops with community-based retired women dancers, employing interactive dance and LLM-powered video generation probes in co-design activities. Findings indicate that age-sensitive adaptations—such as low-barrier keyword input, motion-aligned visual effects, and participatory scaffolds—lowered technical barriers and fostered a sense of authorship. These features enabled retired women to empower their stage, transitioning from passive recipients of stage design to empowered co-creators of performance. We outline design implications for incorporating interactive dance and artificial intelligence-generated content (AIGC) into the cultural practices of retired women, offering broader strategies for age-sensitive creative technologies.
Danlin Zheng, Xiaoying Wei, Quanyu Zhang, Shihui Guo, Mingming Fan 0001
CHI6
2026 Foreword to special section: AniNex workshop 2025
Anil Bas, Shihui Guo
Comput. Graph.3
2026 MAformer: A transformer with mixed attention for RFID seated posture recognition
Luo Xiaoxiangxin, Lan Hao, Lvqing Yang, Yishu Qiu, Shihui Guo, Liu Zhipeng
Expert Syst. Appl.7
2026 RF-OnlineHAR: A Real-Time RFID Activity Recognition Framework With Memory-Augmented Transformer
abstract
Radio Frequency Identification (RFID)-based human activity recognition (HAR) holds strong potential for ubiquitous sensing due to its capabilities of identity-awareness, privacy-preservation, and long-term passive monitoring. However, most existing methods are constrained to offline classification, which significantly limits their deployment in real-time scenarios. In this paper, we propose RF-OnlineHAR, a real-time HAR framework that enables online, frame-by-frame recognition directly on streaming RFID signals. To better capture short-term temporal characteristics, we introduce a novel Tag Activity Representation Frame (TARF) that jointly encodes the dynamics of physical-layer signals and the irregularity of tag responses. Furthermore, we present the Memory-Augmented Streaming Transformer Encoder (MASTE), which integrates local attention with dynamic memory to capture long-range dependencies across sliding windows, enabling context-aware real-time recognition. Experimental results demonstrate that RF-OnlineHAR achieves 98.61% accuracy on the offline CWNU-RDA dataset, outperforming existing state-of-the-art methods, and attains 94.72% mAP on RF-StreamHAR, a dedicated real-time dataset constructed for continuous HAR evaluation. The proposed framework contributes toward bridging the gap between algorithmic development and practical deployment of RFID-based activity recognition systems, and shows promising potential for applications in areas such as smart healthcare and eldercare.
Yishu Qiu, Lvqing Yang, Shaoqin Shen, Shihui Guo
IEEE Internet Things J.7
2025 LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data
abstract
LiDAR-based semantic scene understanding is an important module in the modern autonomous driving perception stack. However, identifying outlier points in a LiDAR point cloud is challenging as LiDAR point clouds lack semantically-rich information. While former SOTA methods adopt heuristic architectures, we revisit this problem from the perspective of Selective Classification, which introduces a selective function into the standard closed-set classification setup. Our solution is built upon the basic idea of abstaining from choosing any inlier categories but learns a point-wise abstaining penalty with a margin-based loss. Apart from learning paradigms, synthesizing outliers to approximate unlimited real outliers is also critical, so we propose a strong synthesis pipeline that generates outliers originated from various factors: object categories, sampling patterns and sizes. We demonstrate that learning different abstaining penalties, apart from point-wise penalty, for different types of (synthesized) outliers can further improve the performance. We benchmark our method on SemanticKITTI and nuScenes and achieve SOTA results.
Shaocong Xu, Pengfei Li 0007, Qianpu Sun, Yang Li 0178, Shihui Guo, Kehua Sheng, Bo Zhang 0106, Li Jiang 0009, Hao Zhao 0002
AAAI6
2025 FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial Sensors
abstract
CHI ’25, Yokohama, Japan
Ruonan Zheng, Jiawei Fang, Xiaoxia Gao, Chengxu Zuo, Shihui Guo, Yiyue Luo
CHI6
2025 MODA: Motion-Drift Augmentation for Inertial Human Motion Analysis
abstract
While data augmentation (DA) has been extensively studied in computer vision, its application to Inertial Measurement Unit (IMU) signals remains largely unexplored, despite IMUs’ growing importance in human motion analysis. In this paper, we present the first systematic study of IMU-specific data augmentation, beginning with a comprehensive analysis that identifies three fundamental properties of IMU signals: their time-series nature, inherent multimodality (rotation and acceleration) and motion-consistency characteristics. Through this analysis, we demonstrate the limitations of applying conventional time-series augmentation techniques to IMU data. We then introduce Motion-Drift Augmentation (MODA), a novel technique that simulates the natural displacement of body-worn IMUs during motion. We evaluate our approach across five diverse datasets and five deep learning settings, including i) fully-supervised, ii) semi-supervised, iii) domain adaptation, iv) domain generalization and v) few-shot learning for both Human Action Recognition (HAR) and Human Pose Estimation (HPE) tasks. Experimental results show that our proposed MODA consistently outperforms existing augmentation methods, with semi-supervised learning performance approaching state-of-the-art fully-supervised methods.
Yinghao Wu, Shihui Guo, Yipeng Qin
CVPR2
2025 MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic Disturbances
abstract
This paper proposes a novel method called MagShield, designed to address the issue of magnetic interference in sparse inertial motion capture (MoCap) systems. Existing Inertial Measurement Unit (IMU) systems are prone to orientation estimation errors in magnetically disturbed environments, limiting their practical application in real-world scenarios. To address this problem, MagShield employs a "detect-then-correct" strategy, first detecting magnetic disturbances through multi-IMU joint analysis, and then correcting orientation errors using human motion priors. MagShield can be integrated with most existing sparse inertial MoCap systems, improving their performance in magnetically disturbed environments. Experimental results demonstrate that MagShield significantly enhances the accuracy of motion capture under magnetic interference and exhibits good compatibility across different sparse inertial MoCap systems.
Yunzhe Shao, Xinyu Yi, Shihui Guo, Jun-Hai Yong, Feng Xu 0005
ICCV4
2025 DistMovGen: A Knowledge Distillation-Enhanced Framework for Real-Time Virtual Character Motion Generation
Xiangren Shi, Fred Charles, Shihui Guo
ICXR4
2025 ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion Capture
abstract
Sparse inertial measurement units (IMUs) provide a portable, low-cost solution for human motion tracking but struggle with error accumulation from drift and sensor noise when estimating joint position through time-based linear acceleration integration (i.e., indirect measurement). To address this, we propose ToF-IP, a novel 3D full-body pose estimation system that integrates Time-of-Flight (ToF) sensors with sparse IMUs. The distinct advantage of our approach is that ToF sensors provide direct distance measurements, effectively mitigating error accumulation without relying on indirect time-based integration. From a hardware perspective, we maintain the portability of existing solutions by attaching ToF sensors to selected IMUs with a negligible volume increase of just 3\%. On the software side, we introduce two novel techniques to enhance multi-sensor integration: (i) a Node-Centric Data Integration strategy that leverages a Transformer encoder to explicitly model both intra-node and inter-node data integration by treating each sensing node as a token; and (ii) a Dynamic Spatial Positional Encoding scheme that encodes the continuously changing spatial positions of wearable nodes as motion-conditioned functions, enabling the model to better capture human body dynamics in the embedding space.Additionally, we contribute a 208-minute human motion dataset from 10 participants, including synchronized IMU-ToF measurements and ground-truth from optical tracking. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches such as PNP, achieving superior accuracy in tracking complex and slow motions like Tai Chi, which remains challenging for inertial-only methods.
Shifan Jiang, Yangqing Hou, Chengxu Zuo, Shihui Guo, Yipeng Qin
NeurIPS6
2025 WristSketcher: Creating 2D Dynamic Sketches in AR With a Sensing Wristband
abstract
Restricted by the limited interaction area of native AR glasses, creating sketches is a challenge in it. Existing solutions attempt to use mobile devices (e.g., tablets) or mid-air hand gestures to expand the interactive spaces and as the 2D/3D sketching input interfaces for AR glasses. Between them, mobile devices allow for accurate sketching but are often heavy to carry. Sketching with bare hands is zero-burden but can be inaccurate due to arm instability. In addition, mid-air sketching can easily lead to social misunderstandings and its prolonged use can cause arm fatigue. In this work, we present WristSketcher, a new AR system based on a flexible sensing wristband that enables users to place multiple virtual plane canvases in the real environment and create 2D dynamic sketches based on them, featuring an almost zero-burden authoring model for accurate and comfortable sketch creation in real-world scenarios. Specifically, we streamlined the interaction space from the mid-air to the surface of a lightweight sensing wristband, and implemented AR sketching and associated interaction commands by developing a gesture recognition method based on the sensing pressure points. We designed a set of interactive gestures consisting of Long Press, Tap and Double Tap based on a heuristic study involving 26 participants. These gestures are correspondingly mapped to various command interactions using a combination of multi-touch and hotspots. Moreover, we endow our WristSketcher with the ability of animation creation, allowing it to create dynamic and expressive sketches. Experimental results demonstrate that our WristSketcher (i) recognizes users’ gesture interactions with a high accuracy of 95.9%; (ii) achieves higher sketching accuracy than Freehand sketching; (iii) achieves high user satisfaction in ease of use, usability and functionality; and (iv) shows innovation potentials in art creation, memory aids, and entertainment applications.
Enting Ying, Tianyang Xiong, Gaoxiang Zhu, Ming Qiu, Yipeng Qin, Shihui Guo
Int. J. Hum. Comput. Interact.6
2025 Fed-Siamese: Few-Shot RFID Human Activity Recognition With Federated Meta-Learning
abstract
In recent years, wireless radio-frequency identification (RFID) technology has shown great potential in human activity recognition (HAR) due to its passive sensing capability, low deployment cost, and strong privacy protection. However, before its widespread application in real-world scenarios, challenges such as heterogeneous data distributions, scarcity of labeled samples, and data security must be addressed. In this paper, we propose Fed-Siamese, a federated meta-learning based HAR system designed to overcome these challenges through several novel contributions. To address the challenge of data distribution heterogeneity across clients, we first model each client’s activity recognition task as an independent learning problem. Building on this, we propose a federated training algorithm and a trusted parameter aggregation strategy based on KL divergence, all within the framework of federated learning and the Reptile meta-learning optimization algorithm. Our approach enables the distributed training of a Siamese network, thereby mitigating the risk of data leakage inherent in centralized batch training. Furthermore, we design a personalized fusion fine-tuning algorithm to balance adaptation efficiency and recognition performance, enabling the system to quickly adapt to new clients using a few labeled samples. Experimental results show that the proposed system outperforms the baselines in recognition performance on both RFID activity datasets, while maintaining stable model convergence under byzantine and backdoor attacks.
Qianwen Mao, Lvqing Yang, Yishu Qiu, Shihui Guo
IEEE Internet Things J.7
2025 TOPIC: A Parallel Association Paradigm for Multi-Object Tracking Under Complex Motions and Diverse Scenes
abstract
Video data and algorithms have been driving advances in multi-object tracking (MOT). While existing MOT datasets focus on occlusion and appearance similarity, complex motion patterns are widespread yet overlooked. To address this issue, we introduce a new dataset called BEE24 to highlight complex motions. Identity association algorithms have long been the focus of MOT research. Existing trackers can be categorized into two association paradigms: single-feature paradigm (based on either motion or appearance feature) and serial paradigm (one feature serves as secondary while the other is primary). However, these paradigms are incapable of fully utilizing different features. In this paper, we propose a parallel paradigm and present the Two rOund Parallel matchIng meChanism (TOPIC) to implement it. The TOPIC leverages both motion and appearance features and can adaptively select the preferable one as the assignment metric based on motion level. Moreover, we provide an Attention-based Appearance Reconstruction Module (AARM) to reconstruct appearance feature embeddings, thus enhancing the representation of appearance features. Comprehensive experiments show that our approach achieves state-of-the-art performance on four public datasets and BEE24. Moreover, BEE24 challenges existing trackers to track multiple similar-appearing small objects with complex motions over long periods, which is critical in real-world applications such as beekeeping and drone swarm surveillance. Notably, our proposed parallel paradigm surpasses the performance of existing association paradigms by a large margin, e.g., reducing false negatives by 6% to 81% compared to the single-feature association paradigm. The introduced dataset and association paradigm in this work offer a fresh perspective for advancing the MOT field. The source code and dataset are available at https://github.com/holmescao/TOPICTrack.
Xiaoyan Cao, Yiyao Zheng, Yao Yao 0006, Hua-Peng Qin, Shihui Guo
IEEE Trans. Image Process.6
2025 Shape-aware Inertial Poser: Motion Tracking for Humans with Diverse Shapes Using Sparse Inertial Sensors
abstract
Human motion capture with sparse inertial sensors has gained significant attention recently. However, existing methods almost exclusively rely on a template adult body shape to model the training data, which poses challenges when generalizing to individuals with largely different body shapes (such as a child). This is primarily due to the variation in IMU-measured acceleration caused by changes in body shape. To fill this gap, we propose Shape-aware Inertial Poser (SAIP), the first solution considering body shape differences in sparse inertial-based motion capture. Specifically, we decompose the sensor measurements related to shape and pose in order to effectively model their joint correlations. Firstly, we train a regression model to transfer the IMU-measured accelerations of a real body to match the template adult body model, compensating for the shape-related sensor measurements. Then, we can easily follow the state-of-the-art methods to estimate the full body motions of the template-shaped body. Finally, we utilize a second regression model to map the joint velocities back to the real body, combined with a shape-aware physical optimization strategy to calculate global motions on the subject. Furthermore, our method relies on body shape awareness, introducing the first inertial shape estimation scheme. This is accomplished by modeling the shape-conditioned IMU-pose correlation using an MLP-based network. To validate the effectiveness of SAIP, we also present the first IMU motion capture dataset containing individuals of different body sizes. This dataset features 10 children and 10 adults, with heights ranging from 110 cm to 190 cm, and a total of 400 minutes of paired IMU-Motion samples. Extensive experimental results demonstrate that SAIP can effectively handle motion capture tasks for diverse body shapes. The code and dataset are available at https://github.com/yinlu5942/SAIP .
Ziying Shi, Yinghao Wu, Xinyu Yi, Feng Xu 0005, Shihui Guo
ACM Trans. Graph.6
2025 Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion Capture
abstract
In this paper, we propose a novel dynamic calibration method for sparse inertial motion capture systems, which is the first to break the restrictive absolute static assumption in IMU calibration, i.e., the coordinate drift R G′ G and measurement offset R BS remain constant during the entire motion, thereby significantly expanding their application scenarios. Specifically, we achieve real-time estimation of R G′ G and R BS under two relaxed assumptions: i) the matrices change negligibly in a short time window; ii) the human movements/IMU readings are diverse in such a time window. Intuitively, the first assumption reduces the number of candidate matrices, and the second assumption provides diverse constraints, which greatly reduces the solution space and allows for accurate estimation of R G′ G and R BS from a short history of IMU readings in real time. To achieve this, we created synthetic datasets of paired R G′ G , R BS matrices and IMU readings, and learned their mappings using a Transformer-based model. We also designed a calibration trigger based on the diversity of IMU readings to ensure that assumption ii) is met before applying our method. To our knowledge, we are the first to achieve implicit IMU calibration (i.e., seamlessly putting IMUs into use without the need for an explicit calibration process), as well as the first to enable long-term and accurate motion capture using sparse IMUs. The code and dataset are available at https://github.com/ZuoCX1996/TIC.
Chengxu Zuo, Xiangren Shi, Xinyu Yi, Feng Xu 0005, Shihui Guo, Yipeng Qin
ACM Trans. Graph.9
2025 AudioGest: Gesture-Based Interaction for Virtual Reality Using Audio Devices
abstract
Current virtual reality (VR) system takes gesture interaction based on camera, handle and touch screen as one of the mainstream interaction methods, which can provide accurate gesture input for it. However, limited by application forms and the volume of devices, these methods cannot extend the interaction area to such surfaces as walls and tables. To address the above challenge, we propose AudioGest, a portable, plug-and-play system that detects the audio signal generated by finger tapping and sliding on the surface through a set of microphone devices without extensive calibration. First, an audio synthesis-recognition pipeline based on micro-contact dynamics simulation is constructed to generate modal audio synthesis from different materials and physical properties. Then the accuracy and effectiveness of the synthetic audio are verified by mixing the synthetic audio with real audio proportionally as the training sets. Finally, a series of desktop office applications are developed to demonstrate the application potential of AudioGest's scalability and versatility in VR scenarios.
Yi Xiao 0009, Mingwei Hu, Hao Sha 0004, Shining Ma, Boyu Gao 0003, Shihui Guo, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.7
2025 Diverse Motion In-Betweening From Sparse Keyframes With Dual Posture Stitching
abstract
In-betweening is a technique for generating transitions given start and target character states. The majority of existing works require multiple (often 10) frames as input, which are not always available. In addition, they produce results that lack diversity, which may not fulfill artists' requirements. Addressing these gaps, our work deals with a focused yet challenging problem: generating diverse and high-quality transitions given exactly two frames (only the start and target frames). To cope with this challenging scenario, we propose a bi-directional motion generation and stitching scheme which generates forward and backward transitions from the start and target frames with two adversarial autoregressive networks, respectively, and stitches them midway between the start and target frames. In contrast to stitching at the start or target frames, where the ground truth cannot be altered, there is no strict midway ground truth. Thus, our method can capitalize on this flexibility and generate high-quality and diverse transitions simultaneously. Specifically, we employ conditional variational autoencoders (CVAEs) to implement our autoregressive networks and propose a novel stitching loss to stitch the bi-directional generated motions around the midway point. Extensive experiments demonstrate that our method achieves higher motion quality and more diverse results than existing methods on the LaFAN1, Human3.6m and AMASS datasets.
Tianxiang Ren, Jubo Yu, Shihui Guo, Yutao Ouyang, Zijiao Zeng, Yazhan Zhang, Yipeng Qin
IEEE Trans. Vis. Comput. Graph.3
2025 VisMocap: Interactive visualization and analysis for multi-source motion capture data
abstract
With the rapid advancement of artificial intelligence, research on enabling computers to assist humans in achieving intelligent augmentation—thereby enhancing the accuracy and efficiency of information perception and processing—has been steadily evolving. Among these developments, innovations in human motion capture technology have been emerging rapidly, leading to an increasing diversity in motion capture data types. This diversity necessitates the establishment of a unified standard for multi-source data to facilitate effective analysis and comparison of their capability to represent human motion. Additionally, motion capture data often suffer from significant noise, acquisition delays, and asynchrony, making their effective processing and visualization a critical challenge. In this paper, we utilized data collected from a prototype of flexible fabric-based motion capture clothing and optical motion capture devices as inputs. Time synchronization and error analysis between the two data types were conducted, individual actions from continuous motion sequences were segmented, and the processed results were presented through a concise and intuitive visualization interface. Finally, we evaluated various system metrics, including the accuracy of time synchronization, data fitting error from fabric resistance to joint angles, precision of motion segmentation, and user feedback.
Lishuang Zhan, Rongting Li, Juncong Lin, Shihui Guo
Vis. Informatics5
2024 Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear Jacket
abstract
Existing wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing, limiting their application in everyday life. In this paper, we introduce Loose Inertial Poser, a novel motion capture solution with high wearing comfortableness, by integrating four Inertial Measurement Units (IMUs) into a loose-wear jacket. Specifically, we address the challenge of scarce loose-wear IMU training data by proposing a Secondary Motion AutoEncoder (SeMo-AE) that learns to model and synthesize the effects of secondary motion between the skin and loose clothing on IMU data. SeMo-AE is leveraged to generate a diverse synthetic dataset of loose-wear IMU data to augment training for the pose estimation network and significantly improve its accuracy. For validation, we collected a dataset with various subjects and 2 wearing styles (zipped and unzipped). Experimental results demonstrate that our approach maintains high-quality real-time posture estimation even in loose-wear scenarios. Our dataset and code are available at: https://github.com/ZuoCX1966/Loose-Inertial-Poser
Chengxu Zuo, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu 0005, Yipeng Qin
CVPR4
2024 SuDA: Support-based Domain Adaptation for Sim2Real Hinge Joint Tracking with Flexible Sensors
abstract
Flexible sensors hold promise for human motion capture (MoCap), offering advantages such as wearability, privacy preservation, and minimal constraints on natural movement. However, existing flexible sensor-based MoCap methods rely on deep learning and necessitate large and diverse labeled datasets for training. These data typically need to be collected in MoCap studios with specialized equipment and substantial manual labor, making them difficult and expensive to obtain at scale. Thanks to the high-linearity of flexible sensors, we address this challenge by proposing a novel Sim2Real solution for hinge joint tracking based on domain adaptation, eliminating the need for labeled data yet achieving comparable accuracy to supervised learning. Our solution relies on a novel Support-based Domain Adaptation method, namely SuDA, which aligns the supports of the predictive functions rather than the instance-dependent distributions between the source and target domains. Extensive experimental results demonstrate the effectiveness of our method and its superiority overstate-of-the-art distribution-based domain adaptation methods in our task.
Jiawei Fang, Haishan Song, Chengxu Zuo, Xiaoxia Gao, Xiaowei Chen 0017, Shihui Guo, Yipeng Qin
ICML6
2024 TeleMotion: A Realtime Humanoid Teleoperation System with Motion Capture
Jiabao Gan, Shihui Guo, Xiangren Shi
ICXR2
2024 A Survey of Smart Wearable Devices For Motion Capture
Shihui Guo, Caihua Huang
ICXR2
2024 SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground Tactility
abstract
Estimating 3D human poses from monocular images is an important research area with many practical applications. However, the depth ambiguity of 2D solutions limits their accuracy in actions where occlusion exits or where slight centroid shifts can result in significant 3D pose variations. In this paper, we introduce a novel multimodal approach to mitigate the depth ambiguity inherent in monocular solutions by integrating spatial-aware pressure information. We first establish a data collection system with a pressure mat and a monocular camera, and construct a large-scale multimodal human activity dataset comprising over 600,000 frames of motion data. Utilizing this dataset, we propose a pressure image reconstruction network to extract pressure priors from monocular images. Subsequently, we introduce a Transformer-based multimodal pose estimation network to combine pressure priors with monocular images, achieving a world mean per joint position error of 51.6mm, outperforming state-of-the-art methods. Extensive experiments demonstrate the effectiveness of our multimodal 3D human pose estimation method across various actions and joints, highlighting the significance of spatial-aware pressure in improving the accuracy of monocular-vision-based methods. Our dataset is available at: https://github.com/LishuangZhan/SATPose.
Lishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo, Boyu Gao 0003, Yipeng Qin
ACM Multimedia4
2024 Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and Modulation
abstract
Transformer models excel at capturing long-range dependencies in sequential data, but lack explicit mechanisms to leverage structural patterns inherent in fixed-length input sequences. In this paper, we propose a novel sequence structure learning and modulation approach that endows Transformers with the ability to model and utilize such fixed-sequence structural properties for improved performance on inertial pose estimation tasks. Specifically, our method introduces a Sequence Structure Module (SSM) that utilizes structural information of fixed-length inertial sensor readings to adjust the input features of transformers. Such structural information can either be acquired by learning or specified based on users' prior knowledge. To justify the prospect of our approach, we show that i) injecting spatial structural information of IMUs/joints learned from data improves accuracy, while ii) injecting temporal structural information based on smooth priors reduces jitter (i.e., improves steadiness), in a spatial-temporal transformer solution for inertial pose estimation. Extensive experiments across multiple benchmark datasets demonstrate the superiority of our approach against state-of-the-art methods and has the potential to advance the design of the transformer architecture for fixed-length sequences.
Yinghao Wu, Chaoran Wang, Shihui Guo, Yipeng Qin
NeurIPS4
2024 DSteganoM: Deep steganography for motion capture data
Qi Wen Gan, Wei-Chuen Yau, Yee Siang Gan, Md. Iftekhar Salam, Shihui Guo, Chin-Chen Chang 0001, Yubing Wu, Luchen Zhou
Expert Syst. Appl.5
2024 RehabFAB: design investigation and needs assessment of displacement-orientated fabric wearable sensors for rehabilitation
Xiaowei Chen 0017, Shihui Guo, Juncong Lin, Minghong Liao, Hongli Fan, Guoliang Luo
Multim. Tools Appl.3
2024 Multi-Label Action Anticipation for Real-World Videos With Scene Understanding
abstract
With human action anticipation becoming an essential tool for many practical applications, there has been an increasing trend in developing more accurate anticipation models in recent years. Most of the existing methods target standard action anticipation datasets, in which they could produce promising results by learning action-level contextual patterns. However, the over-simplified scenarios of standard datasets often do not hold in reality, which hinders them from being applied to real-world applications. To address this, we propose a scene-graph-based novel model SEAD that learns the action anticipation at the high semantic level rather than focusing on the action level. The proposed model is composed of two main modules, 1) the scene prediction module, which predicts future scene graphs using a grammar dictionary, and 2) the action anticipation module, which is responsible for predicting future actions with an LSTM network by taking as input the observed and predicted scene graphs. We evaluate our model on two real-world video datasets (Charades and Home Action Genome) as well as a standard action anticipation dataset (CAD-120) to verify its efficacy. The experimental results show that SEAD is able to outperform existing methods by large margins on the two real-world datasets and can also yield stable predictions on the standard dataset at the same time. In particular, our proposed model surpasses the state-of-the-art methods with mean average precision improvements consistently higher than 65% on the Charades dataset and an average improvement of 40.6% on the Home Action Genome dataset.
Xiucheng Li, Weijun Zhuang, Shihui Guo, Zhijun Li 0002
IEEE Trans. Image Process.5
2024 Full-body Human Motion Reconstruction with Sparse Joint Tracking Using Flexible Sensors
abstract
Human motion tracking is a fundamental building block for various applications including computer animation, human-computer interaction, healthcare, and so on. To reduce the burden of wearing multiple sensors, human motion prediction from sparse sensor inputs has become a hot topic in human motion tracking. However, such predictions are non-trivial as (i) the widely adopted data-driven approaches can easily collapse to average poses, and (ii) the predicted motions contain unnatural jitters. In this work, we address the aforementioned issues by proposing a novel framework which can accurately predict the human joint moving angles from the signals of only four flexible sensors, thereby achieving the tracking of human joints in multi-degrees of freedom. Specifically, we mitigate the collapse to average poses by implementing the model with a Bi-LSTM neural network that makes full use of short-time sequence information; we reduce jitters by adding a median pooling layer to the network, which smooths consecutive motions. Although being bio-compatible and ideal for improving the wearing experience, the flexible sensors are prone to aging which increases prediction errors. Observing that the aging of flexible sensors usually results in drifts of their resistance ranges, we further propose a novel dynamic calibration technique to rescale sensor ranges, which further improves the prediction accuracy. Experimental results show that our method achieves a low and stable tracking error of 4.51 degrees across different motion types with only four sensors.
Xiaowei Chen 0017, Lishuang Zhan, Shihui Guo, Qunsheng Ruan, Guoliang Luo, Minghong Liao, Yipeng Qin
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Enable Natural Tactile Interaction for Robot Dog based on Large-format Distributed Flexible Pressure Sensors
abstract
Touch is an important channel for human-robot interaction, while it is challenging for robots to recognize human touch accurately and make appropriate responses. In this paper, we design and implement a set of large-format distributed flexible pressure sensors on a robot dog to enable natural human-robot tactile interaction. Through a heuristic study, we sorted out 81 tactile gestures commonly used when humans interact with real dogs and 44 dog reactions. A gesture classification algorithm based on ResNet is proposed to recognize these 81 human gestures, and the classification accuracy reaches 98.7%. In addition, an action prediction algorithm based on Transformer is proposed to predict dog actions from human gestures, reaching a 1-gram BLEU score of 0.87. Finally, we compare the tactile interaction with the voice interaction during a freedom human-robot-dog interactive playing study. The results show that tactile interaction plays a more significant role in alleviating user anxiety, stimulating user excitement and improving the acceptability of robot dogs.
Lishuang Zhan, Yancheng Cao, Qitai Chen, Haole Guo, Jiasi Gao, Yiyue Luo, Shihui Guo, Guyue Zhou, Jiangtao Gong
ICRA7
2023 Self-Adaptive Motion Tracking against On-body Displacement of Flexible Sensors
abstract
Flexible sensors are promising for ubiquitous sensing of human status due to their flexibility and easy integration as wearable systems. However, on-body displacement of sensors is inevitable since the device cannot be firmly worn at a fixed position across different sessions. This displacement issue causes complicated patterns and significant challenges to subsequent machine learning algorithms. Our work proposes a novel self-adaptive motion tracking network to address this challenge. Our network consists of three novel components: i) a light-weight learnable Affine Transformation layer whose parameters can be tuned to efficiently adapt to unknown displacements; ii) a Fourier-encoded LSTM network for better pattern identification; iii) a novel sequence discrepancy loss equipped with auxiliary regressors for unsupervised tuning of Affine Transformation parameters.
Chengxu Zuo, Jiawei Fang, Shihui Guo, Yipeng Qin
NeurIPS3
2023 Computational Design of Wiring Layout on Tight Suits with Minimal Motion Resistance
abstract
An increasing number of electronics are directly embedded on the clothing to monitor human status (e.g., skeletal motion) or provide haptic feedback. A specific challenge to prototype and fabricate such a clothing is to design the wiring layout, while minimizing the intervention to human motion. We address this challenge by formulating the topological optimization problem on the clothing surface as a deformation-weighted Steiner tree problem on a 3D clothing mesh. Our method proposed an energy function for minimizing strain energy in the wiring area under different motions, regularized by its total length. We built the physical prototype to verify the effectiveness of our method and conducted user study with participants of both design experts and smart cloth users. On three types of commercial products of smart clothing, the optimized layout design reduced wire strain energy by an average of 77% among 248 actions compared to baseline design, and 18% over the expert design.
Kai Wang 0107, Yinping Zheng, Da Zhou, Shihui Guo, Yipeng Qin, Xiaohu Guo
SIGGRAPH Asia5
2023 Learning Chinese Calligraphy in VR With Sponge-Enabled Haptic Feedback
abstract
Abstract Nowadays, virtual reality (VR) is becoming an important technique for various educational subjects. However, Chinese calligraphy, as a unique artistic form, remains under-explored in terms of learning in a VR configuration. This deficiency is largely due to the challenge to render delicate haptic feedback of pen and brush during the process of writing. To achieve the purpose of haptic rendering, existing works mostly use the professional device (e.g. Phantom), which is expensive and not accessible to common users. Our work presents a novel yet simple approach to render haptic feedback for Chinese calligraphy in VR by using soft and deformable sponge as the medium between the handheld controller and writing surface. We compared three different feedback configurations using on-device vibration and sponge-enabled haptic feedback against the baseline configuration with no force feedback. Based on both the qualitative and quantitative results from user studies, we found that sponge-based haptic feedback not only provided a comfort experience of interactive virtual writing but also accelerated the learning performance of novices. Our approach is low cost, scalable and produces realistic user experience, which offers an alternative solution for future development of training systems for virtual Chinese calligraphy.
Guoliang Luo, Tingsong Lu, Haibin Xia, Shicong Hu, Shihui Guo
Interact. Comput.5
2023 Struct2Hair: A hair shape descriptor for hairstyle modeling
abstract
Abstract In recent years, it becomes possible to extract hair information for hair reconstruction from multiple cameras or monocular camera. Using a single image as the input avoids the high cost setups and complex calibration compared to multiviewed reconstruction. Taking advantage of an extendible hairstyle database, this paper introduced Struct2Hair, a novel single‐viewed hair modelling approach by extracting hair shape descriptor (HSD). The HSD is defined as the fundamental structure‐aware feature, which is a combination of critical shapes in a hairstyle. A complete dataset of critical hair shapes is constructed from a known database of three‐dimensional (3D) hair models. We first analyze the input two‐dimensional (2D) image to extract the orientation information and 2D hair sketch automatically. The extracted information is then used to retrieve the corresponding critical shapes with optimization to build the robust HSD. Finally, the HSD constructs a weighted 3D hair orientation field to guide full‐head hair model generation. Our method can preserve local geometric features of hair and retain the whole shape of the hairstyle globally owing to the HSD, which will benefit further hair editing and stylization.
Wenshu Zhang, Yinyu Nie, Shihui Guo, Jian Chang 0001, Jian J. Zhang 0001, Ruofeng Tong 0001
Comput. Animat. Virtual Worlds3
2023 Creative and Progressive Interior Color Design with Eye-tracked User Preference
abstract
Interior scene colorization is vastly demanded in areas such as personalized architecture design. Existing works either require manual efforts to colorize individual objects or conform to fixed color patterns automatically learned from prior knowledge, whilst neglecting user preference. Quantitatively identifying user preferences is challenging, particularly at the early stage of the design process. The 3D setup also presents new challenges as the inhabitant can observe from any possible viewpoint. We propose a representative view selection method based on visual attention and a progressive preference inference model. We particularly focus on the progressive integration of eye-tracked user preference, which enables the assistance in creativity support and allows the possibility of convergent thinking. A series of user studies have been conducted to validate the effectiveness of the proposed view selection method, preference inference model and the creativity support mechanism.
Shihui Guo, Yubin Shi, Pintong Xiao, Yinan Fu, Juncong Lin, Wei Zeng 0004, Tong-Yee Lee
ACM Trans. Comput. Hum. Interact.1
2022 iGrow: A Smart Agriculture Solution to Autonomous Greenhouse Control
abstract
Agriculture is the foundation of human civilization. However, the rapid increase of the global population poses a challenge on this cornerstone by demanding more food. Modern autonomous greenhouses, equipped with sensors and actuators, provide a promising solution to the problem by empowering precise control for high-efficient food production. However, the optimal control of autonomous greenhouses is challenging, requiring decision-making based on high-dimensional sensory data, and the scaling of production is limited by the scarcity of labor capable of handling this task. With the advances of artificial intelligence (AI), the internet of things (IoT), and cloud computing technologies, we are hopeful to provide a solution to automate and smarten greenhouse control to address the above challenges. In this paper, we propose a smart agriculture solution named iGrow, for autonomous greenhouse control (AGC): (1) for the first time, we formulate the AGC problem as a Markov decision process (MDP) optimization problem; (2) we design a neural network-based simulator incorporated with the incremental mechanism to simulate the complete planting process of an autonomous greenhouse, which provides a testbed for the optimization of control strategies; (3) we propose a closed-loop bi-level optimization algorithm, which can dynamically re-optimize the greenhouse control strategy with newly observed data during real-world production. We not only conduct simulation experiments but also deploy iGrow in real scenarios, and experimental results demonstrate the effectiveness and superiority of iGrow in autonomous greenhouse simulation and optimal control. Particularly, compelling results from the tomato pilot project in real autonomous greenhouses show that our solution significantly increases crop yield (+10.15%) and net profit (+92.70%) with statistical significance compared to planting experts. Our solution opens up a new avenue for greenhouse production. The code is available at https://github.com/holmescao/iGrow.git.
Xiaoyan Cao, Yao Yao 0006, Lanqing Li, Wanpeng Zhang 0002, Zhicheng An, Zhong Zhang 0014, Li Xiao 0009, Shihui Guo, Meihong Wu, Dijun Luo
AAAI8
2022 Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors
abstract
A key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degradations (\eg, downsampling, noise and compression). Most previous works restore such missing details in the image space. To cope with the high diversity of natural images, they either rely on the unstable GANs that are difficult to train and prone to artifacts, or resort to explicit references from high-resolution (HR) images that are usually unavailable. In this work, we propose Feature Matching SR (FeMaSR), which restores realistic HR images in a much more compact feature space. Unlike image-space methods, our FeMaSR restores HR images by matching distorted LR image features to their distortion-free HR counterparts in our pretrained HR priors, and decoding the matched features to obtain realistic HR images. Specifically, our HR priors contain a discrete feature codebook and its associated decoder, which are pretrained on HR images with a Vector Quantized Generative Adversarial Network (VQGAN). Notably, we incorporate a novel semantic regularization in VQGAN to improve the quality of reconstructed images. For the feature matching, we first extract LR features with an LR encoder consisting of several Swin Transformer blocks and then follow a simple nearest neighbour strategy to match them with the pretrained codebook. In particular, we equip the LR encoder with residual shortcut connections to the decoder, which is critical to the optimization of feature matching loss and also helps to complement the possible feature matching errors.Experimental results show that our approach produces more realistic HR images than previous methods. Code will be made publicly available.
Chaofeng Chen, Yipeng Qin, Xiaoming Li 0001, Xiaoguang Han 0001, Shihui Guo
ACM Multimedia7
2022 Augmenting deep land use prediction with randomized simulation
abstract
Abstract Land use information is the basis of various geo‐spatial applications. Traditionally, land use patterns are predicted with agent‐based simulation, suffering from a long convergence process. Deep learning techniques have recently been used for land use classification but not prediction, due to the lack and difficulty of collecting enough training data. This paper proposes a novel paradigm for land use data generation with a randomized simulation strategy. We also design a tailored deep land use prediction model, LUPnet, to demonstrate the usage of the paradigm. Experimental results reveal the effectiveness of our method.
Zhangwu Chen, Lianhui Lin, Shihui Guo, Juncong Lin
Comput. Animat. Virtual Worlds5
2022 C3 Assignment: Camera Cubemap Color Assignment for Creative Interior Design
abstract
Color design for 3D indoor scenes is a challenging problem due to many factors that need to be balanced. Although learning from images is a commonly adopted strategy, this strategy may be more suitable for natural scenes in which objects tend to have relatively fixed colors. For interior scenes consisting mostly of man-made objects, creative yet reasonable color assignments are expected. We propose$C^{3}$C3Assignment, a system providing diverse suggestions for interior color design while satisfying general global and local rules including color compatibility, color mood, contrast, and user preference. We extend these constraints from the image domain to$\mathbb {R}^3$, and formulate 3D interior color design as an optimization problem. The design is accomplished in an omnidirectional manner to ensure a comfortable experience when the inhabitant observes the interior scene from possible positions and directions. We design a surrogate-assisted evolutionary algorithm to efficiently solve the highly nonlinear optimization problem for interactive applications, and investigate the system performance concerning problem complexity, solver convergence, and suggestion diversity. Preliminary user studies have been conducted to validate the rule extension from 2D to 3D and to verify system usability.
Juncong Lin, Pintong Xiao, Yinan Fu, Yubin Shi, Hongran Wang, Shihui Guo, Ying He 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2021 Unsupervised Domain Adaptation for Person Re-identification via Heterogeneous Graph Alignment
abstract
Unsupervised person re-identification (re-ID) is becoming increasingly popular due to its power in real-world systems such as public security and intelligent transportation systems. However, the person re-ID task is challenged by the problems of data distribution discrepancy across cameras and lack of label information. In this paper, we propose a coarse-to-fine heterogeneous graph alignment (HGA) method to find cross-camera person matches by characterizing the unlabeled data as a heterogeneous graph for each camera. In the coarse-alignment stage, we assign a projection for each camera and utilize an adversarial learning based method to align coarse-grained node groups from different cameras into a shared space, which consequently alleviates the distribution discrepancy between cameras. In the fine-alignment stage, we exploit potential fine-grained node groups in the shared space and introduce conservative alignment loss functions to constrain the graph aligning process, resulting in reliable pseudo labels as learning guidance. The proposed domain adaptation framework not only improves model generalization on target domain, but also facilitates mining and integrating the potential discriminative information across different cameras. Extensive experiments on benchmark datasets demonstrate that the proposed approach outperforms the state-of-the-arts.
Minying Zhang, Yidong Li, Shihui Guo, Hongtao Duan 0003, Yimin Long, Yi Jin 0001
AAAI4
2021 Audio-Visual Salient Object Detection
Shuaiyang Cheng, Shihui Guo
ICIC (2)4
2021 Salient object segmentation for image composition: A case study of group dinner photo
Tianxiang Ren, Lianhui Lin, Shihui Guo, Juncong Lin, Minghong Liao, Shujie Deng, Yinyu Nie
Neurocomputing3
2021 Human posture tracking with flexible sensors for motion recognition
abstract
Abstract The integration of conventional clothes with flexible electronics is a promising solution as a future‐generation computing platform. However, the problem of user authentication on this novel platform is still underexplored. This work uses flexible sensors to track human posture and achieves the goal of user authentication. We capture human movement pattern by four stretch sensors around the shoulder and one on the elbow. We introduce the long short‐term memory fully convolutional network (LSTM‐FCN), which directly takes noisy and sparse sensor data as input and verifies its consistency with the user's predefined movement patterns. The method can identify a user by matching movement patterns even if there are large intrapersonal variations. The authentication accuracy of LSTM‐FCN reaches 98.0%, which is 10.7% and 6.5% higher than that of dynamic time warping and dynamic time warping dependent.
Xiaowei Chen 0017, Yong Ma 0005, Shihui Guo, Yipeng Qin, Minghong Liao
Comput. Animat. Virtual Worlds4
2021 Individual-Based Transfer Learning for Dynamic Multiobjective Optimization
abstract
Dynamic multiobjective optimization problems (DMOPs) are characterized by optimization functions that change over time in varying environments. The DMOP is challenging because it requires the varying Pareto-optimal sets (POSs) to be tracked quickly and accurately during the optimization process. In recent years, transfer learning has been proven to be one of the effective means to solve dynamic multiobjective optimization. However, the negative transfer will lead the search of finding the POS to a wrong direction, which greatly reduces the efficiency of solving optimization problems. Minimizing the occurrence of negative transfer is thus critical for the use of transfer learning in solving DMOPs. In this article, we propose a new individual-based transfer learning method, called an individual transfer-based dynamic multiobjective evolutionary algorithm (IT-DMOEA), for solving DMOPs. Unlike existing approaches, it uses a presearch strategy to filter out some high-quality individuals with better diversity so that it can avoid negative transfer caused by individual aggregation. On this basis, an individual-based transfer learning technique is applied to accelerate the construction of an initial population. The merit of the IT-DMOEA method is that it combines different strategies in maintaining the advantages of transfer learning methods as well as avoiding the occurrence of negative transfer; thereby greatly improving the quality of solutions and convergence speed. The experimental results show that the proposed IT-DMOEA approach can considerably improve the quality of solutions and convergence speed compared to several state-of-the-art algorithms based on different benchmark problems.
Min Jiang 0005, Zhenzhong Wang, Shihui Guo, Xing Gao 0004, Kay Chen Tan
IEEE Trans. Cybern.3
2021 A Fast Dynamic Evolutionary Multiobjective Algorithm via Manifold Transfer Learning
abstract
Many real-world optimization problems involve multiple objectives, constraints, and parameters that may change over time. These problems are often called dynamic multiobjective optimization problems (DMOPs). The difficulty in solving DMOPs is the need to track the changing Pareto-optimal front efficiently and accurately. It is known that transfer learning (TL)-based methods have the advantage of reusing experiences obtained from past computational processes to improve the quality of current solutions. However, existing TL-based methods are generally computationally intensive and thus time consuming. This article proposes a new memory-driven manifold TL-based evolutionary algorithm for dynamic multiobjective optimization (MMTL-DMOEA). The method combines the mechanism of memory to preserve the best individuals from the past with the feature of manifold TL to predict the optimal individuals at the new instance during the evolution. The elites of these individuals obtained from both past experience and future prediction will then constitute as the initial population in the optimization process. This strategy significantly improves the quality of solutions at the initial stage and reduces the computational cost required in existing methods. Different benchmark problems are used to validate the proposed algorithm and the simulation results are compared with state-of-the-art dynamic multiobjective optimization algorithms (DMOAs). The results show that our approach is capable of improving the computational speed by two orders of magnitude while achieving a better quality of solutions than existing methods.
Min Jiang 0005, Zhenzhong Wang, Liming Qiu, Shihui Guo, Xing Gao 0004, Kay Chen Tan
IEEE Trans. Cybern.4
2020 Improving Deep Learning based Optical Character Recognition via Neural Architecture Search
abstract
Optical character rcecognition (OCR) is a process of converting images of typed, handwritten or printed text into machine-encoded one. In recent years, the methods represented by deep learning have greatly improved the performance of OCR systems, but the main challenges of such systems are 1) to accurately perform text detection in complex scenes and 2) to identify and set the optimal parameters to optimize the performance of the system. In this paper, we propose an OCR method based on Neural Architecture Search technique, called AutOCR. The characteristic of the proposed method is the automatic design of text detection framework using an evolutionary computation neural architecture search method. This design can not only accurately recognize the text in a complex environment, but also avoid the process of experts participating in parameter adjustment. We compared it with different methods, and the experimental results proved the effectiveness of our method.
Zhenyao Zhao, Min Jiang 0005, Shihui Guo, Zhenzhong Wang, Fei Chao 0001, Kay Chen Tan
CEC3
2020 Sensock: 3D Foot Reconstruction with Flexible Sensors
abstract
Capturing 3D foot models is important for applications such as manufacturing customized shoes and creating clubfoot orthotics. In this paper, we propose a novel prototype, Sensock, to offer a fully wearable solution for the task of 3D foot reconstruction. The prototype consists of four soft stretchable sensors, made from silk fibroin yarn. We identify four characteristic foot girths based on the existing knowledge of foot anatomy, and measure their lengths with the resistance value of the stretchable sensors. A learning-based model is trained offline and maps the foot girths to the corresponding 3D foot shapes. We compare our method with existing solutions using red-green-blue (RGB) or RGBD (RGB-depth) cameras, and show the advantages of our method in terms of both efficiency and accuracy. In the user experiment, we find that the relative error of Sensock is lower than 0.55%. It performs consistently across different trials and is considered comfortable and suitable for long-term wearing.
Hechuan Zhang, Shihui Guo, Juncong Lin, Yating Shi, Yong Ma 0005
CHI3
2020 Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes From a Single Image
abstract
Semantic reconstruction of indoor scenes refers to both scene understanding and object reconstruction. Existing works either address one part of this problem or focus on independent objects. In this paper, we bridge the gap between understanding and reconstruction, and propose an end-to-end solution to jointly reconstruct room layout, object bounding boxes and meshes from a single image. Instead of separately resolving scene understanding and object reconstruction, our method builds upon a holistic scene context and proposes a coarse-to-fine hierarchy with three components: 1. room layout with camera pose; 2. 3D object bounding boxes; 3. object meshes. We argue that understanding the context of each component can assist the task of parsing the others, which enables joint understanding and reconstruction. The experiments on the SUN RGB-D and Pix3D datasets demonstrate that our method consistently outperforms existing methods in indoor layout estimation, 3D object detection and mesh reconstruction.
Yinyu Nie, Xiaoguang Han 0001, Shihui Guo, Yujian Zheng, Jian Chang 0001, Jian J. Zhang 0001
CVPR3
2020 Skeleton-bridged Point Completion: From Global Inference to Local Adjustment
abstract
Point completion refers to complete the missing geometries of objects from partial point clouds. Existing works usually estimate the missing shape by decoding a latent feature encoded from the input points. However, real-world objects are usually with diverse topologies and surface details, which a latent feature may fail to represent to recover a clean and complete surface. To this end, we propose a skeleton-bridged point completion network (SK-PCN) for shape completion. Given a partial scan, our method first predicts its 3D skeleton to obtain the global structure, and completes the surface by learning displacements from skeletal points. We decouple the shape completion into structure estimation and surface reconstruction, which eases the learning difficulty and benefits our method to obtain on-surface details. Besides, considering the missing features during encoding input points, SK-PCN adopts a local adjustment strategy that merges the input point cloud to our predictions for surface refinement. Comparing with previous methods, our skeleton-bridged manner better supports point normal estimation to obtain the full surface mesh beyond point clouds. The qualitative and quantitative experiments on both point cloud and mesh completion show that our approach outperforms the existing methods on various object categories.
Yinyu Nie, Yiqun Lin, Xiaoguang Han 0001, Shihui Guo, Jian Chang 0001, Shuguang Cui, Jian J. Zhang 0001
NeurIPS4
2020 Information hiding in motion data of virtual characters
Shihui Guo, Xing Gao 0004, Minghong Liao, Chin-Chen Chang 0001, Wei-Chuen Yau
Expert Syst. Appl.2
2020 Online tracking of ants based on deep association metrics: method, dataset and evaluation
Xiaoyan Cao, Shihui Guo, Juncong Lin, Wenshu Zhang, Minghong Liao
Pattern Recognit.2
2020 Shallow2Deep: Indoor scene modeling by single image understanding
Yinyu Nie, Shihui Guo, Jian Chang 0001, Xiaoguang Han 0001, Shi-Min Hu 0001, Jian J. Zhang 0001
Pattern Recognit.2
2020 Semi-Supervised Texture Filtering With Shallow to Deep Understanding
abstract
This work proposed a semi-supervised method for automatic texture filtering. Our method leveraged a limited amount of labeled data and a large amount of unlabeled data to train Generative Adversarial Networks (GANs). Separate loss functions were designed for both labeled and unlabeled datasets. Our main contribution is the introduction of knowledge extracted from shallow and deep layers in neural networks. Loss defined within shallow layers preserves the edge, while loss defined within the deep layers identifies the semantic content and conversely removes the small-scale texture variations. This contribution directly addresses the major challenge for texture filtering, distinguishing the structural content from non-structural textures at the pixel level. The extracted information, in our study, improved the content and color consistency before and after the process of filtering, for unlabeled samples in particular. The proposed method offers twofold benefits: first, significant reductions in the amounts of time and effort expended in reconstructing the labeled dataset, especially given the delicate operations required at the pixel level; second, a reduction in over-fitting, in supervised learning with a small amount of labeled data, by utilizing a large amount of unlabeled data. The results confirm that our method can perform comparably with non-learning-based methods, alleviating the demand for the determination of optimal parameter values.
Xing Gao 0004, Shihui Guo, Minghong Liao, Wencheng Wang 0001
IEEE Trans. Image Process.4
2019 Integrating Peridynamics with Material Point Method for Elastoplastic Material Modeling
Yao Lyu, Jinglu Zhang, Jian Chang 0001, Shihui Guo, Jian J. Zhang 0001
CGI4
2019 Accurate and Fast Classification of Foot Gestures for Virtual Locomotion
abstract
This work explores the use of foot gestures for locomotion in virtual environments. Foot gestures are represented as the distribution of plantar pressure and detected by three sparsely-located sensors on each insole. The Long Short-Term Memory model is chosen as the classifier to recognize the performer's foot gesture based on the captured signals of pressure information. The trained classifier directly takes the noisy and sparse input of sensor data, and handles seven categories of foot gestures (stand, walk forward/backward, run, jump, slide left and right) without manual definition of signal features for classifying these gestures. This classifier is capable of recognizing the foot gestures, even with the existence of large sensor-specific, inter-person and intra-person variations. Results show that an accuracy of ~80% can be achieved across different users with different shoe sizes and ~85% for users with the same shoe size. A novel method, Dual-Check Till Consensus, is proposed to reduce the latency of gesture recognition from 2 seconds to 0.5 seconds and increase the accuracy to over 97%. This method offers a promising solution to achieve lower latency and higher accuracy at a minor cost of computation workload. The characteristics of high accuracy and fast classification of our method could lead to wider applications of using foot patterns for human-computer interaction.
JunJun Pan, Zeyong Hu, Juncong Lin, Shihui Guo, Minghong Liao
ISMAR5
2019 Image Composition of Partially Occluded Objects
abstract
Abstract Image composition extracts the content of interest (COI) from a source image and blends it into a target image to generate a new image. In the majority of existing works, the COI is manually extracted and then overlaid on top of the target image. However, in practice, it is often necessary to deal with situations in which the COI is partially occluded by the target image content. In this regard, both tasks of extracting the COI and cropping its occluded part require intensive user interactions, which are laborious and seriously reduce the composition efficiency. This paper addresses the aforementioned challenges by proposing an efficient image composition method. First, we extract the semantic contents of the images by using state‐of‐the‐art deep learning methods. Therefore, the COI can be selected with clicks only, which can greatly reduce the demanded user interactions. Second, according to the user's operations (such as translation or scale) on the COI, we can effectively infer the occlusion relationships between the COI and the contents of the target image. Thus, the COI can be adaptively embedded into the target image without concern about cropping its occluded part. Therefore, the procedures of content extraction and occlusion handling can be significantly simplified, and work efficiency is remarkably improved. Experimental results show that compared to existing works, our method can reduce the number of user interactions to approximately one‐tenth and increase the speed of image composition by more than ten times.
Xuehan Tan, Shihui Guo, Wencheng Wang 0001
Comput. Graph. Forum3
2019 3D sunken relief generation from a single image by feature line enhancement
abstract
Sunken relief is an art form whereby the depicted shapes are sunk into a given flat plane with a shallow overall depth. In this paper, we propose an efficient sunken relief generation algorithm based on a single image by the technique of feature line enhancement. Our method starts from a single image. First, we smoothen the image with morphological operations such as opening and closing operations and extract the feature lines by comparing the values of adjacent pixels. Then we apply unsharp masking to sharpen the feature lines. After that, we enhance and smoothen the local information to obtain an image with less burrs and jaggies. Differential operations are applied to produce the perceptive relief-like images. Finally, we construct the sunken relief surface by triangularization which transforms two-dimensional information into a three-dimensional model. The experimental results demonstrate that our method is simple and efficient.
Meili Wang 0001, Shihui Guo, Jincen Jiang, Hongming Zhang 0002, Jian Chang 0001
Multim. Tools Appl.4
2019 Action snapshot with single pose and viewpoint
Meili Wang 0001, Shihui Guo, Minghong Liao, Dongjian He, Jian Chang 0001, Jian J. Zhang 0001
Vis. Comput.2
2018 Visual saliency-based bas-relief generation with symmetry composition rule
abstract
Abstract This paper presents a novel approach for bas‐relief generation and synthesis. In contrast to previous methods, we divide this problem into two parts: the selection of the best view and arrangement of the relief layout. Taking these into account, we incorporate the visual saliency and photographic composition rules into the bas‐relief generation. Additionally, a nonlinear compression function is used to compress the models, and finally, we implement surface parameterization by directly manipulating the mesh triangles to generate curved surface bas‐relief. We validate our approach through a variety of models. The results indicate that the proposed approach is effective to adapt different types of target surface with topology unchanged. Comparing with conventional methods, our approach is able to effectively produce bas‐relief with a reasonable layout and distinct details.
Meili Wang 0001, Shihui Guo, Jian Chang 0001, Jian J. Zhang 0001
Comput. Animat. Virtual Worlds6
2018 Semantic modeling of indoor scenes with support inference from a single photograph
abstract
Abstract We present an automatic approach for the semantic modeling of indoor scenes based on a single photograph, instead of relying on depth sensors. Without using handcrafted features, we guide indoor scene modeling with feature maps extracted by fully convolutional networks. Three parallel fully convolutional networks are adopted to generate object instance masks, a depth map, and an edge map of the room layout. Based on these high‐level features, support relationships between indoor objects can be efficiently inferred in a data‐driven manner. Constrained by the support context, a global‐to‐local model matching strategy is followed to retrieve the whole indoor scene. We demonstrate that the proposed method can efficiently retrieve indoor objects including situations where the objects are badly occluded. This approach enables efficient semantic‐based scene editing.
Yinyu Nie, Jian Chang 0001, Ehtzaz Chaudhry, Shihui Guo, Philip Andi Smart, Jian J. Zhang 0001
Comput. Animat. Virtual Worlds4
2017 Pose selection for animated scenes and a case study of bas-relief generation
abstract
This paper aims to automate the process of generating a meaningful single still image from a temporal input of scene sequences. The success of our extraction relies on evaluating the optimal pose of characters selection, which should maximize the information conveyed. We define the information entropy of the still image candidates as the evaluation criteria.
Meili Wang 0001, Shihui Guo, Minghong Liao, Dongjian He, Jian Chang 0001, Jian J. Zhang 0001, Zhiyi Zhang 0002
CGI2
2017 Understanding the impact of multimodal interaction using gaze informed mid-air gesture control in 3D virtual objects manipulation
Shujie Deng, Nan Jiang 0006, Jian Chang 0001, Shihui Guo, Jian J. Zhang 0001
Int. J. Hum. Comput. Stud.4
2017 Simulating collective transport of virtual ants
abstract
Abstract This paper simulates the behaviour of collective transport where a group of ants transports an object in a cooperative fashion. Different from humans, the task coordination of collective transport, with ants, is not achieved by direct communication between group individuals, but through indirect information transmission via mechanical movements of the object. This paper proposes a stochastic probability model to model the decision‐making procedure of group individuals and trains a neural network via reinforcement learning to represent the force policy. Our method is scalable to different numbers of individuals and is adaptable to users' input, including transport trajectory, object shape, external intervention, etc. Our method can reproduce the characteristic strategies of ants, such as realign and reposition. The simulations show that with the strategy of reposition, the ants can avoid deadlock scenarios during the task of collective transport.
Shihui Guo, Meili Wang 0001, Gabriel Notman, Jian Chang 0001, Jian J. Zhang 0001, Minghong Liao
Comput. Animat. Virtual Worlds1
2017 Customization and fabrication of the appearance for humanoid robot
Shihui Guo, Hanxiang Xu, Nadia Magnenat-Thalmann, Junfeng Yao
Vis. Comput.1
2015 Adaptive motion synthesis for virtual characters: a survey
Shihui Guo, Richard Southern, Jian Chang 0001, David Greer, Jian J. Zhang 0001
Vis. Comput.1
2014 A novel locomotion synthesis and optimisation framework for insects
Shihui Guo, Jian Chang 0001, Jian J. Zhang 0001
Comput. Graph.1
2014 Locomotion Skills for Insects with Sample-based Controller
abstract
Abstract Natural‐looking insect animation is very difficult to simulate. The fast movement and small scale of insects often challenge the standard motion capture techniques. As for the manual key‐framing or physics‐driven methods, significant amounts of time and efforts are necessary due to the delicate structure of the insect, which prevents practical applications. In this paper, we address this challenge by presenting a two‐level control framework to efficiently automate the modeling and authoring of insects’ locomotion. On the top level, we design a Triangle Placement Engine to automatically determine the location and orientation of insects’ foot contacts, given the user‐defined trajectory and settings, including speed, load, path and terrain etc. On the low‐level, we relate the Central Pattern Generator to the triangle profiles with the assistance of a Controller Look‐Up Table to fast simulate the physically‐based movement of insects. With our approach, animators can directly author insects’ behavior among a wide range of locomotion repertoire, including walking along a specified path or on an uneven terrain, dynamically adjusting to external perturbations and collectively transporting prey back to the nest.
Shihui Guo, Jian Chang 0001, Xiaosong Yang, Wencheng Wang 0001, Jian J. Zhang 0001
Comput. Graph. Forum1
2013 Motion Adaptation With Motor Invariant Theory
abstract
Bipedal walking is not fully understood. Motion generated from methods employed in robotics literature is stiff and is not nearly as energy efficient as what we observe in nature. In this paper, we propose validity conditions for motion adaptation from biological principles in terms of the topology of the dynamic system. This allows us to provide a closed-form solution to the problem of motion adaptation to environmental perturbations. We define both global and local controllers that improve structural and state stability, respectively. Global control is achieved by coupling the dynamic system with a neural oscillator, which preserves the periodic structure of the motion primitive and ensures stability by entrainment. A group action derived from Lie group symmetry is introduced as a local control that transforms the underlying state space while preserving certain motor invariants. We verify our method by evaluating the stability and energy consumption of a synthetic passive dynamic walker and compare this with motion data of a real walker. We also demonstrate that our method can be applied to a variety of systems.
Fangde Liu, Richard Southern, Shihui Guo, Xiaosong Yang, Jian J. Zhang 0001
IEEE Trans. Cybern.3