Weichao Guo

dblp:150/7626 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Simultaneous Decoding of Wrist Angles and Grasp Forces Based on Channel-Wise Cumulative Spike Trains
abstract
Understanding the underlying mechanism of neuromuscular system on motion/force generation is essential for human-machine interfacing. However, simultaneous decoding of wrist angles and grasp forces from neural signals remains an open challenge in the field of neural interfacing. In this study, we proposed a scheme leveraging channel-wise cumulative spike trains (cw-CSTs) of motor units to simultaneously decode wrist angles and grasp forces. Specifically, a spatial spike detection method was utilized to detect cw-CST from surface electromyography, observing as much as possible of motor unit activities. Accordingly, we extracted three neural features to drive the decoders, including a twitch force model-based (cw-MUdrive) and a discharge rate-based (DR-cwCST) neural features derived from cw-CSTs, and DR of motor units (DR-MUST) decomposed by a conventional blind source separation algorithm. Wrist- and hand-specific decoders were built to estimate wrist angles and grasp forces via Gaussian process regression. Experiments were conducted with ten subjects, in which they activated wrist motions and grasp forces concurrently. We evaluated the performance with both accuracy and output stability. Results demonstrated that the cwCST-based neural features outperformed the conventional DR-MUST features with both higher accuracy and stability metrics. Additionally, cw-MUdrive performed better than DR-cwCST in grasp force estimation and comparable to DR-cwCST in wrist angle estimation. The outcome provides an effective solution for simultaneously decoding wrist movements and hand grasp forces, promoting the development of natural control in neural interface.
Yang Yu 0019, Yang Xu 0079, Jiamin Zhao, Dongxuan Li, Weichao Guo, Xinjun Sheng
IEEE J. Biomed. Health Informatics5
2025 MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Device
Chao Wu 0006, Cheng Ji 0002, Li-Pin Chang, Zongwei Zhu, Congming Gao, Weichao Guo, Yanzhi Wang 0001
FAST6
2025 Hierarchical Reinforcement Learning for Articulated Tool Manipulation with Multifingered Hand
abstract
Manipulating articulated tools, such as tweezers or scissors, has rarely been explored in previous research. Unlike rigid tools, articulated tools change their shape dynamically, creating unique challenges for dexterous robotic hands. In this work, we present a hierarchical, goal-conditioned reinforcement learning (GCRL) framework to improve the manipulation capabilities of anthropomorphic robotic hands using articulated tools. Our framework comprises two policy layers: (1) a low-level policy that enables the dexterous hand to manipulate the tool into various configurations for objects of different sizes, and (2) a high-level policy that defines the tool’s goal state and controls the robotic arm for object-picking tasks. We employ an encoder, trained on synthetic pointclouds, to estimate the tool’s affordance states—specifically, how different tool configurations (e.g., tweezer opening angles) enable grasping of objects of varying sizes—from input point clouds, thereby enabling precise tool manipulation. We also utilize a privilege-informed heuristic policy to generate replay buffer, improving the training efficiency of the high-level policy. We validate our approach through real-world experiments, showing that the robot can effectively manipulate a tweezer-like tool to grasp objects of diverse shapes and sizes with a 70.8% success rate. This study highlights the potential of RL to advance dexterous robotic manipulation of articulated tools.
Wei Xu 0040, Yanchao Zhao, Weichao Guo, Xinjun Sheng
IROS3
2024 DACO: Pursuing Ultra-low Power Consumption via DNN-Adaptive CPU-GPU CO-optimization on Mobile Devices
abstract
As Deep Neural Networks (DNNs) become popular in mobile systems, their high computational and memory demands make them major power consumers, especially in limited-budget scenarios. In this paper, we propose DACO, a DNN-Adaptive CPU-GPU CO-optimization technique, to reduce the power consumption of DNNs. First, a resource-oriented classifier is proposed to quantify the computation/memory intensity of DNN models and classify them accordingly. Second, a set of rule-based policies is deduced for achieving the best-suited CPU-GPU system configuration in a coarse-grained manner. Combined with all the rules, a coarse-to-fine CPU-GPU auto-tuning approach is proposed to reach the Pareto-optimal speed and power consumption in DNN inference. Experimental results demonstrate that, compared with the existing approach, DACO could reduce power consumption by up to 71.9% while keeping an excellent DNN inference speed.
Yushu Wu, Chao Wu 0006, Geng Yuan, Yanyu Li, Weichao Guo, Jing Rao, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001
DATE5
2023 Transparent File Deduplication with Reduced Update Cost on Encryption Enabled Mobile Devices
abstract
Data deduplication has been long studied to achieve data reduction. However, deploying deduplication on encryption enabled mobile systems might consume much memory footprint and computation time for hash calculations. Moreover, frequent file updates on deduplicated files could badly degrade the deduplication efficacy due to the increased file-system metadata penalty. Considering the characteristics of mobile devices, an efficient data deduplication method is proposed in this paper. First, it separates the hash calculation into foreground and background stages. The background stage calculates the hash values of potentially duplicate files while the foreground stage quickly hashes the file which is being written using a lightweight hash algorithm. Second, a dual-level node structure is proposed to improve the file update efficacy for deduplicated files, saving more storage space against the file re-splitting. Besides, we implement a superlink call to make deduplication process compatible with file-based encryption. These methods are combined to realize a transparent file deduplication (TFDedup) approach, which eliminates redundant data and reduces the associated cost of file update. Experimental results show that TFDedup succeeds to lower the space consumption when serving file updates by 55.6% and accelerate the deduplication process by 50.3%.
Junbin Ren, Cheng Ji 0002, Weiwei Jin, Weichao Guo, Yajuan Du, Zongwei Zhu
ICPADS5
2018 Feasibility of Wrist-Worn, Real-Time Hand, and Surface Gesture Recognition via sEMG and IMU Sensing
abstract
While most wearable gesture recognition approaches focus on the forearm or fingers, the wrist may be a more suitable location for practical use. We present the design and validation of a real-time gesture recognition wristband based on surface electromyography and inertial measurement unit sensing fusion, which can recognize 8 air gestures and 4 surface gestures with 2 distinct force levels. Ten healthy subjects performed an initial gesture recognition experiment, followed by a second experiment 1 h later and a third experiment 1 day later. Classification accuracies for the initial experiment were 92.6% and 88.8% for air and surface gestures, respectively, and there were no changes in accuracy results during testing 1 h. and 1 day later (p > 0.05). These results demonstrate the feasibility of wrist-based gesture recognition paving the way for potential future integration in to a smart watch or other wrist-worn wearable for intuitive human computer interaction.
Weichao Guo, Haitao Wang 0006, Xinjun Sheng, Peter B. Shull
IEEE Trans. Ind. Informatics3
2017 Toward an Enhanced Human-Machine Interface for Upper-Limb Prosthesis Control With Combined EMG and NIRS Signals
abstract
Advanced myoelectric prosthetic hands are currently limited due to the lack of sufficient signal sources on amputation residual muscles and inadequate real-time control performance. This paper presents a novel human-machine interface for prosthetic manipulation that combines the advantages of surface electromyography (EMG) and near-infrared spectroscopy (NIRS) to overcome the limitations of myoelectric control. Experiments including 13 able-bodied and three amputee subjects were carried out to evaluate both offline classification accuracy (CA) and online performance of the forearm motion recognition system based on three types of sensors (EMG-only, NIRS-only, and hybrid EMG-NIRS). The experimental results showed that both the offline CA and realtime performance for controlling a virtual prosthetic hand were significantly (p <; 0.05) improved by combining EMG and NIRS. These findings suggest that fusion of EMG and NIRS is feasible to improve the control of upper-limb prostheses, without increasing the number of sensor nodes or complexity of signal processing. The outcomes of this study have great potential to promote the development of dexterous prosthetic hands for transradial amputees.
Weichao Guo, Xinjun Sheng, Honghai Liu 0001
IEEE Trans. Hum. Mach. Syst.1
2016 MARS: Mobile Application Relaunching Speed-Up through Flash-Aware Page Swapping
abstract
The approach for fast application relaunching on the current Android system is to cache background applications in memory. This mechanism is limited by the available memory size. In addition, the application state may not be easily recovered. We propose a prototype system, MARS, to enable page swapping and cache more applications. MARS can speed up the application relaunching and restore the application state. As a new page swapping design for optimizing application relaunching, MARS isolates Android runtime Garbage Collection (GC) from page swapping for compatibility and employs several flash-aware techniques for swap-in speedup. Two main components of MARS are page slot allocation and read/write control. Page slot allocation reorganizes page slots in swap area to produce sequential reads and improve the performance of swap-in. Read/Write control addresses the read/write interference issue by reducing concurrent and extra internal writes. Compared to the conventional Linux page swapping, these two components can scale up the read bandwidth up to about 3.8 times. Application tests on a Google Nexus 4 phone show that MARS reduces the launching time of applications by 50 - 80 percent. The modified page swapping mechanism can outperform the conventional Linux page swapping up to four times.
Weichao Guo, Kang Chen 0001, Huan Feng, Yongwei Wu 0001, Rui Zhang 0003
IEEE Trans. Computers1
2015 Bidding for Highly Available Services with Low Price in Spot Instance Market
abstract
Amazon EC2 has built the Spot Instance Marketplace and offers a new type of virtual machine instances called as spot instances. These instances are less expensive but considered failure-prone. Despite the underlying hardware status, if the bidding price is lower than the market price, such an instance will be terminated.
Weichao Guo, Kang Chen 0001, Yongwei Wu 0001
HPDC1
2014 A wireless wearable sEMG and NIRS acquisition system for an enhanced human-computer interface
abstract
Surface electromyography (sEMG) is extensively explored in human-computer interface (HCI); complementary to the electrophysiological activity of the muscles, the hemodynamic information that measured from near infrared spectroscopy (NIRS) is less investigated. Properly combining the sEMG and NIRS would provide a novel approach for HCI applications. This paper presents a multi-channel wireless wearable sEMG and NIRS acquisition system aiming for enhanced human-computer interaction, by providing more information about the muscle activity for subject's motor intention decoding. Extensive tests were carried out to evaluate the system performance. It showed that this novel system proved to be able to capture sEMG signals similar to those of the commercialized sEMG acquisition devices, and had a comparable NIRS sensor performance. Furthermore, simultaneously recording of sEMG and NIRS signals, the system had shown the ability to provide more information about the muscle activities for a better HCI performance. The classification accuracy of 13 hand gesture motions was significantly (P<;0.001) improved by using combined sEMG and NIRS features comparing to sEMG or NIRS features individually, suggesting that the proposed sEMG and NIRS system could be potentially available for an enhanced HCI.
Weichao Guo, Peng-Fei Yao, Xinjun Sheng, Honghai Liu 0001
SMC1
2014 NO2: Speeding up Parallel Processing of Massive Compute-Intensive Tasks
abstract
Large-scale computing frameworks, either tenanted on the cloud or deployed in the high-end local cluster, have become an indispensable software infrastructure to support numerous enterprise and scientific applications. Tasks executed on these frameworks are generally classified into data-intensive and compute-intensive ones. However, most existing frameworks, led by MapReduce, are mainly suitable for data-intensive tasks. Their task schedulers assume that the proportion of data I/O reflects the task progress and state. Unfortunately, this assumption does not apply to most compute-intensive tasks. Due to biased estimation of task progress, traditional frameworks cannot timely cut off outliers and therefore largely prolong execution time when performing compute-intensive tasks. We propose a new framework designed for compute-intensive tasks. By using instrumentation and automatic instrument point selector, our framework estimates the compute-intensive task progress without resorting to data I/O. We employ a clustering method to identify outliers at runtime and perform speculative execution/aborting, speeding up task execution by up to 25%. Moreover, our improvement to bare instrumentation limits overhead within 0.1%, and the aborting-based execution only introduces 10% more average CPU usage. Low overhead and resource consumption make our framework practically usable in the production environment.
Yongwei Wu 0001, Weichao Guo, Jinglei Ren
IEEE Trans. Computers2