Yantong Wang

dblp:192/5064 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OSATG-GPT: Instruction-tuning large language models with open-source atomic tasks in Github
Fanyu Han, Fenglin Bi, Yantong Wang, Mingdong You, Jiaheng Peng, Xiaoya Xia
Expert Syst. Appl.4
2025 Value-Based Proactive Caching for Sensing Data in Vehicular Networks: An Operator's Perspective
abstract
Access to sensing data (SD) is crucial for vehicular networks to ensure safe and efficient transportation services. Given the vast volume of data involved, proactive caching required SD is a pivotal strategy for alleviating network congestion and improving data accessibility. Despite merits, existing studies predominantly address SD caching within a single slot. Therefore, these approaches lack scalability for scenarios involving multi-slots and are not well-suited for network operators who manage resources within a long-term cost budget. Moreover, the oversight of service capacity at caching nodes may result in substantial queuing delays for SD reception. To tackle these limitations, we jointly consider the problem of anchoring SD caching and allocating from an operator's perspective. A value model incorporating both temporal and spacial characteristics is given to estimate the significance of various caching decisions. Subsequently, a stochastic programming model is proposed to optimize the long-term system performance, which is converted into a series of online optimization problem by leveraging the Lyapunov method and linearized via introducing auxiliary variables. To expedite the solution, we provide a binary quantum particle swarm optimization based algorithm with quadratic time complexity. Numerical investigations demonstrate the superiority of proposed algorithms compared with other schemes in terms of energy consumption, response latency, and cache-hit ratio.
Yantong Wang, Jiande Sun 0001
ICC1
2024 Embodied Interaction Design: A Storytelling City Installation
abstract
Embodied interaction is a significant direction within the realm of interaction design. Interaction design does not only enhance efficiency in people’s lives; it focuses on how humans and computer systems interact efficiently and effectively. In embodied interaction, natural interactive behaviors can be employed to enhance experiences. Interaction is now omnipresent; it serves as an intermediary to establish connections and provide reflection, and not just consolidated within products like smartphones and computers. The incorporation of interaction design into physical life, intertwining with human behaviors and spatial environments, poses our contemplation. Thus, under the backdrop of sustainable development goals (SDGs), we adopted the direction of embodied interaction and utilized metaphors to create a tangible installation. We aspire for urban explorers to establish connections with the city in terms of time and space dimensions through interactive experiences, thereby comprehending the multifaceted nature of the city. Throughout the design process, we validated and iterated our designs via prototyping and user testing.
Liyuan Dong, Pin Jia Lai, Yantong Wang
TEI3
2024 EmoTake: Exploring Drivers' Emotion for Takeover Behavior Prediction
abstract
The blossoming semi-automated vehicles allow drivers to engage in various non-driving-related tasks, which may stimulate diverse emotions, thus affecting takeover safety. Though the effects of emotion on takeover behavior have recently been examined, how to effectively obtain and utilize drivers' emotions for predicting takeover behavior remains largely unexplored. We propose EmoTake, a deep learning-empowered system that explores drivers' emotional and physical states to predict takeover readiness, reaction time, and quality. The key enabler is a deep neural framework that extracts drivers' fine-grained body movements from a camera and interprets them into drivers' multi-channel emotional and physical information (e.g., facial expression, and head pose) for prediction. Our study (N = 26) verifies the efficiency of EmoTake and shows that: 1) facial expression benefits prediction; 2) emotions have diverse impacts on takeovers. Our findings provide insights into takeover prediction and in-vehicle emotion regulation.
Yu Gu 0003, Yibing Weng, Yantong Wang, Meng Wang 0037, Guohang Zhuang, Jinyang Huang, Xiaolan Peng, Fuji Ren
IEEE Trans. Affect. Comput.3
2023 Toward Facial Expression Recognition in the Wild via Noise-Tolerant Network
abstract
Facial Expression Recognition (FER) has recently emerged as a crucial area in Human-Computer Interaction (HCI) system for understanding the user’s inner state and intention. However, feature- and label-noise constitute the major challenge for FER in the wild due to the ambiguity of facial expressions worsened by low-quality images. To deal with this problem, in this paper, we propose a simple but effective Facial Expression Noise-tolerant Network (FENN) which explores the inter-class correlations for mitigating ambiguity that usually happens between morphologically similar classes. Specifically, FENN leverages a multivariate normal distribution to model such correlations at the final hidden layer of the neural network to suppress the heteroscedastic uncertainty caused by inter-class label noise. Furthermore, the discriminative ability of deep features is weakened by the subtle differences between expressions and the presence of feature noise. FENN utilizes a feature-noise mitigation module to extract compact intra-class feature representations under feature noise while preserving the intrinsic inter-class relationships. We conduct extensive experiments to evaluate the effectiveness of FENN on both original annotated images and synthetic noisy annotated images from RAF-DB, AffectNet, and FERPlus in-the-wild facial expression datasets. The results show that FENN significantly outperforms state-of-the-art FER methods.
Yu Gu 0003, Huan Yan 0005, Xiang Zhang 0011, Yantong Wang, Yusheng Ji, Fuji Ren
IEEE Trans. Circuits Syst. Video Technol.4
2023 Wital: A COTS WiFi Devices Based Vital Signs Monitoring System Using NLOS Sensing Model
abstract
Vital sign (breathing and heartbeat) monitoring is essential for patient care and sleep disease prevention. Most current solutions are based on wearable sensors or cameras; however, the former could affect sleep quality, while the latter often present privacy concerns. To address these shortcomings, we propose Wital, a contactless vital sign monitoring system based on low-cost and widespread commercial off-the-shelf (COTS) Wi-Fi devices. There are two challenges that need to be overcome. First, the torso deformations caused by breathing/heartbeats are weak. How can such deformations be effectively captured? Second, movements such as turning over affect the accuracy of vital sign monitoring. How can such detrimental effects be avoided? For the former, we propose a non-line-of-sight (NLOS) sensing model for modeling the relationship between the energy ratio of line-of-sight (LOS) to NLOS signals and the vital sign monitoring capability using Ricean K theory and use this model to guide the system construction to better capture the deformations caused by breathing/heartbeats. For the latter, we propose a motion segmentation method based on motion regularity detection that accurately distinguishes respiration from other motions, and we remove periods that include movements such as turning over to eliminate detrimental effects. We have implemented and validated Wital on low-cost COTS devices. The experimental results demonstrate the effectiveness of Wital in monitoring vital signs.
Xiang Zhang 0011, Yu Gu 0003, Huan Yan 0005, Yantong Wang, Mianxiong Dong, Kaoru Ota, Fuji Ren, Yusheng Ji
IEEE Trans. Hum. Mach. Syst.4
2023 Learning From Images: Proactive Caching With Parallel Convolutional Neural Networks
abstract
With the continuous trend of data explosion, delivering packets from data servers to end users causes increased stress on both the fronthaul and backhaul traffic of mobile networks. To mitigate this problem, caching popular content closer to the end-users has emerged as an effective method for reducing network congestion and improving user experience. To find the optimal locations for content caching, many conventional approaches construct various Mixed Integer Linear Programming (MILP) models. However, such methods may fail to support online decision making due to the inherent curse of dimensionality. In this paper, a novel framework for proactive caching is proposed. This framework merges model-based optimization with data-driven techniques by transforming an optimization problem into a grayscale image. For parallel training and simple design purposes, the proposed MILP model is first decomposed into a number of sub-problems and, then, Convolutional Neural Networks (CNNs) are trained to predict content caching locations of these sub-problems. Furthermore, since the MILP model decomposition neglects the network resources (such as caching space and link bandwidth) competition among sub-problems, the CNNs' outputs have the risk to be infeasible solutions. Therefore, two algorithms are provided: the first uses predictions from CNNs as an extra constraint to reduce the number of decision variables; the second employs CNNs' outputs to accelerate local search. Numerical results show that the proposed scheme can reduce 71.6% computation time, whose computation time reaches around 28.9 ms, with only 0.8% additional performance cost compared to the MILP solution, which provides high quality decision making in pseudo real-time.
Yantong Wang, Zhaohui Yang 0001, Walid Saad 0001, Kai-Kit Wong, Vasilis Friderikos
IEEE Trans. Mob. Comput.1
2022 A Minmax Utilization Algorithm for Network Traffic Scheduling of Industrial Robots
abstract
Emerging 5G and beyond wireless industrial virtualized networks are expected to support a significant number of robotic manipulators. Depending on the processes involved, these industrial robots might result in significant volume of multi-modal traffic that will need to traverse the network all the way to the (public/private) edge cloud, where advanced processing, control and service orchestration will be taking place. In this paper, we perform the traffic engineering by capitalizing on the underlying pseudo-deterministic nature of the repetitive processes of robotic manipulators in an industrial environment and propose an integer linear programming (ILP) model to minimize the maximum aggregate traffic in the network. The task sequence and time gap requirements are also considered in the proposed model. To tackle the curse of dimensionality in ILP, we provide a random search algorithm with quadratic time complexity. Numerical investigations reveal that the proposed scheme can reduce the peak data rate up to 53.4% compared with the nominal case where robotic manipulators operate in an uncoordinated fashion, resulting in significant improvement in the utilization of the underlying network resources.
Yantong Wang, Vasilis Friderikos, Sebastian Andraos
ICC1
2022 Mitigating Label-Noise for Facial Expression Recognition in the Wild
abstract
Label-noise constitutes a major challenge for facial expression recognition in the wild due to the ambiguity of facial expressions worsened by low-quality images. To deal with this problem, we propose a simple but effective Label-noise Robust Network (LRN) which explores the inter-class correlations for mitigating ambiguity that usually happens between morphologically similar classes. Specifically, LRN leverages a multivariate normal distribution to model such correlations at the final hidden layer of the neural network to suppress the heteroscedastic uncertainty caused by inter-class label noise. Furthermore, LRN utilizes a confidence-based label-free loss to extract compact intra-class feature representations under label noise while preserving the intrinsic inter-class relationships. Experiments on three in-the-wild facial expression datasets demonstrates the superiority of our method.
Huan Yan 0005, Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Yusheng Ji, Fuji Ren
ICME4
2022 SpiroFi: Contactless Pulmonary Function Monitoring using WiFi Signal
abstract
Human pulmonary function declines with age. Elders, especially those with lung or cardiovascular diseases, yearn for daily lung function tests for timely diagnosis and treatment. However, current clinical spirometers are cumbersome and ex-pensive while home-use portable ones’ accuracy is questionable. Moreover, both kinds require contact measurements and could cause cross infection, especially hazardous for contagious diseases like COVID-19. To this end, we propose SpiroFi, a contactless system that leverages WiFi Channel State Information (CSI) for convenient yet accurate Pulmonary Function Testing (PFT) out of clinic. The key enabler underlying SpiroFi is a set of algorithms that can extract chest wall movement from WiFi signal variations and interpret such information into lung function indices. We have realized SpiroFi on low-cost commodity WiFi devices and tested it in a home-like site where it achieves 2.55% monitoring error over healthy youths. Then, with the Ethics Committee (EC) approval, we conducted a 2-month clinic study in a city hospital over elders with basic diseases. SprioFi still yields 6.05% monitoring error despite elders’ degenerated pulmonary function and body control. Also, the correlation between lung function and age as well as chronic diseases has been revealed, highlighting the importance of daily PFT for the elderly.
Yu Gu 0003, Meng Wang 0001, Peng Zhao 0024, Yantong Wang, Hao Zhou 0001, Yusheng Ji, Celimuge Wu
IWQoS4
2022 Energy-efficient proactive caching with multipath routing
Yantong Wang, Vasilis Friderikos
Comput. Networks1
2022 WiGRUNT: WiFi-Enabled Gesture Recognition Using Dual-Attention Network
abstract
Gestures constitute an important form of nonverbal communication where bodily actions are used for delivering messages alone or in parallel with spoken words. Recently, there exists an emerging trend of WiFi sensing-enabled gesture recognition due to its inherent merits like remote sensing, non-line-of-sight covering, and privacy-friendly. However, current WiFi-based approaches mainly reply on domain-specific training since they don’t know “where to look” and “when to look.” To this end, we propose WiGRUNT, a WiFi-enabled gesture recognition system using dual-attention network, to mimic how a keen human being intercepting a gesture regardless of the environment variations. The key insight is to train the network to dynamically focus on the domain-independent features of a gesture on the WiFi channel state information via a spatial-temporal dual-attention mechanism. WiGRUNT roots in a deep residual network (ResNet) backbone to evaluate the importance of spatial-temporal clues and exploit their inbuilt sequential correlations for fine-grained gesture recognition. We evaluate WiGRUNT on the open Widar3 dataset and show that it significantly outperforms its state-of-the-art rivals by achieving the best-ever performance in-domain or cross-domain.
Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Meng Wang 0001, Huan Yan 0005, Yusheng Ji, Zhi Liu 0002, Jianhua Li 0003, Mianxiong Dong
IEEE Trans. Hum. Mach. Syst.3
2022 Secure User Authentication Leveraging Keystroke Dynamics via Wi-Fi Sensing
abstract
User authentication plays a critical role in access control of a man-machine system, where the knowledge factor, such as a personal identification number, constitutes the most widely used authentication element. However, knowledge factors are usually vulnerable to the spoofing attack. Recently, the inheritance factor, such as fingerprints, emerges as an efficient alternative resilient to malicious users, but it normally requires special equipment. To this end, in this article, we propose WiPass, a device-free authentication system only leveraging the pervasive Wi-Fi infrastructure to explore keystroke dynamics (manner and rhythm of keystrokes) captured by the channel state information to recognize legitimate users while rejecting spoofers. However, it remains an open challenge to characterize the behavioral features hidden in the human subtle motions, such as keystrokes. Therefore, we build a signal enhancement model using Ricean distribution to amplify user keystroke dynamics and a hybrid learning model for user authentication, which consists of two parts, i.e., convolutional neural network based feature extraction and support vector machine based classification. The former relies on visualizing the channel responses into time-series images to learn the behavioral features of keystrokes in energy and spectrum domains, whereas the latter exploits such behavioral features for user authentication. We prototype WiPass on the low-cost off-the-shelf Wi-Fi devices and verify its performance. Empirical results show that WiPass achieves on average 92.1% authentication accuracy, 5.9% false accept rate, and 6.3% false reject rate in three real environments.
Yu Gu 0003, Yantong Wang, Meng Wang 0037, Zulie Pan, Zhi Liu 0002, Mianxiong Dong
IEEE Trans. Ind. Informatics2
2020 Caching as an Image Characterization Problem using Deep Convolutional Neural Networks
abstract
Caching of popular content closer to the mobile user can significantly increase overall user experience as well as network efficiency by decongesting backbone network segments in the case of congestion episodes. In order to find the optimal caching locations, many conventional approaches rely on solving a complex optimization problem that suffers from the curse of dimensionality, which may fail to support online decision making. In this paper we propose a framework to amalgamate model based optimization with data driven techniques by transforming an optimization problem to a grayscale image and train a convolutional neural network (CNN) to predict optimal caching location policies. The rationale for the proposed modelling comes from CNN's superiority to capture features in grayscale images reaching human level performance in image recognition problems. The CNN is trained with optimal solutions and numerical investigations reveal that the performance can increase by more than 400% compared to powerful randomized greedy algorithms. To this end, the proposed technique seems as a promising way forward to the holy grail aspect in resource orchestration which is providing high quality decision making in real time.
Yantong Wang, Vasilis Friderikos
ICC1
2019 A Contactless and Fine-Grained Sleep Monitoring System Leveraging WiFi Channel Response
abstract
How can we effectively log a fine-grained sleep record consisting of still postures and in-place motions for the sleep disorder diagnosis without any specialized hardware? Existing sensor-based or vision-based solutions are either obstructive to use or rely on particular devices. This paper introduces SleepGuardian, a Radio Frequency (RF) based sleep monitoring system leveraging only omnipresent WiFi signals to provide a silent (unobtrusive and free of privacy concerns) yet loyal (finegrained and reliable) logging service. The key to SleepGuardian is to model the energy feature of wireless channel as a Gaussian Mixture Model (GMM) to adaptively recognize motions happened during sleep. We prototype SleepGuardian with off-the-shelf WiFi devices and evaluate it in an office. Experimental results over 11 subjects with several artificial and real periods of sleep demonstrate that SleepGuardian is effective since it achieves 100% overall accuracy (ACC), 0% false negative rate (FNR) and 0.64 s mean absolute error (MAE) on average. Considering that SleepGuardian is compatible with existing WiFi infrastructure, it constitutes a low-cost yet promising solution for sleep monitoring.
Yu Gu 0003, Yantong Wang, Zhi Liu 0002, Yusheng Ji, Jie Li 0002
ICC3
2019 Proactive Caching in Mobile Networks with Delay Guarantees
abstract
The explosive growth of mobile data traffic and the envisioned delay sensitive applications in 5G networks ranging from high definition video streaming with strict playout deadlines to multi modal tactile/haptic with kinaesthetic feedback that require some form of edge cloud cache support makes mobility a challenge. In this paper, we propose a Proactive Caching with Delay Guarantees (PCDG) approach to enhance the supporting of seamless mobility within 5G networks that are Information-Centric Networking (ICN)-aware. The proposed scheme is designed to cache contents into a set of potential edge clouds with delay guarantees and to achieve a trade-off among caching, redirection and missing cost. In particular, this approach consider the delay constraints in mobile network, especially the queuing delay in network links and edge clouds, which are modeled as M/M/1 and M/M/c queuing systems respectively. We formulate and linearize this problem as a Mixed Integer Linear Programming (MILP) model and compare the performance with other techniques. The result obtained from simulation reveal that the proposed PCDG scheme lead to a significantly lower total cost and higher satisfied probability albeit higher computational/complexity cost.
Yantong Wang, Gao Zheng 0001, Vasilis Friderikos
ICC1