Weiwei Jiang 0001

dblp:30/9639-1 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-4413-2497ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 4 since 2021Computer networks · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Understanding the Effects of Interaction on Emotional Experiences in VR
abstract
Virtual reality has been effectively used for eliciting emotions, yet most research focuses on the intensity of affective responses rather than on how interaction influences those experiences. To address this gap, we advance a validated VR emotion-elicitation dataset through two key extensions. First, we add a new high-arousal, high-valence scene and validate its effectiveness in a within-subject study (N=24). Second, we incorporate interactive elements into each scene, creating both interactive and non-interactive versions to examine the impact of interaction on emotional responses. We evaluate interaction through a multimodal approach combining subjective ratings and physiological signals to capture both conscious and unconscious affective responses. Our evaluation study (N=84) shows that interaction not only amplifies emotions but modulates them in context, supporting coping in negative scenes and enhancing enjoyment in positive scenes. These findings highlight the potential of scene-tailored interaction for different applications, where regulating emotions is as important as eliciting them.
Zheyuan Kuang, Tinghui Li 0001, Weiwei Jiang 0001, Sven Mayer, Flora D. Salim, Benjamin Tag, Anusha Withana, Zhanna Sarsenbayeva
CHI3
2026 EA-APO: A Universal Proactive Defense Against Facial Manipulation
abstract
The advent of deep learning has accelerated the development of facial manipulation techniques, particularly face-swapping and face attribute editing, raising serious concerns about privacy and identity-related misuse. Existing proactive defense methods predominantly target attribute editing and often generalize poorly to face-swapping models, making it difficult to provide effective protection across both tasks within a unified framework. To bridge this gap, we propose a generalized defense framework, Epoch-Adaptive Adversarial Perturbation Optimization (EA-APO). Specifically, EA-APO introduces a proactive defense mechanism that establishes optimal adversarial paths by optimizing perturbations on a white-box surrogate model to enhance adversarial transferability, and applies the resulting perturbations to source face images to disrupt both face swapping and face attribute editing, even against previously unseen target models in black-box settings. This approach mitigates identity feature tampering while adapting to changes in visual attributes and preserving high-quality adversarial examples. Experimental results show the generalization of our method across multiple face-swapping and attribute-editing models, including commercial ones, while also maintaining strong defense under various common post-processing operations and real-world social media transmission conditions, underscoring its potential for real-world deployment.
Lizhi Xiong, Ziqiang Li 0001, Weiwei Jiang 0001, Zhangjie Fu 0001, Zhihua Xia
IEEE Trans. Inf. Forensics Secur.4
2025 Is Artificial Intelligence Generated Image Detection a Solved Problem?
abstract
The rapid advancement of generative models, such as GANs and Diffusion models, has enabled the creation of highly realistic synthetic images, raising serious concerns about misinformation, deepfakes, and copyright infringement. Although numerous Artificial Intelligence Generated Image (AIGI) detectors have been proposed, often reporting high accuracy, their effectiveness in real-world scenarios remains questionable. To bridge this gap, we introduce AIGIBench, a comprehensive benchmark designed to rigorously evaluate the robustness and generalization capabilities of state-of-the-art AIGI detectors. AIGIBench simulates real-world challenges through four core tasks: multi-source generalization, robustness to image degradation, sensitivity to data augmentation, and impact of test-time pre-processing. It includes 23 diverse fake image subsets that span both advanced and widely adopted image generation techniques, along with real-world samples collected from social media and AI art platforms. Extensive experiments on 11 advanced detectors demonstrate that, despite their high reported accuracy in controlled settings, these detectors suffer significant performance drops on real-world data, limited benefits from common augmentations, and nuanced effects of pre-processing, highlighting the need for more robust detection strategies. By providing a unified and realistic evaluation framework, AIGIBench offers valuable insights to guide future research toward dependable and generalizable AIGI detection.
Ziqiang Li 0001, Jiazhen Yan, Ziwen He, Weiwei Jiang 0001, Lizhi Xiong, Zhangjie Fu 0001
NeurIPS5
2025 EarOE: Enabling Body-Channel Voice Interaction Interface on Earphones via Occlusion Effect
abstract
Nowadays, voice input on earphones has become one of the most paramount human-computer interaction approaches. Traditional voice interaction is built on the air channel, which is highly noise-susceptible and suffers from being falsely triggered by nearby competing users. In our work, we design a noise-resistant voice interaction interface on earphones, namedEarOE, which takes advantage of the narrow-bandwidth body channel to reconstruct high-fidelity audible speech. Although promising, directly taking the body channel as a voice interaction interface is nontrivial since the limited bandwidth of body-channel speech causes the original timbre information and linguistic content to be lost. To address these issues, we employ an electro-acoustic (EA) model for occlusion effect-based cross-channel correlation analysis. Based on that, we carefully design an attention-based encoder-decoder network to embrace cross-channel correlation for high-quality wide-bandwidth spectrum synthesis. To accommodate individual differences and improve model generalization, we implement a physics-based data augmentation strategy to expand the scale of the training dataset. Through extensive real-world experiments with 28 participants,EarOE can achieve an average Mel-cepstral distance of 8.08, an average modulation spectra distance of 0.82, and an average log-spectral distance of 11.16, outperforming existing solutions.
Feiyu Han, You Zuo, Weiwei Jiang 0001, Dawei Yan 0005, Panlong Yang, Yubo Yan
IEEE Internet Things J.3
2025 SEER: Knowledge-driven semantic image restoration with vision-language diffusion alignment
Shengliang Wu, Xin He 0017, Yong Xu 0001, Yujun Zhu, Weiwei Jiang 0001, Heju Li
Knowl. Based Syst.6
2024 An Immersive and Interactive VR Dataset to Elicit Emotions
abstract
Images and videos are widely used to elicit emotions; however, their visual appeal differs from real-world experiences. With virtual reality becoming more realistic, immersive, and interactive, we envision virtual environments to elicit emotions effectively, rapidly, and with high ecological validity. This work presents the first interactive virtual reality dataset to elicit emotions. We created five interactive virtual environments based on corresponding validated 360° videos and validated their effectiveness with 160 participants. Our results show that our virtual environments successfully elicit targeted emotions. Compared with the existing methods using images or videos, our dataset allows virtual reality researchers and practitioners to integrate their designs effectively with emotion elicitation settings in an immersive and interactive way.
Weiwei Jiang 0001, Maximiliane Windl, Benjamin Tag, Zhanna Sarsenbayeva, Sven Mayer
IEEE Trans. Vis. Comput. Graph.1
2023 PAssTrack: Practical and Accurate Passive Human Tracking System Using Commodity Wi-Fi
abstract
In this paper, we present PAssTrack, a Wi-Fi based passive human tracking system which is adapted to the practical antenna spacing of most commodity Wi-Fi access points (APs) and achieves accurate tracking results. We mainly enhance our system in the following four aspects. Firstly, we modify 2D MUSIC algorithm for estimating parameters including angle of arrival (AoA) and relative time of flight (rToF) of dynamic human reflection path. And we analyze the superiority of our algorithm compared with the state-of-the-art solutions. Secondly, we leverage the estimated rToF and the continuity of AoA to resolve the angle ambiguity caused by antenna spacing larger than half of the wavelength. Thirdly, we optimize the mesh model based on Fresnel zone theory to a dual antenna version for fine-grained velocity estimation. Lastly, we introduce hologram for target localization in order to compensate the errors of estimated path parameters using velocity estimates, and reduce the impact of outliers through kernel density estimation (KDE). The experiments are conducted in two real-world indoor environments, and the results demonstrate the advantages of PAssTrack in aspects of better tracking accuracy than the state-of-the-arts, easy and effective calibration, robustness under the case of blocked transceivers and adaptiveness to different antenna spacing.
Boxiao Zhang, Panlong Yang, Yubo Yan, Xin He 0017, Weiwei Jiang 0001
MSN5
2023 Human Activity Recognition Using Smartphones With WiFi Signals
abstract
In this article, we present a work using a smartphone with an off-the-shelf WiFi router for human activity recognition with various scales. The router serves as a hotspot for transmitting WiFi packets. The smartphone is configured with customized firmware and developed software for capturing WiFi channel state information (CSI) data. We extract the features from the CSI data associated with specific human activities, and utilize the features to classify the activities using machine learning models. To evaluate the system performance, we test 20 types of human activities with different scales including seven small motions, four medium motions, and nine big motions. We recruit 60 participants and spend 140 hours for data collection at various experimental settings, and have 36 000 data points collected in total. Furthermore, for comparison, we adopt three distinct machine learning models, including convolutional neural networks (CNNs), decision tree, and long short-term memory. The results demonstrate that our system can predict these human activities with an overall accuracy of 97.25%. Specifically, our system achieves a mean accuracy of 97.57% for recognizing small-scale motions that are particularly useful for gesture recognition. We then consider the adaptability of the machine learning algorithms in classifying the motions, where CNN achieves the best predicting accuracy. As a result, our system enables human activity recognition in a more ubiquitous and mobile fashion that can potentially enhance a wide range of applications such as gesture control, sign language recognition, etc.
Guiping Lin, Weiwei Jiang 0001, Sicong Xu, Xiaobo Zhou 0003, Yujun Zhu, Xin He 0017
IEEE Trans. Hum. Mach. Syst.2
2023 Near-infrared Imaging for Information Embedding and Extraction with Layered Structures
abstract
Non-invasive inspection and imaging techniques are used to acquire non-visible information embedded in samples. Typical applications include medical imaging, defect evaluation, and electronics testing. However, existing methods have specific limitations, including safety risks (e.g., X-ray), equipment costs (e.g., optical tomography), personnel training (e.g., ultrasonography), and material constraints (e.g., terahertz spectroscopy). Such constraints make these approaches impractical for everyday scenarios. In this article, we present a method that is low-cost and practical for non-invasive inspection in everyday settings. Our prototype incorporates a miniaturized near-infrared spectroscopy scanner driven by a computer-controlled 2D-plotter. Our work presents a method to optimize content embedding, as well as a wavelength selection algorithm to extract content without human supervision. We show that our method can successfully extract occluded text through a paper stack of up to 16 pages. In addition, we present a deep-learning-based image enhancement model that can further improve the image quality and simultaneously decompose overlapping content. Finally, we demonstrate how our method can be generalized to different inks and other layered materials beyond paper. Our approach enables a wide range of content embedding applications, including chipless information embedding, physical secret sharing, 3D print evaluations, and steganography.
Weiwei Jiang 0001, Difeng Yu, Chaofan Wang 0001, Zhanna Sarsenbayeva, Niels van Berkel, Jorge Gonçalves 0001, Vassilis Kostakos
ACM Trans. Graph.1
2022 Hand Hygiene Quality Assessment Using Image-to-Image Translation
Chaofan Wang 0001, Kangning Yang, Weiwei Jiang 0001, Jing Wei 0002, Zhanna Sarsenbayeva, Jorge Gonçalves 0001, Vassilis Kostakos
MICCAI (8)3
2022 LF-SWIPT: Outage Analysis for SWIPT Relaying Networks Using Lossy Forwarding With QoS Guaranteed
abstract
We analyze the outage performance of a lossy forwarding (LF) relaying system with the simultaneous wireless information and power transfer (SWIPT) capability. In the system of LF with SWIPT (LF-SWIPT), a source broadcasts its message to both a relay and a destination. A relay node with SWIPT functionality harvests energy and decodes information from the source signal. The energy is split into two parts for information processing and message forwarding, respectively. For information processing, the relay attempts to decode the incoming source signal and forms an estimate. Unlike the existing decode-and-forward SWIPT system (DF-SWIPT), the estimate is always forwarded using the harvested energy. The destination performs joint decoding to recover the message with the signals received from both the source node and the relay node. We derive the outage probability for the LF-SWIPT system based on the theorem ofsource coding with side information. The simulation results demonstrate that the proposed system achieves significant gains (around 1–2 dB) compared to the DF-SWIPT system. We further evaluate the impact of the distance and the power splitting (PS) strategy on the system performance using simulations. Finally, we build an optimization algorithm on the PS ratio by maximizing the admissible region from the theoretical perspective.
Guiping Lin, Yike Zhou, Weiwei Jiang 0001, Xin He 0017, Xiaobo Zhou 0003, Guodong He, Panlong Yang
IEEE Internet Things J.3
2022 Understanding How to Administer Voice Surveys through Smart Speakers
abstract
Smart speakers have become exceedingly popular and entered many people's homes due to their ability to engage users with natural conversations. Researchers have also looked into using smart speakers as an interface to collect self-reported health data through conversations. Responding to surveys prompted by smart speakers requires users to listen to questions and answer in voice without any visual stimuli. Compared to traditional web-based surveys, where users can see questions and answers visually, voice surveys may be more cognitively challenging. Therefore, to collect reliable survey data, it is important to understand what types of questions are suitable to be administered by smart speakers. We selected five common survey questionnaires and deployed them as voice surveys and web surveys in a within-subject study. Our 24 participants answered questions using voice and web questionnaires in one session. They then repeated the same study session after 1 week to provide a "retest'' response. Our results suggest that voice surveys have comparable reliability to web surveys. We find that, when using 5-point or 7-point scales, voice surveys take about twice as long as web surveys. Based on objective measurements, such as response agreement and test-retest reliability, and subjective evaluations of user experience, we recommend that researchers consider adopting the binary scale and 5-point numerical scales for voice surveys on smart speakers.
Jing Wei 0002, Weiwei Jiang 0001, Chaofan Wang 0001, Difeng Yu, Jorge Gonçalves 0001, Tilman Dingler, Vassilis Kostakos
Proc. ACM Hum. Comput. Interact.2
2021 User Trust in Assisted Decision-Making Using Miniaturized Near-Infrared Spectroscopy
abstract
We investigate the use of a miniaturized Near-Infrared Spectroscopy (NIRS) device in an assisted decision-making task. We consider the real-world scenario of determining whether food contains gluten, and we investigate how end-users interact with our NIRS detection device to ultimately make this judgment. In particular, we explore the effects of different nutrition labels and representations of confidence on participants’ perception and trust. Our results show that participants tend to be conservative in their judgment and are willing to trust the device in the absence of understandable label information. We further identify strategies to increase user trust in the system. Our work contributes to the growing body of knowledge on how NIRS can be mass-appropriated for everyday sensing tasks, and how to enhance the trustworthiness of assisted decision-making systems.
Weiwei Jiang 0001, Zhanna Sarsenbayeva, Niels van Berkel, Chaofan Wang 0001, Difeng Yu, Jing Wei 0002, Jorge Gonçalves 0001, Vassilis Kostakos
CHI1
2020 Does Smartphone Use Drive our Emotions or vice versa? A Causal Analysis
abstract
In this paper, we demonstrate the existence of a bidirectional causal relationship between smartphone application use and user emotions. In a two-week long in-the-wild study with 30 participants we captured 502,851 instances of smartphone application use in tandem with corresponding emotional data from facial expressions. Our analysis shows that while in most cases application use drives user emotions, multiple application categories exist for which the causal effect is in the opposite direction. Our findings shed light on the relationship between smartphone use and emotional states. We furthermore discuss the opportunities for research and practice that arise from our findings and their potential to support emotional well-being.
Zhanna Sarsenbayeva, Gabriele Marini, Niels van Berkel, Chu Luo, Weiwei Jiang 0001, Kangning Yang, Greg Wadley, Tilman Dingler, Vassilis Kostakos, Jorge Gonçalves 0001
CHI5
2020 GuardRider: Reliable WiFi Backscatter Using Reed-Solomon Codes With QoS Guarantee
abstract
The WiFi backscatter communications offer ultralow power and ubiquitous connections for IoT systems. Caused by the intermittent-nature of the WiFi traffics, state-of-the-art WiFi backscatter communications are not reliable for backscatter link or simple for the tag to do the adaptive transmission. In order to build reliable WiFi backscatter communications, we present GuardRider, a WiFi backscatter system that enables backscatter communications to improve the quality of service (QoS). The key contribution of GuardRider is an optimization algorithm of designing RS codes to follow the statistical knowledge of WiFi traffics and adjust backscatter transmission. With GuardRider, the reliable baskscatter link is guaranteed and a backscatter tag is able to adaptively transmit information without heavily listening to the excitation channel, by taking QoS into account. We built a hardware prototype of GuardRider using a customized tag with FPGA implementation. Both the simulations and field experiments verify that GuardRider could achieve notably gains in bit error rate and frame error rate, which are a hundredfold reduction in simulations and around 99% in filed experiments. Our system is able to achieve around 700 kbps throughput.
Xin He 0017, Weiwei Jiang 0001, Meng Cheng 0001, Xiaobo Zhou 0003, Panlong Yang, Brian M. Kurkoski
IWQoS2
2019 Effect of Ambient Light on Mobile Interaction
Zhanna Sarsenbayeva, Niels van Berkel, Weiwei Jiang 0001, Danula Hettiachchi, Vassilis Kostakos, Jorge Gonçalves 0001
INTERACT (3)3
2014 Pulse: low bitrate wireless magnetic communication for smartphones
abstract
We present Pulse, a wireless magnetic communication protocol for smartphones. Pulse is designed for off-the-shelf Android smartphones with magnetometers, and encodes data in magnetic fields. We present the design and evaluation of Pulse in various conditions (e.g., different voltages, number of transfer channels). The system provides security due to its short range (~1 cm), it can reach a speed of up to 44 bits per second, and it is possible to run it on most mobile phones with a magnetometer. We present our evaluation and discuss practical use cases where Pulse can be used today.
Weiwei Jiang 0001, Denzil Ferreira, Jani Ylioja, Jorge Gonçalves 0001, Vassilis Kostakos
UbiComp1