Tengxiang Zhang

dblp:201/8078 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Counteracting the Delayed Conversions in OCPC with Survival Analysis
abstract
As an emerging advertising pricing method, optimized cost-per-click (OCPC) has attracted much research interest. In OCPC, the platform employs intelligent bidding strategies to optimize the advertising performance. Although widely adopted, existing bidding strategies overlook the delayed conversion phenomenon in OCPC, that is, the platform needs to wait for a period to receive the corresponding conversion signal after a click. Ignoring such delayed conversions causes the bidding strategies to overestimate the cost-per-action (average cost of a conversion, CPA), bid low, and finally hurt the platform's revenue. Moreover, the characteristics of the OCPC scenario make estimating the conversion probabilities for the delayed conversions more difficult. To address these issues, this paper proposes SurvBid (bidding with Survival Analysis) which aims to predict the convert probabilities for those delayed conversions in OCPC scenario. The CPA can then be accurately estimated and used to guide existing bidding methods to make more accurate bids. To meet the needs of different advertising platforms, we provide two versions of SurvBid, SurvBid-M (SurvBid with multitask survival model) and SurvBid-C (SurvBid with Cox survival model) with theoretical results to guide the model selection. Both online and offline experiments demonstrate that SurvBid can improve the platform's revenue and advertisers' conversions.
Chenxuan He, Xiao Zhang 0034, Yichao Wang 0002, Tengxiang Zhang, Zhenhua Dong, Jun Xu 0001
WSDM5
2024 Unsupervised Human Activity Recognition Via Large Language Models and Iterative Evolution
abstract
Human activity recognition (HAR) is crucial for health monitoring and disease diagnosis in Internet-of-Things environments. However, existing HAR approaches either suffer from poor accuracy or achieve high accuracy at the expense of costly manual annotations. To overcome the challenge above, we propose a novel method named LLMIE-UHAR that that leverages LLMs and Iterative Evolution to realize Unsupervised HAR. Specifically, with our designed prompt engineering mechanism, we employ large language models to fuse both contextual and semantic information, and annotate key samples selected by a clustering algorithm. Moreover, LLMIE-UHAR enhances the recognition accuracy with iterative evolution of clustering algorithm, large language models and the neural network based recognition model. Experiments conducted on the public ARAS datasets show the efficiency of our method, achieving an accuracy of 96.00%. This highlights the practical value of our approach.
Jiayuan Gao, Yingwei Zhang 0002, Yiqiang Chen 0001, Tengxiang Zhang, Boshi Tang
ICASSP4
2024 GestureGPT: Toward Zero-Shot Free-Form Hand Gesture Understanding with Large Language Model Agents
abstract
Existing gesture interfaces only work with a fixed set of gestures defined either by interface designers or by users themselves, which introduces learning or demonstration efforts that diminish their naturalness. Humans, on the other hand, understand free-form gestures by synthesizing the gesture, context, experience, and common sense. In this way, the user does not need to learn, demonstrate, or associate gestures. We introduce GestureGPT, a free-form hand gesture understanding framework that mimics human gesture understanding procedures to enable a natural free-form gestural interface. Our framework leverages multiple Large Language Model agents to manage and synthesize gesture and context information, then infers the interaction intent by associating the gesture with an interface function. More specifically, our triple-agent framework includes a Gesture Description Agent that automatically segments and formulates natural language descriptions of hand poses and movements based on hand landmark coordinates. The description is deciphered by a Gesture Inference Agent through self-reasoning and querying about the interaction context (e.g., interaction history, gaze data), which is managed by a Context Management Agent. Following iterative exchanges, the Gesture Inference Agent discerns the user’s intent by grounding it to an interactive function. We validated our framework offline under two real-world scenarios: smart home control and online video streaming. The average zero-shot Top-1/Top-5 grounding accuracies are 44.79%/83.59% for smart home tasks and 37.50%/73.44% for video streaming tasks. We also provide an extensive discussion that includes rationale for model selection, generalizability, and future research directions for a practical system etc.
Tengxiang Zhang, Chun Yu, Shengdong Zhao 0001, Yiqiang Chen 0001
Proc. ACM Hum. Comput. Interact.3
2024 InputJump: Augmented reality-facilitated cross-device input fusion based on spatial and semantic information
abstract
The proliferation of computing devices requires seamless cross-device interactions. Augmented reality (AR) headsets can facilitate interactions with existing computers owing to their user-centered views and natural inputs. In this study, we propose InputJump, a user-centered cross-device input fusion method that maps multi-modal cross-device inputs to interactive elements on graphical interfaces. The input jump calculates the spatial coordinates of the input target positions and the interactive elements within the coordinate system of the AR headset. It also extracts semantic descriptions of inputs and elements using large language models (LLMs). Two types of information from different inputs (e.g., gaze, gesture, mouse, and keyboard) were fused to map onto an interactive element. The proposed method is explained in detail and implemented on both an AR headset and a desktop PC. We then conducted a user study and extensive simulations to validate our proposed method. The results showed that InputJump can accurately associate a fused input with the target interactive element, enabling a more natural and flexible interaction experience.
Tengxiang Zhang, Yukang Yan, Yiqiang Chen 0001
Virtual Real. Intell. Hardw.3
2023 Modaldrop: Modality-Aware Regularization for Temporal-Spectral Fusion in Human Activity Recognition
abstract
Although most of existing works for sensor-based Human Activity Recognition rely on the temporal view, we argue that the spectral view also provides complementary prior and accordingly benchmark a standard multi-view framework with extensive experiments to demonstrate its consistent superiority over single-view opponents. We then delve into the intrinsic mechanism of the multi-view representation fusion, and propose ModalDrop as a novel modality-aware regularization method to learn and exploit representations of both views effectively. We demonstrate its advantage over existing representation fusion alternatives with comprehensive experiments and ablations. The improvements are consistent for various settings and are orthogonal with different backbones. We also discuss its potential application for other related tasks regarding representation or modality fusion. The source code is available on https://github.com/studyzx/ModalDrop.git.
Yiqiang Chen 0001, Benfeng Xu, Tengxiang Zhang
ICASSP4
2022 Easily-add battery-free wireless sensors to everyday objects: system implementation and usability study
Tengxiang Zhang, Zi Qian, Hsuan-Wei Fan, Jie Ren 0017, Yuntao Wang 0001, Yuanchun Shi
CCF Trans. Pervasive Comput. Interact.1
2021 What can "drag & drop" tell? Detecting mild cognitive impairment by hand motor function assessment under dual-task paradigm
Yingwei Zhang 0002, Yiqiang Chen 0001, Hanchao Yu, Zeping Lv, Xiaodong Yang 0005, Chunyu Hu 0001, Tengxiang Zhang
Int. J. Hum. Comput. Stud.7
2020 MoveVR: Enabling Multiform Force Feedback in Virtual Reality using Household Cleaning Robot
abstract
Haptic feedback can significantly enhance the realism and immersiveness of virtual reality (VR) systems. In this paper, we propose MoveVR, a technique that enables realistic, multiform force feedback in VR leveraging commonplace cleaning robots. MoveVR can generate tension, resistance, impact and material rigidity force feedback with multiple levels of force intensity and directions. This is achieved by changing the robot's moving speed, rotation, position as well as the carried proxies. We demonstrated the feasibility and effectiveness of MoveVR through interactive VR gaming. In our quantitative and qualitative evaluation studies, participants found that MoveVR provides more realistic and enjoyable user experience when compared to commercially available haptic solutions such as vibrotactile haptic systems.
Yuntao Wang 0001, Zichao (Tyson) Chen, Hanchuan Li, Zhengyi Cao, Huiyi Luo, Tengxiang Zhang, Ke Ou, John Raiti, Chun Yu, Shwetak N. Patel, Yuanchun Shi
CHI6
2020 ThermalRing: Gesture and Tag Inputs Enabled by a Thermal Imaging Smart Ring
abstract
The heterogeneous and ubiquitous input demands in smart spaces call for an input device that can enable rich and spontaneous interactions. We propose ThermalRing, a thermal imaging smart ring using low-resolution thermal camera for identity-anonymous, illumination-invariant, and power-efficient sensing of both dynamic and static gestures. We also design ThermalTag, thin and passive thermal imageable tags that reflect the heat from the human hand. ThermalTag can be easily made and applied onto everyday objects by users. We develop sensing techniques for three typical input demands: drawing gestures for device pairing, click and slide gestures for device control, and tag scan gestures for quick access. The study results show that ThermalRing can recognize nine drawing gestures with an overall accuracy of 90.9%, detect click gestures with an accuracy of 94.9%, and identify among six ThermalTags with an overall accuracy of 95.0%. Finally, we show the versatility and potential of ThermalRing through various applications.
Tengxiang Zhang, Yinshuai Zhang, Ke Sun 0003, Yuntao Wang 0001, Yiqiang Chen 0001
CHI1
2017 BitID: Easily Add Battery-Free Wireless Sensors to Everyday Objects
abstract
Radio-Frequency Identification (RFID) systems are becoming increasingly used within smart environments. In this paper, we propose BitID, a passive Ultra-High Frequency (UHF) RFID based sensing technique that can easily be made using off-the-shelf tags. BitID can be added to everyday objects to enable sensing and control capabilities. With a simple shorting mechanism, BitID is able to differentiate between two states of the object to which it is attached (for example, whether a door is open or closed). We explain the working principle of BitID, and demonstrate how to build and apply it to target objects. We also show that by using a three- layered system architecture, BitID can be used for various applications, including event detection, energy monitoring, fitness tracking, human behavior tracking and control.
Tengxiang Zhang, Nicholas Becker, Yuntao Wang 0001, Yuanchun Shi
SMARTCOMP1