Zhihang Zhu

dblp:327/1500 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BuyMate: Making AI Interventions Effective in Promoting Rational Consumption in Live Commerce
abstract
Live commerce platforms frequently employ algorithmic recommendations and time-limited promotions to trigger impulsive purchases, challenging rational consumer decision-making. While existing research has identified manipulative design patterns in live commerce, significant gaps remain in understanding consumer psychological motivations and developing counter-persuasion interventions. We conducted a multi-stage formative study involving surveys (N = 116), interviews (N = 21), and co-design workshops (N = 16) to explore user preferences for rational consumption support systems. Informed by these insights, we designed BuyMate, which provides gentle, real-time rational interventions through product comparison and persuasive speech reframing. A user evaluation (N = 35) demonstrates that the system effectively reduces impulsive purchases, enhances decision autonomy, and promotes sustainable consumption. This work contributes an AI-driven counter-persuasion approach, identifies user-centered principles for adaptive interventions, and offers practical guidance for responsible AI in digital commerce.
Yishan Liu, Zhihang Zhu, Xuerui Ma, Tianyang Feng, Qingfei Zhao, XinZhi Zhang, Haipeng Mi
CHI3
2024 Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time Variations
abstract
Spatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high computational costs and limited performance in real-world scenes. This paper presents an innovative Wavelet-based SpatioTemporal (WaST) framework, which extracts and adaptively controls both low and high-frequency components at image and feature levels via 3D discrete wavelet transform for faster processing while maintaining high-quality predictions. We propose a Time-Frequency Aware Translator uniquely crafted to efficiently learn short- and long-range spatiotemporal information by individually modeling spatial frequency and temporal variations. Meanwhile, we design a wavelet-domain High-Frequency Focal Loss that effectively supervises high-frequency variations. Extensive experiments across various real-world scenarios, such as driving scene prediction, traffic flow prediction, human motion capture, and weather forecasting, demonstrate that our proposed WaST achieves state-of-the-art performance over various spatiotemporal prediction methods.
Xuesong Nie, Yunfeng Yan, Siyuan Li 0002, Cheng Tan 0012, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Stan Z. Li, Donglian Qi
AAAI7
2024 PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative Decoding
abstract
Predictive learning models, which aim to predict future frames based on past observations, are crucial to constructing world models. These models need to maintain low-level consistency and capture high-level dynamics in unannotated spatiotemporal data. Transitioning from frame-wise to token-wise prediction presents a viable strategy for addressing these needs. How to improve token representation and optimize token decoding presents significant challenges. This paper introduces PredToken, a novel predictive framework that addresses these issues by decoupling space-time tokens into distinct components for iterative cascaded decoding. Concretely, we first design a “decomposition, quantization, and reconstruction” schema based on VQGAN to improve the token representation. This scheme disentangles low- and high-frequency representations and employs a dimension-aware quantization model, allowing more low-level details to be preserved. Building on this, we present a “coarse-to-fine iterative decoding” method. It leverages dynamic soft decoding to refine coarse tokens and static soft decoding for fine tokens, enabling more high-level dynamics to be captured. These designs make Pred-Token produce high-quality predictions. Extensive experiments demonstrate the superiority of our method on various real-world spatiotemporal predictive benchmarks. Furthermore, PredToken can also be extended to other visual generative tasks to yield realistic outcomes.
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
CVPR5
2024 SAMP: Adapting Segment Anything Model for Pose Estimation
abstract
Segment Anything Model (SAM) exhibits superior performance for segmentation. Many follow-up works explore adapting this powerful model to specific domains. However, those works mainly focus on different sub-tasks of segmentation. The cross-task generalization ability of SAM is still not explored. In this paper, we propose SAMP (SAM for Pose), which makes the first attempt to adapt SAM for pose estimation. We observe that SAM could segment different human parts with specific prompts, proving that it contains the knowledge to understand the human structure. Considering that localizing keypoints requires fine-grained perceptual capabilities, we design a Detail-aware Adapter (DA-Adapter), which complements the features of the SAM encoder with multi-scale feature fusion and multi-level supervision. Experimental results demonstrate that SAMP achieves novel state-of-the-art against previously specifically designed pose estimation methods. Specifically, with ViT-B backbone, SAMP achieves 78.1% AP on the COCO val2017, 77.1% AP on the COCO test-dev2017, and 70.5% AP on the CrowdPose dataset.
Zhihang Zhu, Yunfeng Yan, Haoyuan Jin, Xuesong Nie, Donglian Qi, Xi Chen 0119
ICME1
2024 Object-Level Pseudo-3D Lifting for Distance-Aware Tracking
abstract
Multi-object tracking (MOT) is a pivotal task for media interpretation, where reliable motion and appearance cues are essential for cross-frame identity preservation. However, limited by the inherent perspective properties of 2D space, the crowd density and frequent occlusions in real-world scenes expose the fragility of these cues. We observe the natural advantage of objects being well-separated in high-dimensional space and propose a novel 2D MOT framework, "Detecting-Lifting-Tracking'' (DLT). Initially, a pre-trained detector is employed to capture 2D object information. Secondly, we introduce a Mamba Distance Estimator to obtain the distances of objects to a monocular camera with temporal consistency, achieving object-level pseudo-3D lifting. Finally, we thoroughly explore distance-aware tracking via pseudo-3D information. Specifically, we introduce a Score-Distance Hierarchical Matching and Short-Long Terms Association to enhance accurate and robust association capability. Even without appearance cues, our DLT achieves state-of-the-art performance on MOT17, MOT20, and DanceTrack, demonstrating its potential to address occlusion challenges.
Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
ACM Multimedia5
2024 Triplet Attention Transformer for Spatiotemporal Predictive Learning
abstract
Spatiotemporal predictive learning offers a self-supervised learning paradigm that enables models to learn both spatial and temporal patterns by predicting future sequences based on historical sequences. Mainstream methods are dominated by recurrent units, yet they are limited by their lack of parallelization and often underperform in real-world scenarios. To improve prediction quality while maintaining computational efficiency, we propose an innovative triplet attention transformer designed to capture both inter-frame dynamics and intra-frame static features. Specifically, the model incorporates the Triplet Attention Module (TAM), which replaces traditional recurrent units by exploring self-attention mechanisms in temporal, spatial, and channel dimensions. In this configuration: (i) temporal tokens contain abstract representations of inter-frame, facilitating the capture of inherent temporal dependencies; (ii) spatial and channel attention combine to refine the intra-frame representation by performing fine-grained interactions across spatial and channel dimensions. Alternating temporal, spatial, and channel-level attention allows our approach to learn more complex short-and long-range spatiotemporal dependencies. Extensive experiments demonstrate performance surpassing existing recurrent-based and recurrent-free methods, achieving state-of-the-art under multi-scenario examination including moving object trajectory prediction, traffic flow prediction, driving scene prediction, and human motion capture.
Xuesong Nie, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Yunfeng Yan, Donglian Qi
WACV4
2024 ScopeViT: Scale-Aware Vision Transformer
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
Pattern Recognit.5
2024 AHOR: Online Multi-Object Tracking With Authenticity Hierarchizing and Occlusion Recovery
abstract
Despite extensive exploration of more powerful multi-object tracking (MOT) frameworks, the impact of frequent occlusion has remained a formidable challenge. In this work, we present a novel MOT framework with Authenticity Hierarchizing and Occlusion Recovery (AHOR), that strikingly handles occlusion and demonstrates superior precision and adaptability. Specifically, through an in-depth analysis of the classical tracking-by-detection (TBD) paradigm, we fully upgrade three aspects. Firstly, we propose an Existence Score that provides a more accurate depiction of detection authenticity under occlusion, enhancing the effectiveness and robustness of the hierarchical association. Secondly, we present an ingeniously devised pre-processing method in conjunction with a Recovery Intersection over Union (RIoU) for location similarity measurement, addressing the adverse effects of occlusion-induced disparity between visible and true object regions. Lastly, we introduce an Occluded Person Re-identification Module (ODReID) that extracts appearance features from the restricted visible region, overcoming the critical dependence on object quality. Results of extensive experiments demonstrate that our AHOR achieves state-of-the-art performance on MOT17, MOT20, DanceTrack, and VisDrone test sets.
Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
IEEE Trans. Circuits Syst. Video Technol.5
2022 Stock Prediction Based on Deep Learning and its Application in Pairs Trading
abstract
This paper implements and analyses the effectiveness of three recent deep neural networks (RNN, LSTM, TCN) for stock price prediction as well as a financial method, ARIMA, in the context of time-series U.S. stock data. We have found that TCN model has the smallest error among all the deep neural networks. We also make comparisons between different industries of the stock and the proposed models can perform equally satisfying results on different industries. Many of our models, like TCN and LSTM, can generate satisfying outcomes. This is mainly reflected in the low RMSE and MAPE values. In most cases, the performance of deep learning models is better than traditional financial models for stock price prediction. We find that the combination of deep learning model and pairs trading strategy could be used for gaining more profits. The models we used and corresponding evaluation methods could be adopted to help individuals to better allocate their assets and increase their profit return.
Zhihang Zhu, Minghao Liu 0013, Changjiang Zhang, Bingyan Han
ISNCC3