Yitong Zhu

dblp:362/2290 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 67% Trustworthy machine learning · 33%
Human-computer interaction and pervasive computing
2 papers
Immersive interaction · 79% Wearable and physiological sensing · 21%
Computer graphics and multimedia
2 papers
Image and video processing · 54% Virtual and augmented reality · 46%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › sentiment analysis
multimodal sentiment analysis
1.012026
QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis · ACL (1) 2026
Machine learning › Trustworthy machine learning
robustness
1.012026
QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis · ACL (1) 2026
Natural language and speech › Information extraction and text analysis
sentiment analysis
1.012026
QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis · ACL (1) 2026
Image and video processing › motion estimation
optical flow
1.012026
Flow-Aware Diffusion for Real-Time VR Restoration: Mitigating Cybersickness With Enhanced Spatiotemporal Coherence · IEEE Trans. Vis. Comput. Graph. 2026
Immersive interaction › virtual reality › cybersickness
cybersickness mitigation
1.012026
Flow-Aware Diffusion for Real-Time VR Restoration: Mitigating Cybersickness With Enhanced Spatiotemporal Coherence · IEEE Trans. Vis. Comput. Graph. 2026
Immersive interaction
virtual reality
1.012026
Flow-Aware Diffusion for Real-Time VR Restoration: Mitigating Cybersickness With Enhanced Spatiotemporal Coherence · IEEE Trans. Vis. Comput. Graph. 2026
Virtual and augmented reality › cybersickness
cybersickness prediction
0.912025
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference · ACM Multimedia 2025
Wearable and physiological sensing
eye tracking
0.312025
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference · ACM Multimedia 2025
Wearable and physiological sensing › motion capture
head dynamics sensing
0.312025
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

optical flow attenuation · 2.0deep learning-based frame modulation · 2.0graph neural network · 1.7difference attention module · 1.7cross-modal alignment · 1.7mixture of experts · 1.0
YearPublicationVenuePosition
2026 QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis
abstract
Yitong Zhu, Yuxuan Jiang, Guanxuan Jiang, Bojing Hou, Peng Yuan Zhou, Ge Lin, Yuyang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yitong Zhu, Guanxuan Jiang, Bojing Hou, Peng Yuan Zhou, Ge Lin 0001
ACL (1)1
2026 Flow-Aware Diffusion for Real-Time VR Restoration: Mitigating Cybersickness With Enhanced Spatiotemporal Coherence
abstract
Cybersickness remains a critical barrier to the widespread adoption of Virtual Reality (VR), particularly in scenarios involving intense or artificial motion cues.Among the key contributors is excessive optical flow-perceived visual motion that, when unmatched by vestibular input, leads to sensory conflict and discomfort. While previous efforts have explored geometric or hardware-based mitigation strategies, such methods often rely on predefined scene structures, manual tuning, or intrusive equipment. In this work, we propose U-MAD, a lightweight, real-time, AI-based solution that suppresses perceptually disruptive optical flow directly at the image level. Unlike prior handcrafted approaches, this method learns to attenuate high-intensity motion patterns from rendered frames without requiring mesh-level editing or scene-specific adaptation. Designed as a plug-and-play module, U-MAD integrates seamlessly into existing VR pipelines and generalizes well to procedurally generated environments. The experiments show that U-MAD consistently reduces average optical flow and enhances temporal stability across diverse scenes. A user study further supports the finding that reducing visual motion defects can improve perceptual comfort and alleviate cybersickness symptoms. These findings demonstrate that perceptually guided modulation of optical flow provides an effective and scalable approach to creating more user-friendly immersive experiences.
Yitong Zhu, Guanxuan Jiang, Zhuowen Liang, Yuyang Wang 0002
IEEE Trans. Vis. Comput. Graph.1
2025 CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation
abstract
Gaze estimation methods typically encounter significant performance degradation in generalized tasks due to the domain mismatch between the source and target domains. Existing approaches attempt to utilize various domain generalization techniques. However, their generalization capabilities are limited since they are constrained to a single visual modality. Notably, large-scale contrastive language-image pre-training (CLIP) models have been widely applied to downstream visual tasks for their robust generalization capabilities, but the potential of CLIP for regression tasks has not been fully explored. To bridge this gap, we introduce a novel framework called CR-CLIP, which endows CLIP with the capability to generalize gaze estimation. Specifically, we convert gaze labels into textual descriptions and achieve alignment between images and text signals with gaze cues, thereby extracting generalized gaze-related features. To enhance the model’s understanding of the numerical relationships of gaze directions, we propose a novel regression loss function based on image-text similarity. Additionally, we fine-tune the model on the original gaze dataset, achieving high precision in generalized gaze estimation. Experimental results show that our proposed method achieves state-of-the-art performance on four generalized gaze estimation tasks.
Yitong Zhu, Xurong Xie, Naiming Yao, Hui Chen 0020, Feng Tian 0001
ICASSP1
2025 Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
abstract
Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.
Yitong Zhu, Zhuowen Liang, Tangyao Li, Yuyang Wang 0002
ACM Multimedia1
2025 Latent Information in the Evolving Energy Structure: The Intersection of Artificial Intelligence Development and Social Progress
abstract
Recent studies indicate that Artificial Intelligence (AI) technology, characterized by its integration of information and communication technology attributes, exerts a multifaceted influence on the energy system. The authors employ Difference-in-Differences (DID) and Triple-Difference (DDD) models to investigate the effects of AI. The research initially demonstrates that AI may exhibit dual attributes of constraining energy structure. Specifically, the rapid development of AI tends to increase the proportion of fossil fuel-based electricity generation while optimizing the consumption of renewable energy. Furthermore, the degradation of energy structure stems from the surge in electricity consumption by AI, which the renewable energy consumption capacity is unable to satisfy. Lastly, industrial agglomeration effects and the construction of the digital economy have positive impacts on development of renewable energy; technological innovation aids in mitigating the negative shocks to the energy structure. This study provides a new perspective on the role of AI technology in energy transition.
Boqiang Lin, Yitong Zhu
J. Glob. Inf. Manag.2