Yang Tian 0008

dblp:64/5869-8 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-6633-8579ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Metaphorical Visual Question Answering: Benchmark and Knowledge-Enhanced Metaphor Understanding Method
abstract
Fact and common-sense reasoning grounded in metaphorical imagery constitute a more challenging form of visual question answering (VQA). Under this form, models typically cannot obtain answers directly from images. Models first need to identify and comprehend the metaphorical components within the image, subsequently integrating prior knowledge to establish the mappings between the target and source domains depicted in the image. To evaluate this capability, we propose a VQA benchmark based on metaphorical images (METAVQA), which measures the understanding of the model of metaphorical images in the form of VQA. Experimental results indicate that interpreting metaphorical images requires robust prior knowledge and a strong ability to understand abstract components, which remains a challenge for LLMs. To improve the metaphorical understanding capability of LLMs, we propose a knowledge-enhanced method (KEMU). This method utilizes LLMs to extract image triplets, retrieve implicit metaphorical knowledge, and reason through questions step-by-step using chain of thought prompting, and pretrained classifier to match the refined reasoning output to the provided answer choices. Validated on the METAVQA dataset, KEMU outperforms LLMs in both one-hop and multi-hop questions. Our code and benchmark can be seen inhttps://github.com/VILAN-Lab/METAVQA.
Qingbao Huang, Peihang He, Pijian Li, Yang Tian 0008, Yi Cai 0001, Qing Li 0001
IEEE Trans. Multim.5
2025 GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
abstract
Large Language Models (LLMs) face significant limitations when applied to large-scale graphs, struggling with context constraints and inflexible reasoning. We introduce GraphChain, a novel framework enabling LLMs to analyze large graphs by orchestrating dynamic sequences of specialized tools, mimicking human exploratory processes. GraphChain incorporates two core technical contributions: (1) Progressive Graph Distillation, a reinforcement learning approach that learns to generate tool sequences balancing task relevance and intermediate state compression, thereby overcoming LLM context limitations. (2) Structure-aware Test-Time Adaptation (STTA), a mechanism using a lightweight, self-supervised adapter conditioned on graph spectral properties to efficiently adapt a frozen LLM policy to diverse graph structures via soft prompts without retraining. Experiments show GraphChain significantly outperforms prior methods, enabling scalable and adaptive LLM-driven graph analysis.
Chunyu Wei, Wenji Hu, Xingjia Hao, Yunhai Wang, Yang Tian 0008, Yueguo Chen
NeurIPS7
2025 BoundaryScreen: Summoning the Home Screen in VR via Walking Outward
abstract
A safety boundary wall in VR is a virtual barrier that defines a safe area, allowing users to navigate and interact without safety concerns. However, existing implementations neglect to utilize the safety boundary wall's large surface for displaying interactive information. In this work, we propose the BoundaryScreen technique based on the "walking outward" metaphor to add interactivity to the safety boundary wall. Specifically, we augment the safety boundary wall by placing the home screen on it. To summon the home screen, the user only needs to walk outward until it appears. Results showed that (i) participants significantly preferred BoundaryScreen in the outermost two-step-wide ring-shaped section of a circular safety area; and (ii) participants exhibited strong "behavioral inertia" for walking, i.e., after completing a routine activity involving constant walking, participants significantly preferred to use the walking-based BoundaryScreen technique to summon the home screen.
Yang Tian 0008, Xingjia Hao, Jianchun Su, Wei Sun 0050, Yangjian Pan, Yunhai Wang, Minghui Sun 0001, Teng Han, Ningjiang Chen
IEEE Trans. Vis. Comput. Graph.1
2025 SummonBrush: Enhancing Touch Interaction on Large XR User Interfaces by Augmenting Users' Hands with Virtual Brushes
abstract
Touch interaction is one of the fundamental interaction paradigms in XR, as users have become very familiar with touch interactions on physical touchscreens. However, users typically need to perform extensive arm movements for engaging with XR user interfaces much larger than mobile device touchscreens. We propose the SummonBrush technique to facilitate easy access to hidden windows while interacting with large XR user interfaces, requiring minimal arm movements. The SummonBrush technique adds a virtual brush to the index fingertip of a user's hand. Upon making contact with a virtual user interface, the brush bends and diverges and ink starts to diffuse in it. The more the brush bends and diverges, the more the ink diffuses. The user can summon hidden windows or background applications in situ, which is achieved by firstly pressing the brush against the user interface to make ink fully fill the brush and then perform swipe gestures. Also, the user can press the brush against the thumbtails of background applications in situ to quickly cycle them through. Ecological studies showed that SummonBrush significantly reduced the arm movement time by 39% and 34% in summoning hidden windows and activating/closing background applications, respectively, leading to a significant decrease in reported physical demand.
Yang Tian 0008, Zhao Su, Tianren Luo, Teng Han, Shengdong Zhao 0001, Boyu Gao 0003, Dangxiao Wang
IEEE Trans. Vis. Comput. Graph.1
2025 AmplitudeArrow: On-the-Go AR Menu Selection Using Consecutive Simple Head Gestures and Amplitude Visualization
abstract
Heads-up computing aims to provide synergistic digital assistance that minimally interferes with users' on-the-go daily activities. Currently, the input modalities of heads-up computing are mainly voice and finger gestures. In this work, we propose and evaluate the AmplitudeArrow (AA) technique designed for on-the-go AR menu selection to demonstrate that consecutive simple head gestures can also be an effective input modality for heads-up computing. Specifically, AA arranges menu icons into one/two row(s). To select a target icon, the user first makes their head yaw to pre-select the target icon or the column containing it and then makes their head pitch to make the arrow in the target icon expand until the arrow covers the target icon completely, i.e., the pitch amplitude surpasses the selection confirmation threshold. User studies indicated that AA demonstrated robust resistance to walking-caused head perturbation and external factors such as other people/obstacles, delivering high accuracy (error rate $< $< 5$\%$%) and fast speed ($< $< 1.5s per selection) when there were no more than six icon columns (twelve icons) distributed horizontally and evenly in a menu area with a horizontal visual angle of $43^{\circ }$43∘.
Yang Tian 0008, Yukang Yan, Shengdong Zhao 0001, Xiaojuan Ma, Yuanchun Shi
IEEE Trans. Vis. Comput. Graph.1
2024 XDrain: Effective log parsing in log streams using fixed-depth forest
Yang Tian 0008, Siyu Yu, Donghui Gao, Yifan Wu 0002, Suqun Huang, Xiaochun Hu, Ningjiang Chen
Inf. Softw. Technol.2
2024 Kine-Appendage: Enhancing Freehand VR Interaction Through Transformations of Virtual Appendages
abstract
Kinesthetic feedback, the feeling of restriction or resistance when hands contact objects, is essential for natural freehand interaction in VR. However, inducing kinesthetic feedback using mechanical hardware can be cumbersome and hard to control in commodity VR systems. We propose the kine-appendage concept to compensate for the loss of kinesthetic feedback in virtual environments, i.e., a virtual appendage is added to the user's avatar hand; when the appendage contacts a virtual object, it exhibits transformations (rotation and deformation); when it disengages from the contact, it recovers its original appearance. A proof-of-concept kine-appendage technique, BrittleStylus, was designed to enhance isomorphic typing. Our empirical evaluations demonstrated that (i) BrittleStylus significantly reduced the uncorrected error rate of naive isomorphic typing from 6.53% to 1.92% without compromising the typing speed; (ii) BrittleStylus could induce the sense of kinesthetic feedback, the degree of which was parity with that induced by pseudo-haptic (+ visual cue) methods; and (iii) participants preferred BrittleStylus over pseudo-haptic (+ visual cue) methods because of not only good performance but also fluent hand movements.
Yang Tian 0008, Hualong Bai, Shengdong Zhao 0001, Chi-Wing Fu, Chun Yu, Haozhao Qin, Qiong Wang 0001, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.1
2023 Automatically detecting human-object interaction by an instance part-level attention deep framework
Fenglian Chen, Yang Tian 0008
Pattern Recognit.3
2020 Virtually-Extended Proprioception: Providing Spatial Reference in VR through an Appended Virtual Limb
abstract
Selecting targets directly in the virtual world is difficult due to the lack of haptic feedback and inaccurate estimation of egocentric distances. Proprioception, the sense of self-movement and body position, can be utilized to improve virtual target selection by placing targets on or around one's body. However, its effective scope is limited closely around one's body. We explore the concept of virtually-extended proprioception by appending virtual body parts mimicking real body parts to users' avatars, to provide spatial reference to virtual targets. Our studies suggest that our approach facilitates more efficient target selection in VR as compared to no reference or using an everyday object as reference. Besides, by cultivating users' sense of ownership on the appended virtual body part, we can further enhance target selection performance. The effects of transparency and granularity of the virtual body part on target selection performance are also discussed.
Yang Tian 0008, Yuming Bai, Shengdong Zhao 0001, Chi-Wing Fu, Tianpei Yang, Pheng-Ann Heng
CHI1
2019 Bas-Relief Modeling from Normal Layers
abstract
Bas-relief is characterized by its unique presentation of intrinsic shape properties and/or detailed appearance using materials raised up in different degrees above a background. However, many bas-relief modeling methods could not manipulate scene details well. We propose a simple and effective solution for two kinds of bas-relief modeling (i.e., structure-preserving and detail-preserving) which is different from the prior tone mapping alike methods. Our idea originates from an observation on typical 3D models, which are decomposed into a piecewise smooth base layer and a detail layer in normal field. Proper manipulation of the two layers contributes to both structure-preserving and detail-preserving bas-relief modeling. We solve the modeling problem in a discrete geometry processing setup that uses normal-based mesh processing as a theoretical foundation. Specifically, using the two-step mesh smoothing mechanism as a bridge, we transfer the bas-relief modeling problem into a discrete space, and solve it in a least-squares manner. Experiments and comparisons to other methods show that (i) geometry details are better preserved in the scenario with high compression ratios, and (ii) structures are clearly preserved without shape distortion and interference from details.
Mingqiang Wei, Yang Tian 0008, Wai-Man Pang, Charlie C. L. Wang, Mingyong Pang, Jun Wang 0039, Harry Qin, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.2