Xinda Liu

dblp:191/2562 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0002-1981-7243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Walking in the Wild: Safe and Natural Redirected Walking in Open Physical Spaces
abstract
Redirected Walking (RDW) enables continuous locomotion in virtual environments (VEs) within limited physical spaces. However, classic RDW methods rely on static settings with tight boundaries, which can reduce their applicability in large shared physical spaces where boundary constraints are not dominant, especially under dynamic and multi-user conditions. To overcome this, we introduce a new task: Redirected Walking in Open Physical Spaces (OPSRDW), which allows users to navigate expansive VEs safely and naturally despite dynamic obstacles and without requiring a fixed boundary model. We further propose Dynamic Control and Redirection with Safety Constraints (DyCoRe), which formulates OPS-RDW as a constrained optimization problem. DyCoRe uses Dynamic Control Barrier Functions to model real-time collision avoidance constraints and solves an online Quadratic Programming problem to compute optimal velocities that minimize path deviation while ensuring safety. These velocities are mapped into real-time redirection gains, guiding users along natural paths while reducing collision risk. A relaxation mechanism is incorporated to handle infeasible scenarios. Extensive simulations suggest consistent improvements in safety and obstacle clearance over state-of-theart methods. User studies demonstrate that DyCoRe significantly improves navigation continuity, reduces the number of resets, and tends to reduce perceived discomfort compared to classic RDW strategies. DyCoRe provides an efficient and learning-free solution for safe and natural VR locomotion in open physical spaces with dynamic obstacles and multiple users.
Xinda Liu, Guoqiang Yang, Yunchen Li, Jian Wu 0033, Guohua Geng, Lili Wang 0006
VR1
2026 From SVG to DSVG: Leveraging Foveal Visual Cues to Mitigate Cybersickness in Redirected Walking
abstract
ABSTRACT Redirected Walking (RDW) effectively extends the navigable area of virtual environments but frequently induces cybersickness due to vestibular‐visual conflicts. To mitigate this, this study proposes a hierarchical visual guidance framework. We first introduce a Steering Visual Guidance (SVG) model, which employs a grid‐based pattern to provide a stable visual reference frame. While effective, our evaluation revealed that the static, full‐field nature of SVG can occlude the user's view and reduce visual clarity. To address this limitation, we optimized the model into Dynamic Steering Visual Guidance (DSVG). Grounded in the physiological principles of central and peripheral vision, DSVG dynamically renders stability cues exclusively within the foveal region, fading them in the periphery based on retinal sensitivity. A user study demonstrates that both SVG and DSVG significantly reduce subjective discomfort and physiological markers of sickness compared to a control condition without cues. Crucially, DSVG achieves sickness mitigation comparable to SVG while reducing visual occlusion by approximately 70%. These findings suggest that DSVG offers a robust, occlusion‐minimizing solution for deploying RDW in constrained physical spaces.
Xinda Liu, Yongbo Tang, Xiaoning Liu 0001, Pengbo Zhou, Guohua Geng
Comput. Animat. Virtual Worlds1
2026 Knowledge-injected prompt tuning with semantic regularization for fine-grained image recognition
Xinda Liu, Pengbo Zhou, Guohua Geng
Knowl. Based Syst.2
2026 Attention-guided multi-scale local reconstruction for point clouds via masked autoencoder self-supervised learning
Xin Cao 0004, Jiaxu Shi, Linzhi Su, Xinda Liu, Kang Li 0005
Multim. Syst.5
2026 Structured-condensed prompt tuning in vision-language models for fine-grained image recognition
Xinda Liu, Weiqing Min, Guohua Geng, Shuqiang Jiang
Pattern Recognit.1
2026 From Structure to Semantics: Hypergraph-Based AR Assembly Guidance with LLM-Mediated Narration
abstract
Effective Augmented Reality (AR) guidance for complex assembly faces a dual challenge: the inability of conventional liaison graphs to represent procedural logic, and the cognitive burden imposed by visual instructions. We argue that the solution requires a more expressive structure to overcome these representational deficits and a narration approach to mediate instruction complexity. Our method first employs an assembly hypergraph to capture the task's hierarchical information, from which an A* search algorithm generates an optimal assembly path. Then a Large Language Model (LLM)-mediated narration workflow is designed to address the ergonomic deficiencies of the machine-centric path. It employs an optimizer to improve fluency, followed by a narrator that crafts the steps into an intuitive instruction narration. A within-subjects user study (N = 24) revealed a progressive enhancement from our method's components. The transition from a liaison-graph baseline to the hypergraph alone improved objective outcomes by reducing task time and errors and improving subjective ratings (SUS, NASA-TLX, TAM, and ARI). Subsequently, augmenting the LLM-mediated narration maintained these gains while lowering cognitive load and elevating user experience and usability. Our findings indicate the value of our AR assembly design and discuss the opportunities of using LLM as a mediation layer for better user interaction.
Xinda Liu, Jiaju Xu, Jian Wu 0033, Guohua Geng, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.1
2025 CMFF: Cross-modal feature fusion network for robust point cloud completion
Pengbo Zhou, Xinda Liu, Longquan Yan, Guohua Geng
Neural Networks4
2025 Scene-Aware Foveated Neural Radiance Fields
abstract
Foveated rendering provides an idea for improving the image synthesis performance of neural radiance fields (NeRF) methods. In this article, we propose a scene-aware foveated neural radiance fields method to synthesize high-quality foveated images in complex VR scenes at high frame rates. First, we construct a multi-ellipsoidal neural representation to enhance the neural radiance field's representation capability in salient regions of complex VR scenes based on the scene content. Then, we introduce a uniform sampling based foveated neural radiance field framework to improve the foveated image synthesis performance with one-pass color inference, and improve the synthesis quality by leveraging the foveated scene-aware objective function. Our method synthesizes high-quality binocular foveated images at the average frame rate of 66 frames per second ($FPS$FPS) in complex scenes with high occlusion, intricate textures, and sophisticated geometries. Compared with the state-of-the-art foveated NeRF method, our method achieves significantly higher synthesis quality in both the foveal and peripheral regions with 1.41-1.46× speedup. We also conduct a user study to prove that the perceived quality of our method has a high visual similarity with the ground truth.
Xuehuai Shi, Lili Wang 0006, Xinda Liu, Jian Wu 0033, Zhiwen Shao
IEEE Trans. Vis. Comput. Graph.3
2024 Where Should a Virtual Guide Stand in a VR Museum?
Xinda Liu, Jian Wu 0033, Lili Wang 0006, Guohua Geng
ICXR1
2024 Multi-granularity sequence generation for hierarchical image classification
abstract
Hierarchical multi-granularity image classification is a challenging task that aims to tag each given image with multiple granularity labels simultaneously. Existing methods tend to overlook that different image regions contribute differently to label prediction at different granularities, and also insufficiently consider relationships between the hierarchical multi-granularity labels. We introduce a sequence-to-sequence mechanism to overcome these two problems and propose a multi-granularity sequence generation (MGSG) approach for the hierarchical multi-granularity image classification task. Specifically, we introduce a transformer architecture to encode the image into visual representation sequences. Next, we traverse the taxonomic tree and organize the multi-granularity labels into sequences, and vectorize them and add positional information. The proposed multi-granularity sequence generation method builds a decoder that takes visual representation sequences and semantic label embedding as inputs, and outputs the predicted multi-granularity label sequence. The decoder models dependencies and correlations between multi-granularity labels through a masked multi-head self-attention mechanism, and relates visual information to the semantic label information through a cross-modality attention mechanism. In this way, the proposed method preserves the relationships between labels at different granularity levels and takes into account the influence of different image regions on labels with different granularities. Evaluations on six public benchmarks qualitatively and quantitatively demonstrate the advantages of the proposed method. Our project is available at https://github.com/liuxindazz/mgsg .
Xinda Liu, Lili Wang 0006
Comput. Vis. Media1
2024 Real-scene-constrained virtual scene layout synthesis for mixed reality
Runze Fan, Lili Wang 0006, Xinda Liu, Sio Kei Im, Chan-Tong Lam
Vis. Comput.3
2023 Feature-Suppressed Contrast for Self-Supervised Food Pre-training
abstract
Most previous approaches for analyzing food images have relied on extensively annotated datasets, resulting in significant human labeling expenses due to the varied and intricate nature of such images. Inspired by the effectiveness of contrastive self-supervised methods in utilizing unlabelled data, weiqing explore leveraging these techniques on unlabelled food images. In contrastive self-supervised methods, two views are randomly generated from an image by data augmentations. However, regarding food images, the two views tend to contain similar informative contents, causing large mutual information, which impedes the efficacy of contrastive self-supervised learning. To address this problem, we propose Feature Suppressed Contrast (FeaSC) to reduce mutual information between views. As the similar contents of the two views are salient or highly responsive in the feature map, the proposed FeaSC uses a response-aware scheme to localize salient features in an unsupervised manner. By suppressing some salient features in one view while leaving another contrast view unchanged, the mutual information between the two views is reduced, thereby enhancing the effectiveness of contrast learning for self-supervised food pre-training. As a plug-and-play module, the proposed method consistently improves BYOL and SimSiam by 1.70% ~ 6.69% classification accuracy on four publicly available food recognition datasets. Superior results have also been achieved on downstream segmentation tasks, demonstrating the effectiveness of the proposed method.
Xinda Liu, Linhu Liu, Jiang Tian, Lili Wang 0006
ACM Multimedia1
2022 Monitoring of Surface Intraday Frozen Duration Based on Multi-Temporal Passive Microwave Remote Sensing Observations
abstract
The surface freezing-thawing cycle is an important indicator of global climate change. In current studies, more attention has been paid on the monitoring of daily freezing-thawing changes, but less on the intra-day freezing-thawing changes. While the intraday freeze-thaw cycles are more sensitive to climate change. Thus, in this study, the multi-temporal observations derived from DMSP F15, F16, F17 and F18 in the period of 2015.01.01~2016.12.31 were used to monitoring the intraday freezing and thawing varations. A unified discriminant equation for four kinds of satellites was established. The ground soil temperature measurements from four networks (Naqu, Maqu, Pali and Ali) in the Qinghai-Tibet Plateau were used for the validation. The results show that there is a strong consistency between the intraday frozen duration calculated based on the F/T discrimination results and the in-situ measurements, with a maximum standard error of 12.6 hours in Ali and a minimum standard error of 4.34 hours in Naqu.
Xiaokang Kou, Zhenyi Zu, Shuaibing Han, Xinda Liu, Tianliang Wang
IGARSS6
2022 Distant Object Manipulation with Adaptive Gains in Virtual Reality
abstract
Object Manipulation is a fundamental interaction in virtual reality (VR). The efficiency and accuracy of object manipulation are important to provide immersion to users. We propose a manipulation method with adaptive gains to improve the efficiency and accuracy of object manipulation in VR applications. First, we introduce manipulation gains. We then design an experiment to collect user behavior during manipulation to determine fitting functions for calculating manipulation gain. At last, we design a user study to evaluate the performance of our distant object manipulation method with adaptive gains. The results show that, compared with the state of the art methods, our method has a significant improvement in the completion time, and the manipulation accuracy of the tasks. Moreover, our method significantly increases usability and reduces task load.
Lili Wang 0006, Shuai Luan, Xuehuai Shi, Xinda Liu
ISMAR5
2022 Transformer with peak suppression and knowledge guidance for fine-grained image recognition
Xinda Liu, Lili Wang 0006, Xiaoguang Han 0001
Neurocomputing1
2021 A Study on the Detection of Deformation of Tuotuohe Area on the Qinghai-Tibet Plateau
abstract
In permafrost regions, the ground surface deformation is closely related to the ice-water phase transition process in active layer and underground ice. It is of great significance to carry out ground surface deformation detection research for understanding the development of permafrost in the Qinghai- Tibet Plateau. The InSAR technology has been proved to be an effective method for monitoring frozen soil deformation. However, due to the limitation of spatial coverage and revisit cycle of SAR data, few scholars had paid attention on the unstable permafrost regions with less ice content in the previous studies. Thus, a less ice content permafrost region loceted in Tuotuohe was taken as the study area to carry out surface deformation detection research based on SBAS-InSAR technology, and the field measurement was carried out to do the verification. The result showed that there was a good agreement between them, and the errors are 0.9mm, 2.6mm and 2.8mm respectively. The deformation detected by SBAS-InSAR method well reflects the frost heaving and thaw subsidence trend along with seasonal changes. Considering the small content of underground ice, this trend mainly reflects the seasonal freezing-thawing variation of the active layer. This study further confirms the applicability of InSAR technology in unstable permafrost regions.
Xiaokang Kou, Xinda Liu, Tianliang Wang, Shuang Yan
IGARSS2
2021 Plant Disease Recognition: A Large-Scale Benchmark Dataset and a Visual Region and Loss Reweighting Approach
abstract
Plant disease diagnosis is very critical for agriculture due to its importance for increasing crop production. Recent advances in image processing offer us a new way to solve this issue via visual plant disease analysis. However, there are few works in this area, not to mention systematic researches. In this paper, we systematically investigate the problem of visual plant disease recognition for plant disease diagnosis. Compared with other types of images, plant disease images generally exhibit randomly distributed lesions, diverse symptoms and complex backgrounds, and thus are hard to capture discriminative information. To facilitate the plant disease recognition research, we construct a new large-scale plant disease dataset with 271 plant disease categories and 220,592 images. Based on this dataset, we tackle plant disease recognition via reweighting both visual regions and loss to emphasize diseased parts. We first compute the weights of all the divided patches from each image based on the cluster distribution of these patches to indicate the discriminative level of each patch. Then we allocate the weight to each loss for each patch-label pair during weakly-supervised training to enable discriminative disease part learning. We finally extract patch features from the network trained with loss reweighting, and utilize the LSTM network to encode the weighed patch feature sequence into a comprehensive feature representation. Extensive evaluations on this dataset and another public dataset demonstrate the advantage of the proposed method. We expect this research will further the agenda of plant disease recognition in the community of image processing.
Xinda Liu, Weiqing Min, Shuhuan Mei, Lili Wang 0006, Shuqiang Jiang
IEEE Trans. Image Process.1
2017 Modality-specific and hierarchical feature learning for RGB-D hand-held object recognition
Xiong Lv, Xinda Liu, Xiangyang Li 0002, Shuqiang Jiang, Zhiqiang He 0002
Multim. Tools Appl.2
2017 Being a Supercook: Joint Food Attributes and Multimodal Content Modeling for Recipe Retrieval and Exploration
abstract
This paper considers the problem of recipe-oriented image-ingredient correlation learning with multi-attributes for recipe retrieval and exploration. Existing methods mainly focus on food visual information for recognition while we model visual information, textual content (e.g., ingredients), and attributes (e.g., cuisine and course) together to solve extended recipe-oriented problems, such as multimodal cuisine classification and attribute-enhanced food image retrieval. As a solution, we propose a multimodal multitask deep belief network ($\mathrm{M}^{3}$TDBN) to learn joint image-ingredient representation regularized by different attributes. By grouping ingredients into visible ingredients (which are visible in the food image, e.g., “chicken” and “mushroom”) and nonvisible ingredients (e.g., “salt” and “oil”),$\mathrm{M}^{3}$TDBN is capable of learning both midlevel visual representation between images and visible ingredients and nonvisual representation. Furthermore, in order to utilize different attributes to improve the intermodality correlation,$\mathrm{M}^{3}$TDBN incorporates multitask learning to make different attributes collaborate each other. Based on the proposed$\mathrm{M}^{3}$TDBN, we exploit the derived deep features and the discovered correlations for three extended novel applications: 1) multimodal cuisine classification; 2) attribute-augmented cross-modal recipe image retrieval; and 3) ingredient and attribute inference from food images. The proposed approach is evaluated on the constructed Yummly dataset and the evaluation results have validated the effectiveness of the proposed approach.
Weiqing Min, Shuqiang Jiang, Huayang Wang, Xinda Liu, Luis Herranz
IEEE Trans. Multim.5
2016 RGB-D scene classification via heterogeneous model fusion
abstract
We study the problem of scene classification for RGB-D images in this paper. Firstly we analyze the difference between the RGB and depth images. And then based on the difference, an efficient method is implemented to make use of the RGB and depth images and make a well fusion for the RGB and depth features. Focusing on the difference of modality between the RGB and depth images, we propose a method to learn features from color and depth separately using the heterogeneous model. Especially we use the deep ConvNet model with shallow finetuning for RGB images and the relatively shallow ConvNet model with deep finetuning, which can adequately extract different characteristics of the two modalities. After obtaining the discriminative features for each modality, a multiple fully-connected layers connected with a soft-max classifier is trained to harness the complementary relationship between the two modalities. Experimental evaluations on two publicly RGB-D datasets validate the effectiveness of the proposed method.
Xinda Liu, Xueming Wang, Shuqiang Jiang
ICIP1