Qixiang Ma

dblp:265/2401 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error Correction
abstract
Existing Grammatical Error Correction (GEC) systems suffer from limited reference diversity, leading to underestimated evaluation and restricted model generalization. To address this issue, we introduce the Judge of Edit-Level Validity (JELV), an automated framework to validate correction edits from grammaticality, faithfulness, and fluency. Using our proposed human-annotated Pair-wise Edit-level Validity Dataset (PEVData) as benchmark, JELV offers two implementations: a multi-turn LLM-as-Judges pipeline achieving 90% agreement with human annotators, and a distilled DeBERTa classifier with 85% precision on valid edits. We then apply JELV to reclassify misjudged false positives in evaluation and derive a comprehensive evaluation metric by integrating false positive decoupling and fluency scoring, resulting in state-of-the-art correlation with human judgments. We also apply JELV to filter LLM-generated correction candidates, expanding the BEA19's single-reference dataset containing 38,692 source sentences. Retraining top GEC systems on this expanded dataset yields measurable performance gains. JELV provides a scalable solution for enhancing reference diversity and strengthening both evaluation and model generalization.
Yuhao Zhan, Qixiang Ma, Zhiqi Yang, Fei Wu 0001
AAAI4
2026 Motion Hierarchical Gaussian for Dynamic Control in VR
abstract
Intuitive motion control is essential for virtual reality, allowing users to manipulate objects naturally while receiving realistic and responsive visual feedback. 3D Gaussian splatting provides real-time, photorealistic scene rendering, making it promising for virtual reality applications. Still, it falls short in accurate motion control of dynamic objects due to its unstructured global motion representation and redundant motion learning. To address these problems, we propose a motion hierarchical Gaussian based dynamic control method. First, a motion hierarchical Gaussian representation is introduced and initialized with semantic and deformation information. Then a motion hierarchical decomposition method is proposed to optimize the local motion in the representation. The representation is next optimized by a local motion analysis based refinement method. We also design a set of motion control operations for the motion hierarchical Gaussian. Experimental results show that our method achieves high-precision motion reconstruction, accurate motion decomposition, real-time, intuitively and immersive VR motion control.
Runze Fan, Jian Wu 0033, Qixiang Ma, Zhikai Wen, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.3
2026 Interaction-Aware Shared Scene Synthesis for VR Telepresence
Zhangyao Tan, Qixiang Ma, Runze Fan, Sio Kei Im, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.2
2025 A 3D Simulation Platform for Fuel Handling and Storage Systems
Qixiang Ma, Zhikai Wen, Min Zhang 0005, Jian Wu 0033, Lili Wang 0006
ICXR1
2025 Beyond strong labels: Weakly-supervised learning based on Gaussian pseudo labels for the segmentation of ellipse-like vascular structures in non-contrast CTs
Qixiang Ma, Adrien Kaladji, Huazhong Shu, Guanyu Yang 0001, Antoine Lucas, Pascal Haigron
Medical Image Anal.1
2025 Fov-GS: Foveated 3D Gaussian Splatting for Dynamic Scenes
abstract
Rendering quality and performance greatly affect the user's immersion in VR experiences. 3D Gaussian Splatting-based methods can achieve photo-realistic rendering with speeds of over 100 fps in static scenes, but the speed drops below 10 fps in monocular dynamic scenes. Foveated rendering provides a possible solution to accelerate rendering without compromising visual perceptual quality. However, 3DGS and foveated rendering are not compatible. In this paper, we propose Fov-GS, a foveated 3D Gaussian splatting method for rendering dynamic scenes in real time. We introduce a 3D Gaussian forest representation that represents the scene as a forest. To construct the 3D Gaussian forest, we propose a 3D Gaussian forest initialization method based on dynamic-static separation. Subsequently, we propose a 3D Gaussian forest optimization method based on deformation field and Gaussian decomposition to optimize the forest and deformation field. To achieve real-time dynamic scene rendering, we present a 3D Gaussian forest rendering method based on HVS models. Experiments demonstrate that our method not only achieves higher rendering quality in the foveal and salient regions compared to the SOTA methods but also dramatically improves rendering performance, achieving up to 11.33X speedup. We also conducted a user study, and the results prove that the perceptual quality of our method has a high visual similarity with the ground truth.
Runze Fan, Jian Wu 0033, Xuehuai Shi, Lizhi Zhao, Qixiang Ma, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.5
2025 SGSG: Stroke-Guided Scene Graph Generation
abstract
3D scene graph generation is essential for spatial computing in Extended Reality (XR), providing structured semantics for task planning and intelligent perception. However, unlike instance-segmentation-driven setups, generating semantic scene graphs still suffer from limited accuracy due to coarse and noisy point cloud data typically acquired in practice, and from the lack of interactive strategies to incorporate users' spatialized and intuitive guidance. We identify three key challenges: designing controllable interaction forms, involving guidance in inference, and generalizing from local corrections. To address these, we propose SGSG, a Stroke-Guided Scene Graph generation method that enables users to interactively refine 3D semantic relationships and improve predictions in real time. We propose three types of strokes and a lightweight SGstrokes dataset tailored for this modality. Our model integrates stroke guidance representation and injection for spatio-temporal feature learning and reasoning correction, along with intervention losses that combine consistency-repulsive and geometry-sensitive constraints to enhance accuracy and generalization. Experiments and the user study show that SGSG outperforms state-of-the-art methods 3DSSG and SGFN in overall accuracy and precision, surpasses JointSSG in predicate-level metrics, and reduces task load across all control conditions, establishing SGSG as a new benchmark for interactive 3D scene graph generation and semantic understanding in XR. Implementation resources are available at: https://github.com/Sycamore-Ma/SGSG-runtime.
Qixiang Ma, Runze Fan, Lizhi Zhao, Jian Wu 0033, Sio Kei Im, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.1
2024 ViP-Fluid: Visual Perception Driven Method for VR Fluid Rendering
abstract
The demand for fluid simulation and rendering in virtual reality (VR) is increasing. However, achieving high visual quality while maintaining real-time efficiency remains a challenge. Traditional foveated rendering methods balance the simulation quality in the foveated region but neglect the physical realism in the peripheral areas, and fail to account for the perceptual degradation caused by frame rate fluctuations during adaptive updates. To address these challenges, we propose a novel visual perception driven fluid rendering method ViP-Fluid, which further enhances rendering quality while balancing efficiency. Our approach employs a spatiotemporal saliency model for multi-granularity simulation and rendering of Lagrangian fluid systems, and introduces a Perception Threshold for Physical Process Elapsing (PTPE) metric, which guides our temporal acceleration strategy. Through a series of objective experiments, we demonstrate the advantages of our method in rendering quality and performance efficiency. ViP-Fluid demonstrates superior metrics not only in the foveated region but also in the salient and overall regions, achieving up to 2.15 times speed-up compared to the high-resolution Position Based Fluids (PBF) benchmark. Subsequent user experiments further validate the visual perception advantages of ViP-Fluid over both traditional and state-of-the-art methods, confirming the spatiotemporal fidelity of our acceleration strategy as well as a user preference for our approach.
Qixiang Ma, Jian Wu 0033, Runze Fan, Xuehuai Shi
ISMAR1
2024 Scene-aware Foveated Rendering
abstract
We propose a new scene-aware foveated rendering method, which incorporates the scene awareness and characteristics of the human visual system into the mapping-based foveated rendering framework. First, we generate the conservative visual importance map that encodes the visual features of the scene, visual acuity, and gaze motion. Second, we construct the pixel size control map using a convolution kernel method. Third, we utilize the pixel size control map to guide the foveated rendering. At last, a temporal coherent refinement strategy is used to maintain the smooth foveated rendering for the adjacent frames. Compared to the state-of-the-art mapping-based foveated rendering methods using the same compression ratio, our method achieves smaller MSE, higher PSNR, and SSIM in the fovea, periphery, salient regions, and the whole image. We also conducted user studies, and the results proved that the perceptual quality of our method has a high visual similarity with the around truth rendered with the full resolution.
Runze Fan, Xuehuai Shi, Kangyu Wang, Qixiang Ma, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.4
2024 SMigraPH: a perceptually retained method for passive haptics-based migration of MR indoor scenes
Qixiang Ma, Lili Wang 0006, Wei Ke 0001, Sio Kei Im
Vis. Comput.1
2023 Cloud-Edge-End Collaboration Personalized Semi-supervised Federated Learning for Visual Localization
abstract
Deep learning-based visual localization methods use convolutional neural networks to directly regress the position of a target. However, previous studies only consider localization in a single scene and neglect personalized localization in multiple scenes. Furthermore, changes in the scene result in a reduced accuracy due to the model’s lack of adaptability. Moreover, traditional centralized training methods pose data privacy concerns. In this paper, we propose a personalized semi-supervised federated learning framework with cloud-edge-end collaboration, called FedVL. The hierarchical architecture extends single-scene localization to multiple scenes, while the federated learning mechanism ensures data privacy. In this framework, we apply personalized federated learning to achieve scene-specific model and employ semi-supervised federated learning to allow the localization model to adapt to scene changes. Experiments conducted on indoor and outdoor datasets demonstrate the effectiveness of this approach.
Qixiang Ma, Zhe Zhang 0043, Zhenhan Zhu, Yanchao Zhao
ICPADS1
2023 Lambertian-based adversarial attacks on deep-learning-based underwater side-scan sonar image classification
Qixiang Ma, Longyu Jiang, Wenxue Yu
Pattern Recognit.1
2022 CrowdLoc: Robust image indoor localization with edge-assisted crowdsensing
Maoxing Tang, Yanchao Zhao, Qixiang Ma, Jiangshan Hao, Bing Chen 0002
J. Syst. Archit.3
2021 Disocclusion Headlight for Selection Assistance in VR
abstract
We introduce the disocclusion headlight, a method for VR selection assistance based on alleviating occlusions at the center of the user's field of view. The user's visualization of the VE is modified to reduce overlap between objects. This way, selection candidate objects have larger image footprints, which facilitates selection. The modification is confined to the center of the frame, with continuity to the periphery of the frame which is rendered conventionally. The selection assistance is provided automatically, without any interaction from the user. Furthermore, our method disoccludes without destroying the local spatial relationships between selection candidates, which allows solving complex selection queries based on the relative position of objects. We have tested our method on three selection tasks, where we compared it to two state-of-the-art VR selection techniques, i.e., the alpha cursor and the flower cone. Our method showed significant advantages in terms of shorter task completion times, and of fewer selection errors.
Lili Wang 0006, Qixiang Ma, Voicu Popescu
VR3
2020 Training with Noise Adversarial Network: A Generalization Method for Object Detection on Sonar Image
abstract
Object detection tasks for sonar image confront two major challenges, scarcity of dataset and perturbation of noise, which cause overfitting to models. The state-of-the-art object detection designed for optical images cannot address the issues because of the inherent differentiation between the optical image and sonar image. To tackle this problem, in this paper, we propose an adversarial training method to generalize the detector by introducing perturbation with specific noise property of sonar images during training stage. We design a sideway network which we name Noise Adversarial Network (NAN). The NAN is embedded into the state-of-the-art detector to generate adversarial examples which serve as assistant decision-making items to predict both class and bounding box, aiming to improve the generalization and noise robustness of the detector. To provide prior knowledge of noise perturbation to NAN, we also design a Noise Block (NB) for introducing noise in the upstream layers, which further improves noise robustness. Following the Faster R-CNN framework, the results of our experiments indicate a 8.9% mAP boost on our sonar image dataset. The detector equipped with NAN and NB also outperforms the baseline on noised test sets. Furthermore, it gains a 2.4% mAP boost on the optical image dataset PASCAL VOC 2007.
Qixiang Ma, Longyu Jiang, Wenxue Yu, Fangjin Xu
WACV1