VLDB 2026 Research / reviewers in the wild / expert
Xinqi Liu
dblp:72/9634
· DBLP profile ↗
22ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Box-Loc: Box-aware module localization using point cloud and deep learning in modular integrated construction
Xinqi Liu |
Adv. Eng. Informatics | 2 |
| 2026 | Analogy-Augmented Uncertainty-Aware Monocular Visual OdometryabstractVisual odometry (VO) is a critical component of autonomous robot systems, enabling precise pose estimation from visual inputs. Learning-based VO methods are increasingly recognized for their robustness in challenging scenarios, including dynamic environments, motion blur, and low-light conditions. However, their performance is constrained by both the diversity of the data and its utilization rate. To overcome these limitations, we propose an end-to-end monocular VO system incorporating a novel learning-based end-to-end VO framework and multiple analogy augmentation strategies. We introduce the Context Attention Uncertainty-aware VO Network (CUVO), which prioritizes semantically rich regions and mitigating interference from high-uncertainty areas to enhance attentional focus and pose estimation accuracy. Furthermore, our analogy augmentation methods—temporal reversal, random rotation, and geometric mirroring—enhance image pairs and compute corresponding true pose transformations, significantly increasing training data quantity and diversity. Simultaneously, an analogous loss is applied to ensure consistency between the original and augmented data. Extensive experiments demonstrate that CUVO significantly enhances VO performance, outperforming previous end-to-end VO methods on TartanAir and KITTI datasets. By leveraging analogy augmentation strategy to expand training data under limited data conditions (27k), zero-shot capability of CUVO degrades by up to 29.5% on TartanAir and 23.3% on KITTI. Our work introduces the first image-to-pose data augmentation method tailored for VO and establishes CUVO as a robust system for advancing learning-based visual odometry. Jituo Li, Shunwang Sun, Tingxi Xue, Xinqi Liu, Jialu Zhang 0006, Huixu Dong, Guodong Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | An Offline RL-Based Dialogue System for Early Alzheimer's Detection: A Case Study on the Cookie Theft Picture TestabstractIt is of great importance to develop a dialogue system that captures linguistic features for the non-invasive early detection of populations at risk for Alzheimer's Disease (AD). This paper implements an LLM-based automated Cookie Theft Picture (CTP) Test dialogue system, developed using an offline reinforcement self-training framework. We first fine-tune a large language model on a limited real-world clinician-patient dialogue dataset. To address the problems in the initial model, such as asking irrelevant or overly general questions, or failing to probe further when necessary, we introduce an offline reinforcement self-training framework that alternates between generating highquality multi-turn dialogues and refining the model, enabling iterative self-improvement. To support AD detection, a goal-aware reward function is designed and integrated into the framework to guide the model in asking more focused and purposeful questions. In addition, we enhance the sample filtering mechanism and incorporate Direct Preference Optimization (DPO) for dialogue policy optimization, further reducing the generation of ineffective or off-target dialogues. A patient simulator is also developed to support the dialogue training process. Experiments conducted on a public dataset demonstrate that our method achieves an end-to-end AD detection accuracy of 94.80 %, significantly outperforming multiple baseline models and approaching the performance of the MMSE scale$(96.24 \%)$on the same dataset. Xinqi Liu |
BIBM | 2 |
| 2025 | VGA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos
Xinqi Liu, Chenming Wu |
CVM (2) | 1 |
| 2025 | MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingabstractThe vision-language tracking task aims to perform object tracking based on various modality references. Existing Transformer-based vision-language tracking methods have made remarkable progress by leveraging the global modeling ability of self-attention. However, current approaches still face challenges in effectively exploiting the temporal information and dynamically updating reference features during tracking. Recently, the State Space Model (SSM), known as Mamba, has shown astonishing ability in efficient long-sequence modeling. Particularly, its state space evolving process demonstrates promising capabilities in memorizing multimodal temporal information with linear complexity. Witnessing its success, we propose a Mamba-based vision-language tracking model to exploit its state space evolving ability in temporal space for robust multimodal tracking, dubbed MambaVLT. In particular, our approach mainly integrates a time-evolving hybrid state space block and a selective locality enhancement block, to capture contextual information for multimodal modeling and adaptive reference feature update. Besides, we introduce a modality-selection module that dynamically adjusts the weighting between visual and language references, mitigating potential ambiguities from either reference type. Extensive experimental results show that our method performs favorably against state-of-the-art trackers across diverse benchmarks. Xinqi Liu, Li Zhou 0017, Zikun Zhou, Jianqiu Chen, Zhenyu He 0001 |
CVPR | 1 |
| 2025 | VisitFrequency-Diffusion: Leveraging Recurrent Visits for Long-Term Individual Trajectory ForecastingabstractIndividual trajectory prediction plays a crucial role in intelligent transportation systems. While existing methods demonstrate strong performance in short-term forecasting (e.g., minute-level predictions), they are limited in modeling long-term patterns (day-level predictions). The key challenge is capturing both the periodic regularity and stochastic variability of urban mobility. To bridge this gap, we propose VF-Diffusion, a novel framework for long-term individual trajectory prediction with three key innovations: (1) A direction-sensitive diffusion model that generates baseline trajectories by learning motion trends; (2) A trajectory rectification module that refines spatial displacements using historical median coordinates; and (3) A frequency-sensitive mechanism that identifies high-frequency visit locations, predicts their temporal sequences via an ensemble model, and integrates them with the baseline trajectory. By combining generative modeling with a frequency-sensitive mechanism, VF-Diffusion fills a critical gap in existing methods, offering the ability to predict new visiting areas and improve trajectory accuracy. Extensive experiments on Beijing Wi-Fi trajectory data show that our method outperforms four baselines, achieving about 90% accuracy for predictions within a 1 km threshold. It particularly excels in areas with frequent and periodic visits. This framework advances trajectory prediction by enabling multi-day forecasting, a previously underexplored capability, and offers practical solutions for enhancing smart city infrastructure. Shuhui Gong, Xinqi Liu, Jiahao Lv 0003, Jilin Hu, Hongbin Pei |
SIGSPATIAL/GIS | 3 |
| 2025 | DAFU-CAD: Depth-assisted Feature Unraveling for Sketch-based Robust CAD ModelingabstractSketching is a quick ideation and multimedia tool for effectively expressing design intent. By translating simple strokes into CAD models, it allows non-expert users to create editable designs, reducing the learning curve associated with traditional CAD software. However, current sketch-based CAD modeling methods are often limited to basic shapes and require structured inputs, making them less robust when dealing with varied sketch styles. To overcome these challenges, we propose a novel sketch-based modeling framework DAFU-CAD, that is both efficient and robust. Our approach features a Depth-Assisted and Feature-Unraveling sketch classification module that categorizes sketches into corresponding modeling operations, independent of their drawing style. A parameter regression and optimization module then estimates the modeling parameters, ensuring consistent and stable model reconstruction across different sketch inputs. To support this, we compile a diverse sketch dataset with a range of modeling categories and abstraction levels. Experimental results show that our method outperforms existing approaches in terms of both robustness and versatility. Xinqi Liu, Zhiliang He, Jialu Zhang 0006, Chenming Wu, Guodong Lu, Jituo Li |
ACM Multimedia | 2 |
| 2025 | Online semantic embedding correlation for discrete cross-media hashing
Fan Yang 0071, Fumin Ma, Xiaojian Ding, Xinqi Liu |
Expert Syst. Appl. | 6 |
| 2025 | Fault resilient on-device batched DNN inference for mobile devices with ARM TrustZone
Longlong Liao, Wenbin Zeng, Xinqi Liu, Yuanlong Yu 0001 |
J. Syst. Archit. | 3 |
| 2025 | Online Asymmetric Supervised Discrete Cross-Modal Hashing for Streaming Multimedia Data
Fan Yang 0071, Xinqi Liu, Fumin Ma, Xiaojian Ding, Kaixiang Wang 0001 |
Pattern Recognit. | 2 |
| 2025 | Learning Pose Controllable Human Reconstruction With Dynamic Implicit Fields From a Single ImageabstractRecovering a user-special and controllable human model from a single RGB image is a nontrivial challenge. Existing methods usually generate static results with an image consistent subject's pose. Our work aspires to achieve pose-controllable human reconstruction from a single image by learning a dynamic (multi-pose) implicit field. We first construct a feature-embedded human model (FEHM) as a bridge to propagate image features to different pose spaces. Based on FEHM, we then encode three pose-decoupled features. Global image features represent user-specific shapes in images and replace widely used pixel-aligned ways to avoid unwanted shape-pose entanglement. Spatial color features propagate FEHM-embedded image cues into 3D pose space to provide spatial high-frequency guidance. Spatial geometry features improve reconstruction robustness by using the surface shape of the FEHM as the prior. Finally, new implicit functions are designed to predict the dynamic human implicit fields. For effective supervision, a realistic human avatar dataset, SimuSCAN, with 1000+ models is constructed using a low-cost hierarchical mesh registration method. Extensive experiments demonstrate that our method achieves the state-of-the-art reconstruction level. Jituo Li, Xinqi Liu, Guodong Lu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Reconstructing Complex Shaped Clothing From a Single Image With Feature Stable Unsigned Distance FieldsabstractSingle-view clothing reconstruction usually relies on topologically fixed clothing templates to reduce the problem complexity, but this strategy also makes the reconstructed clothing shape contours simple and lack diversity. In this article, we propose a novel clothing reconstruction method to generate complex shape contours and open clothing mesh from a single image. At the heart of our work is an implicit unsigned distance field condition on clothing-oriented and pose-stable spatial shape features to represent the clothing from the image. This feature can provide spatially aligned clothing shape priors to improve the pose robustness. It is based on a type-generic clothing template derived from the mainstream clothing generative model to avoid tedious template design and switching. To output open clothing mesh results from noisy clothing unsigned distance fields, we develop a two-stage clothing mesh extraction method. It takes the point clouds as an intermediate representation and produces smooth, plausible and editable clothing mesh results. To provide effective supervision, we construct a pose-rich and shape-complete clothing scan dataset by enhancing clothing pose diversity and complementing missing clothing geometry caused by occlusion. Extensive experiments demonstrate that our method achieves state-of-the-art levels. More importantly, we provide a simple but effective, and low-cost way to reconstruct complex shape contours clothing from a single image. Xinqi Liu, Jituo Li, Guodong Lu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Sketch2Seq: Reconstruct CAD Models From Feature-Based Sketch SegmentationabstractSketch-based modeling studies reconstructing models from sketches automatically, allowing users visualize design concepts rapidly. Generating CAD models based on user sketches helps reduce the learning curve for novice users, which promotes the everyday use of CAD software, and expands its reach to non-professional groups. While various algorithms study automatically generating models from single sketch or line drawing, they often produce non-editable models or editable models limited to simple extrusion operations. To improve this issue, we propose a novel sketch-based modeling system, Sketch2Seq, which generates complex, semantic, and editable CAD models. Our system eliminates the need for additional annotations from users and produces models that support subsequent application in commercial software. The core of our method lies in understanding users' design intent from CAD sketches. We design a novel sketch segmentation network for identifying diverse operation features in CAD sketches, which utilizes geometric features of strokes and different levels of topological connections. Additionally, to tackle the segmentation task, a dataset for CAD sketch segmentation is introduced. Comparative experiments and ablation evaluations prove the effectiveness of the proposed method. Based on segmentation result, coarse CAD sequences are generated and progressively executed. Meanwhile, the orders and parameters of the CAD sequences are optimized with context models and input sketches. All algorithms are integrated into a user interface. Experiments and evaluations validate the feasibility and superiority of our entire system which is able to reconstruct more complex features and achieve better results for longer sequence. Jituo Li, Ziqin Xu, Jialu Zhang 0006, Xinqi Liu, Guodong Lu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | TexOct: Generating Textures of 3D Models with Octree-based DiffusionabstractThis paper focuses on synthesizing high-quality and complete textures directly on the surface of 3D models within 3D space. 2D diffusion-based methods face challenges in generating 2D texture maps due to the infinite possibilities of UV mapping for a given 3D mesh. Utilizing point clouds helps circumvent variations arising from diverse mesh topologies and UV mappings. Nevertheless, achieving dense point clouds to accurately represent texture details poses a challenge due to limited computational resources. To address these challenges, we propose an efficient octree-based diffusion pipeline called TexOct. Our method starts by sampling a point cloud from the surface of a given 3D model, with each point containing texture noise values. We utilize an octree structure to efficiently represent this point cloud. Additionally, we introduce an innovative octree-based diffusion model that leverages the denoising capabilities of the Denoising Diffusion Probabilistic Model (DDPM). This model gradually reduces the texture noise on the octree nodes, resulting in the restoration of fine texture. Experimental results on ShapeNet demonstrate that TexOct effectively generates high-quality 3D textures in both unconditional and text / image-conditional scenarios. Jialun Liu, Chenming Wu, Xinqi Liu, Haotian Peng, Chen Zhao 0011, Haocheng Feng, Jingtuo Liu, Errui Ding |
CVPR | 3 |
| 2024 | A Study of Slope Path Tracking for Tracked Vehicles in Hilly Mountainous AreasabstractThe paper proposes a tracked vehicle slope path track control algorithm based on MPC hilly mountainous area motor differential control. First, a tracked vehicle slope driving slip model is established to achieve the accurate estimation slope position. Second, the desired acceleration and desired angular acceleration are obtained as inputs by incorporating the intended linear and angular velocities, and the actual motivation required of the tracked vehicle when steering the vehicle on the slope is estimated by the MPC, and then the driving force is obtained as the actual control signal through the road surface response and kinematics solving to achieve the slope control of tracked vehicles. The actual control signal is obtained from the driving force through the road surface response and kinematics solution to realize the tracked vehicle slope path tracking control. The test outcomes demonstrate that our path tracking algorithm exhibits superior performance compared to the proportion integral differential (PID) control method during the vehicle's travel on a slope. Boyang Wang 0009, Zhaoguo Zhang, Faan Wang, Xinqi Liu, Kaiting Xie, Chang Ni |
INDIN | 4 |
| 2024 | Fault-tolerant deep learning inference on CPU-GPU integrated edge devices with TEEs
Hongjian Xu, Longlong Liao, Xinqi Liu, Shuguang Chen, Zhixuan Liang, Yuanlong Yu 0001 |
Future Gener. Comput. Syst. | 3 |
| 2024 | Modeling Realistic Clothing From a Single Image Under Normal GuideabstractWe propose a robust and highly realistic clothing modeling method to generate a 3D clothing model with visually consistent clothing style and wrinkles distribution from a single RGB image. Notably, this entire process only takes a few seconds. Our high-quality clothing results benefit from the idea of combining learning and optimization, making it highly robust. First, we use the neural networks to predict the normal map, a clothing mask, and a learning-based clothing model from input images. The predicted normal map can effectively capture high-frequency clothing deformation from image observations. Then, by introducing a normal-guided clothing fitting optimization, the normal maps are used to guide the clothing model to generate realistic wrinkles details. Finally, we utilize a clothing collar adjustment strategy to stylize clothing results using predicted clothing masks. An extended multi-view version of the clothing fitting is naturally developed, which can further improve the realism of the clothing without tedious effort. Extensive experiments have proven that our method achieves state-of-the-art clothing geometric accuracy and visual realism. More importantly, it is highly adaptable and robust to in-the-wild images. Further, our method can be easily extended to multi-view inputs to improve realism. In summary, our method can provide a low-cost and user-friendly solution to achieve realistic clothing modeling. Xinqi Liu, Jituo Li, Guodong Lu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Wrinkles Realistic Clothing Reconstruction by Combining Implicit and Explicit Method
Xinqi Liu, Jituo Li, Guodong Lu |
Comput. Aided Des. | 1 |
| 2023 | Generating High-Fidelity Texture in RGB-D Reconstruction using Patches Density Regularization
Xinqi Liu, Jituo Li, Guodong Lu |
Comput. Aided Des. | 1 |
| 2023 | Robust and automatic clothing reconstruction based on a single RGB image
Xinqi Liu, Jituo Li, Guodong Lu, Shihai Xing |
Comput. Graph. | 1 |
| 2023 | Improving RGB-D-based 3D reconstruction by combining voxels and points
Xinqi Liu, Jituo Li, Guodong Lu |
Vis. Comput. | 1 |
| 2022 | Reconstruction of Colored Soft Deformable Objects Based on Self-Generated Template
Jituo Li, Xinqi Liu, Haijing Deng, Guodong Lu, Jin Wang 0015 |
Comput. Aided Des. | 2 |