Yuqing Lan

dblp:11/5833 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Label noise learning based SAR target classification method
Hongqiang Wang 0004, Yuqing Lan, Fuzhan Yue, Zhenghuan Xia, Tao Zhang 0023
Neural Networks2
2026 RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction
abstract
The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it substantially improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the whole mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes. Project page can be found at https://lanlan96.github.io/RemixFusion/ .
Yuqing Lan, Chenyang Zhu 0002, Shuaifeng Zhi, Jiazhao Zhang, Zhoufeng Wang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004
ACM Trans. Graph.1
2025 OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask Merging
abstract
Online zero-shot 3D instance segmentation of a progressively reconstructed scene is both a critical and challenging task for embodied applications. With the success of visual foundation models (VFMs) in the image domain, leveraging 2D priors to address 3D online segmentation has become a prominent research focus. Since segmentation results provided by 2D priors often require spatial consistency to be lifted into final 3D segmentation, an efficient method for identifying spatial overlap among 2D masks is essential—yet existing methods rarely achieve this in real time, mainly limiting its use to offline approaches. To address this, we propose an efficient method that lifts 2D masks generated by VFMs into a unified 3D instance using a hashing technique. By employing voxel hashing for efficient 3D scene querying, our approach reduces the time complexity of costly spatial overlap queries from O(n2) to O(n). Accurate spatial associations further enable 3D merging of 2D masks through simple similarity-based filtering in a zero-shot manner, making our approach more robust to incomplete and noisy data. Evaluated on the ScanNet200 and SceneNN benchmarks, our approach achieves state-of-the-art performance in online, zero-shot 3D instance segmentation with leading efficiency. The project page is at https://yjtang249.github.io/OnlineAnySeg.
Jiazhao Zhang, Yuqing Lan, Yulan Guo, Dezun Dong, Chenyang Zhu 0002, Kai Xu 0004
CVPR3
2025 Unsupervised Fact Error Correction Modeling by Using Span-Level Contrastive Learning
Yuqing Lan, Zhenghao Liu 0001, Yu Gu 0002, Ge Yu 0001
DASFAA (2)1
2025 Fitness-Driven Evolutionary Federated Learning: Adaptive Client Selection and Dynamic Population for Communication Efficiency
Yichun Yu, Yuqing Lan, Zhihuan Xing
ICONIP (4)2
2025 Enhancing Person Re-Identification with TF-SS: A Transformer Model based on Shifted Windows Division and Skipping Sub-Layer Techniques
abstract
Person re-identification (Re-ID) involves recognizing a specific individual from images captured across different devices, which has splendid application prospects. Due to factors such as motion changes and target occlusion, the accuracy of person re-identification algorithms in complex scenes will be reduced, and existing algorithms fail to solve such problems properly. On this basis, this paper proposes a person re-identification model named Transformer Shifted Window with Skipping Sub-Layer (TF-SS) based on the shifted window division and skipping sub-layer methods. Firstly, a relative position encoding structure is introduced into the Transformer to improve the model's ability to capture spatio-temporal features. Secondly, the standard multi-head self-attention module of the Transformer is replaced with a multi-head self-attention module based on shifted windows, which can perform better dense predictions on hierarchical feature maps with linear computational complexity. Then, an adversarial training algorithm is introduced into the semantic embedding layer to enhance the robustness of the model. Finally, the skipping sublayer method is introduced. By randomly omitting sub-layers to introduce perturbations into the training, a greater constraint effect is exerted on the sub-layers, further enhancing the accuracy of the model. In this paper, the effectiveness of TF-SS is verified through extensive ablation experiments and comparative analyses, and it outperforms the current most advanced person re-identification methods in terms of mean Average Precision (mAP) and Rank@5 and Rank@10 accuracies.
Yuqing Lan, Xinyan Lu 0002
IJCNN2
2025 BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion
abstract
Abstract Open‐vocabulary 3D object detection has gained significant interest due to its critical applications in autonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial computational overhead and memory constraints, hindering real‐time deployment in downstream tasks. To address this, we propose a novel reconstruction‐free online framework tailored for memory‐efficient and real‐time 3D detection. Specifically, given streaming posed RGB‐D video input, we leverage Cubify Anything as a pre‐trained visual foundation model (VFM) for single‐view 3D object detection, coupled with CLIP to capture open‐vocabulary semantics of detected objects. To fuse all detected bounding boxes across different views into a unified one, we employ an association module for correspondences of multi‐views and an optimization module to fuse the 3D bounding boxes of the same instance. The association module utilizes 3D Non‐Maximum Suppression (NMS) and a box correspondence matching module. The optimization module uses an IoU‐guided efficient random optimization technique based on particle filtering to enforce multi‐view consistency of the 3D bounding boxes while minimizing computational complexity. Extensive experiments on CA‐1M and ScanNetV2 datasets demonstrate that our method achieves state‐of‐the‐art performance among online methods. Benefiting from this novel reconstruction‐free paradigm for 3D object detection, our method exhibits great generalization abilities in various scenarios, enabling real‐time perception even in environments exceeding 1000 square meters.
Yuqing Lan, Chenyang Zhu 0002, Zhirui Gao, Jiazhao Zhang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004
Comput. Graph. Forum1
2025 Multi-Evidence Based Fact Verification via A Confidential Graph Neural Network
abstract
Fact verification tasks aim to identify the integrity of textual contents according to the truthful corpus. Existing fact verification models usually build a fully connected reasoning graph, which regards claim-evidence pairs as nodes and connects them with edges. They employ the graph to propagate the semantics of the nodes. Nevertheless, the noisy nodes usually propagate their semantics via the edges of the reasoning graph, which misleads the semantic representations of other nodes and amplifies the noise signals. To mitigate the propagation of noisy semantic information, we introduce a Confidential Graph Attention Network (CO-GAT), which proposes a node masking mechanism for modeling the nodes. Specifically, CO-GAT calculates the node confidence score by estimating the relevance between the claim and evidence pieces. Then, the node masking mechanism uses the node confidence scores to control the noise information flow from the vanilla node to the other graph nodes. CO-GAT achieves a 73.59% FEVER score on the FEVER dataset and shows the generalization ability by broadening the effectiveness to the science-specific domain.
Yuqing Lan, Zhenghao Liu 0001, Yu Gu 0002, Xiaoyuan Yi, Xiaohua Li 0004, Liner Yang, Ge Yu 0001
IEEE Trans. Big Data1
2024 BDEL: A Backdoor Attack Defense Method Based on Ensemble Learning
Zhihuan Xing, Yuqing Lan, Yin Yu, Yichun Yu
PRICAI (1)2
2024 LLM-enhanced Scene Graph Learning for Household Rearrangement
Qijin She, Zhinan Yu, Yuqing Lan, Chenyang Zhu 0002, Ruizhen Hu, Kai Xu 0004
SIGGRAPH Asia5
2022 DisARM: Displacement Aware Relation Module for 3D Detection
abstract
We introduce Displacement Aware Relation Module (DisARM), a novel neural network module for enhancing the performance of 3D object detection in point cloud scenes. The core idea is extracting the most principal contextual information is critical for detection while the target is incomplete or featureless. We find that relations between proposals provide a good representation to describe the context. However, adopting relations between all the object or patch proposals for detection is inefficient, and an imbalanced combination of local and global relations brings extra noise that could mislead the training. Rather than working with all relations, we find that training with relations only between the most representative ones, or an-chors, can significantly boost the detection performance. Good anchors should be semantic-aware with no ambiguity and able to describe the whole layout of a scene with no redundancy. To find the anchors, we first perform a preliminary relation anchor module with an objectness-aware sampling approach and then devise a displacement based module for weighing the relation importance for better utilization of contextual information. This lightweight relation module leads to significantly higher accuracy of object instance detection when being plugged into the state-of-the-art detectors. Evaluations on the public benchmarks of real-world scenes show that our method achieves the state-of-the-art performance on both SUN RGB-D and Scan-Net V2. The code and models are publicly available at https://github.com/YaraDuan/DisARM.
Yao Duan, Chenyang Zhu 0002, Yuqing Lan, Renjiao Yi, Xinwang Liu 0002, Kai Xu 0004
CVPR3
2022 ARM3D: Attention-based relation module for indoor 3D object detection
abstract
Relation contexts have been proved to be useful for many challenging vision tasks. In the field of 3D object detection, previous methods have been taking the advantage of context encoding, graph embedding, or explicit relation reasoning to extract relation contexts. However, there exist inevitably redundant relation contexts due to noisy or low-quality proposals. In fact, invalid relation contexts usually indicate underlying scene misunderstanding and ambiguity, which may, on the contrary, reduce the performance in complex scenes. Inspired by recent attention mechanism like Transformer, we propose a novel 3D attention-based relation module (ARM3D). It encompasses object-aware relation reasoning to extract pair-wise relation contexts among qualified proposals and an attention module to distribute attention weights towards different relation contexts. In this way, ARM3D can take full advantage of the useful relation contexts and filter those less relevant or even confusing contexts, which mitigates the ambiguity in detection. We have evaluated the effectiveness of ARM3D by plugging it into several state-of-the-art 3D object detectors and showing more accurate and robust detection results. Extensive experiments show the capability and generalization of ARM3D on 3D object detection. Our source code is available at https://github.com/lanlan96/ARM3D .
Yuqing Lan, Yao Duan, Chenyi Liu, Chenyang Zhu 0002, Yueshan Xiong, Hui Huang 0004, Kai Xu 0004
Comput. Vis. Media1
2021 3DRM: Pair-wise relation module for 3D object detection
Yuqing Lan, Yao Duan, Hui Huang 0004, Kai Xu 0004
Comput. Graph.1
2018 The Rapid Extraction of Suspicious Traffic from Passive DNS
Tianning Zang, Yuqing Lan
ICISSP3
2018 Revealing the densest communities of social networks efficiently through intelligent data space reduction
Yu-Chu Tian, Yuqing Lan, Fenglian Li
Expert Syst. Appl.3
2014 Research on technology of desktop virtualization based on SPICE protocol and its improvement solutions
Yuqing Lan
Frontiers Comput. Sci.1