Siqi Tan

dblp:330/2416 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 94% Robot manipulation · 6%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d reconstruction › object reconstruction
3d reassembly
0.912025
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures · ICCV 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025
Computer vision › 3D vision › geometric estimation
3d registration
0.912025
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures · ICCV 2025
Computer vision › 3D vision › 3d reconstruction › multi-view reconstruction
sparse-view reconstruction
0.912025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025
Computer vision › 3D vision
visual localization
0.912025
Adversarial Exploitation of Data Diversity Improves Visual Localization · ICCV 2025
Robotics › Robot manipulation
tactile sensing
0.312025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025

Methods — techniques the papers use, named apart from their topics

hierarchical optimization · 0.9fracture-aware pretraining · 0.9foundation model priors · 0.9flow matching · 0.9adversarial exploitation · 0.93d gaussian splatting · 0.9
YearPublicationVenuePosition
2025 GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
abstract
3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models trained on synthetic datasets can generalize to real-world fractures where breakage patterns are more complex. To bridge this gap, we propose GARF, a generalizable 3D reassembly framework for real-world fractures. GARF leverages fracture-aware pretraining to learn fracture features from individual fragments, with flow matching enabling precise 6-DoF alignments. At inference time, we introduce one-step preassembly, improving robustness to unseen objects and varying numbers of fractures. In collaboration with archaeologists, paleoanthropologists, and ornithologists, we curate Fractura, a diverse dataset for vision and learning communities, featuring real-world fracture types across ceramics, bones, eggshells, and lithics. Comprehensive experiments have shown our approach consistently outperforms state-of-the-art methods on both synthetic and real-world datasets, achieving 82.87\% lower rotation error and 25.15\% higher part accuracy. This sheds light on training on synthetic data to advance real-world 3D puzzle solving, demonstrating its strong generalization across unseen object shapes and diverse fracture types. GARF's code, data and demo are available at https://ai4ce.github.io/GARF/.
Sihang Li 0001, Grace Chen, Siqi Tan, Irving Fang, Kristof Zyskowski, Shannon P. McPherron, Radu Iovita, Chen Feng 0002
ICCV5
2025 Adversarial Exploitation of Data Diversity Improves Visual Localization
Sihang Li 0001, Siqi Tan, Bowen Chang, Chen Feng 0002, Yiming Li 0003
ICCV2
2025 Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
abstract
Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observations from vision and tactile sensors. FusionSense addresses three key challenges: (i) How can robots efficiently acquire robust global shape information about the surrounding scene and objects? (ii) How can robots strategically select touch points on the object using geometric and commonsense priors? (iii) How can partial observations such as tactile signals improve the overall representation of the object? Our framework employs 3D Gaussian Splatting as a core representation and incorporates a hierarchical optimization strategy involving global structure construction, object visual hull pruning and local geometric constraints. This advancement results in fast and robust perception in environments with traditionally challenging objects that are transparent, reflective, or dark, enabling more downstream manipulation or navigation tasks. Experiments on real-world data suggest that our framework outperforms previously state-of-the-art sparse-view methods. All code and data are open-sourced on the project website.
Irving Fang, Kairui Shi, Xujin He, Siqi Tan, Hanwen Zhao, Hung-Jui Huang, Wenzhen Yuan 0001, Chen Feng 0002
ICRA4
2023 OA-Bug: An Olfactory-Auditory Augmented Bug Algorithm for Swarm Robots in a Denied Environment
abstract
Searching in a denied environment is challenging for swarm robots as no assistance from GNSS, mapping, data sharing, and central processing is allowed. However, using olfactory and auditory signals to cooperate like animals could be an important way to improve the collaboration of swarm robots. In this paper, an Olfactory-Auditory augmented Bug algorithm (OA-Bug) is proposed for a swarm of autonomous robots to explore a denied environment. A simulation environment is built to measure the performance of OA-Bug. The coverage of the search task can reach 96.93% using OA-Bug, which is significantly improved compared with a similar algorithm, SGBA [1]. Furthermore, experiments are conducted on real swarm robots to prove the validity of OA-Bug. Results show that OA-Bug can improve the performance of swarm robots in a denied environment. Video: https://youtu.be/vj9cRiSmgeM.
Siqi Tan, Ruitao Jing, Mufan Zhao, Quan Quan
IROS1