Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Irving Fang

dblp:284/8283 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 81% Image recognition and object detection · 14% Robot manipulation · 5%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d reconstruction › object reconstruction
3d reassembly
0.912025
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures · ICCV 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025
Computer vision › 3D vision › geometric estimation
3d registration
0.912025
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures · ICCV 2025
Computer vision › 3D vision › 3d reconstruction › multi-view reconstruction
sparse-view reconstruction
0.912025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025
Computer vision › 3D vision
egocentric vision
0.812024
EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot Interaction · ICRA 2024
Computer vision › Image recognition and object detection
image classification
0.812024
LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images · CVPR 2024
Human-robot interaction
intent prediction
0.812024
EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot Interaction · ICRA 2024
Robotics › Robot manipulation
tactile sensing
0.312025
Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction · ICRA 2025
Computational social science and digital humanities
archaeology
0.212024
LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images · CVPR 2024

Methods — techniques the papers use, named apart from their topics

pretrained vision models · 1.5large pre-trained model · 1.5human prior knowledge · 1.5few-shot learning · 1.5hierarchical optimization · 0.9fracture-aware pretraining · 0.9foundation model priors · 0.9flow matching · 0.93d gaussian splatting · 0.9
YearPublicationVenuePosition
2025 GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
abstract
3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models trained on synthetic datasets can generalize to real-world fractures where breakage patterns are more complex. To bridge this gap, we propose GARF, a generalizable 3D reassembly framework for real-world fractures. GARF leverages fracture-aware pretraining to learn fracture features from individual fragments, with flow matching enabling precise 6-DoF alignments. At inference time, we introduce one-step preassembly, improving robustness to unseen objects and varying numbers of fractures. In collaboration with archaeologists, paleoanthropologists, and ornithologists, we curate Fractura, a diverse dataset for vision and learning communities, featuring real-world fracture types across ceramics, bones, eggshells, and lithics. Comprehensive experiments have shown our approach consistently outperforms state-of-the-art methods on both synthetic and real-world datasets, achieving 82.87\% lower rotation error and 25.15\% higher part accuracy. This sheds light on training on synthetic data to advance real-world 3D puzzle solving, demonstrating its strong generalization across unseen object shapes and diverse fracture types. GARF's code, data and demo are available at https://ai4ce.github.io/GARF/.
Sihang Li 0001, Grace Chen, Siqi Tan, Irving Fang, Kristof Zyskowski, Shannon P. McPherron, Radu Iovita, Chen Feng 0002
ICCV7
2025 Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
abstract
Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observations from vision and tactile sensors. FusionSense addresses three key challenges: (i) How can robots efficiently acquire robust global shape information about the surrounding scene and objects? (ii) How can robots strategically select touch points on the object using geometric and commonsense priors? (iii) How can partial observations such as tactile signals improve the overall representation of the object? Our framework employs 3D Gaussian Splatting as a core representation and incorporates a hierarchical optimization strategy involving global structure construction, object visual hull pruning and local geometric constraints. This advancement results in fast and robust perception in environments with traditionally challenging objects that are transparent, reflective, or dark, enabling more downstream manipulation or navigation tasks. Experiments on real-world data suggest that our framework outperforms previously state-of-the-art sparse-view methods. All code and data are open-sourced on the project website.
Irving Fang, Kairui Shi, Xujin He, Siqi Tan, Hanwen Zhao, Hung-Jui Huang, Wenzhen Yuan 0001, Chen Feng 0002
ICRA1
2025 VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
abstract
Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interpret long-horizon human demonstration videos to generate a sequence of robot task plans in natural language. To achieve this, we propose SeeDo, an agent that integrates keyframe selection module, visual prompting module, and a VLM interpreter into a pipeline that enables the VLM to "see" human demonstrations and generate step-by-step plans for robots to "do" them. To evaluate, we curate a benchmark of long-horizon human demonstration videos of pick-and-place tasks in three diverse categories and designed comprehensive evaluation metrics. The experiments demonstrate SeeDo’s superior performance in generating subtask planning in natural language from long-horizon human demo videos. Experiments show SeeDo outperforms state-of-the-art video VLMs in generating subtask plans. By further integrating SeeDo with low-level action primitive functions and language model programs, we validated SeeDo in both simulated and real-world deployments. The code, demos, prompts and data can be found at ai4ce.github.io/SeeDo.
Beichen Wang, Juexiao Zhang, Shuwen Dong, Irving Fang, Chen Feng 0002
IROS4
2024 LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images
abstract
Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material, which is critical for understanding archaeological artifacts, material interactions, tool functionalities, and dental records. However, this challenging task goes beyond the well-studied image classification problem for common objects. It is affected by many confounders owing to the complex wear mechanism and microscopic imaging, which makes it difficult even for human experts to identify the worked material successfully. In this paper, we investigate the following three questions on this unique vision task for the first time:(i) How well can state-of-the-art pre-trained models (like DINOv2) generalize to the rarely seen domain? (ii) How can few-shot learning be exploited for scarce microscopic images? (iii) How do the ambiguous magnification and sensing modality influence the classification accuracy? To study these, we collaborated with archaeologists and built the first open-source and the largest LUWA dataset containing 23,130 microscopic images with different magnifications and sensing modalities. Extensive experiments show that existing pretrained models notably outperform human experts but still leave a large gap for improvements. Most importantly, the LUWA dataset provides an underexplored opportunity for vision and learning communities and complements existing image classification problems on common objects.
Irving Fang, Akshat Kaushik, Alice Rodriguez, Hanwen Zhao, Juexiao Zhang, Zhuo Zheng, Radu Iovita, Chen Feng 0002
CVPR2
2024 EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot Interaction
abstract
A robot’s ability to anticipate the 3D action target location of a hand’s movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic action classification or 2D target region prediction, we argue that predicting the action target’s 3D coordinate could pave the way for more versatile downstream robotics tasks, especially given the increasing prevalence of headset devices. This study expands EgoPAT3D, the sole dataset dedicated to egocentric 3D action target prediction. We augment both its size and diversity, enhancing its potential for generalization. Moreover, we substantially enhance the baseline algorithm by introducing a large pre-trained model and human prior knowledge. Remarkably, our novel algorithm can now achieve superior prediction outcomes using solely RGB images, eliminating the previous need for 3D point clouds and IMU input. Furthermore, we deploy our enhanced baseline algorithm on a real-world robotic platform to illustrate its practical utility in straightforward HRI tasks. The demonstrations showcase the real-world applicability of our advancements and may inspire more HRI use cases involving egocentric vision. All code and data are open-sourced and can be found on the project website.
Irving Fang, Jianghan Zhang, Xibo He, Weibo Gao, Yiming Li 0003, Chen Feng 0002
ICRA1