Yifan Xing

dblp:93/10423 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AuthGuard: Generalizable Deepfake Detection via Language Guidance
abstract
Existing deepfake detection techniques struggle to keep-up with the ever-evolving novel, unseen forgeries methods. This limitation stems from their reliance on statistical artifacts learned during training, which are often tied to specific generation processes that may not be representative of samples from new, unseen deepfake generation methods encountered at test time. We propose that incorporating language guidance can improve deepfake detection generalization by integrating human-like commonsense reasoning – such as recognizing logical inconsistencies and perceptual anomalies – alongside statistical cues. To achieve this, we train an expert deepfake vision encoder by combining discriminative classification with image-text contrastive learning, where the text is generated by generalist MLLMs using few-shot prompting. This allows the encoder to extract both language-describable, commonsense deepfake artifacts and statistical forgery artifacts from pixel-level distributions. To further enhance robustness, we integrate data uncertainty learning into vision-language contrastive learning, mitigating noise in image-text supervision. Our expert vision encoder seamlessly interfaces with an LLM, further enabling more generalized and interpretable deepfake detection while also boosting accuracy. The resulting framework, AuthGuard, achieves state-of-the-art deepfake detection accuracy in both in-distribution and out-of-distribution settings, achieving AUC gains of 6.15% on the DFDC dataset and 16.68% on the DF40 dataset. Additionally, AuthGuard significantly enhances deepfake reasoning, improving performance by 24.69% on the DDVQA dataset.
Guangyu Shen, Tianchen Zhao, Zheng Zhang 0001, Dongsheng An, Zhuowen Tu, Yifan Xing
WACV8
2025 Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
abstract
Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of their face data with service providers. In this work, we introduce a novel method in which the face anti-spoofing model can be adapted by the client itself to a target domain at test time using only a small sample of data while keeping model parameters and training data inaccessible to the client. Specifically, we develop a prototype-based base model and an optimal transport-guided adaptor that enables adaptation in either a lightweight training or training-free fashion, without updating base model’s parameters. Furthermore, we propose geodesic mixup, an optimal transport-based synthesis method that generates augmented training data along the geodesic path between source prototypes and target data distribution. This allows training a lightweight classifier to effectively adapt to target-specific characteristics while retaining essential knowledge learned from the source domain. In cross-domain and cross-attack settings, compared with recent methods, our method achieves average relative improvements of 19.17% in HTER and 8.58% in AUC, respectively.
Zhuowei Li 0002, Tianchen Zhao, Xuanbai Chen, Alessandro Bergamo, Anil K. Jain 0001, Yifan Xing
CVPR10
2025 Model Diagnosis and Correction via Linguistic and Implicit Attribute Editing
abstract
How can we troubleshoot a deep visual model, i.e. understand why it makes certain mistakes and further take action to correct its behavior? We design a ${\mathbf{M}}$ odel ${\mathbf{D}}$ iagnosis and ${\mathbf{C}}$ orrection system (MDC), an automated framework that analyzes the pattern of errors, proposes candidate causes of attributes, conducts hypothesis testing via attribute editing, and ultimately generates counterfactual training samples to improve the performance of the model. Unlike previous methods, in addition to the linguistic attributes, our method also incorporates the analysis for implicit causal attributes, those cannot to be accurately described by natural language. To achieve this, we propose an image editing module capable of leveraging both implicit and linguistic attributes to generate counterfactual images depicting error patterns and further experimentally validate causality relationships. Lastly, we enrich the training set with synthetic samples depicting verified causal attributes and retrain the model, further boosting accuracy and robustness. Extensive experiments on both generalized and specialized domains demonstrate the superiority of MDC in model diagnosis and correction. Specifically, we achieve an average relative improvement of 62.01% in HTER for face security application over state-of-the-art methods.
Xuanbai Chen, Tianchen Zhao, Pietro Perona, Yifan Xing
CVPR7
2025 Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
abstract
This work presents a simple yet effective workflow for automatically scaling instruction-following data to elicit pixel-level grounding capabilities of VLMs under complex instructions. In particular, we address five critical real-world challenges in text-instruction-based grounding: hallucinated references, multi-object scenarios, reasoning, multi-granularity, and part-level references. By leveraging knowledge distillation from a pre-trained teacher model, our approach generates high-quality instruction-response pairs linked to existing pixel-level annotations, minimizing the need for costly human annotation. The resulting dataset, Ground-V, captures rich object localization knowledge and nuanced pixel-level referring expressions. Experiment results show that models trained on Ground-V exhibit substantial improvements across diverse grounding tasks. Specifically, incorporating Ground-V during training directly achieve an average accuracy boost of 4.4% for LISA and a 7.9% for PSALM across six benchmarks on the gIoU metric. It also sets new state-of-the-art results on standard benchmarks such as RefCOCO/+/g. Notably, on gRefCOCO, we achieve an N-Acc of 83.3%, exceeding the previous state-of-the-art by more than 20%.
Yongshuo Zong, Dongsheng An, Linghan Xu, Zhuowen Tu, Yifan Xing, Onkar Dabeer
CVPR8
2025 Salient Concept-Aware Generative Data Augmentation
abstract
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73\% and 6.5\% under conventional and long-tail settings, respectively.
Tianchen Zhao, Xuanbai Chen, Dongsheng An, Zhuowen Tu, Yifan Xing
NeurIPS8
2025 Human perception faithful curve reconstruction based on persistent homology and principal curve
abstract
Reconstructing curves that align with human visual perception from a noisy point cloud presents a significant challenge in the field of curve reconstruction. A specific problem involves reconstructing curves from a noisy point cloud sampled from multiple intersecting curves, ensuring that the reconstructed results align with the Gestalt principles and thus produce curves faithful to human perception. This task involves identifying all potential curves from a point cloud and reconstructing approximating curves, which is critical in applications such as trajectory reconstruction, path planning, and computer vision. In this study, we propose an automatic method that utilizes the topological understanding provided by persistent homology and the local principal curve method to separate and approximate the intersecting closed curves from point clouds, ultimately achieving successful human perception faithful curve reconstruction results using B-spline curves. This technique effectively addresses noisy data clouds and intersections, as demonstrated by experimental results.
Yu Chen 0104, Yifan Xing
Graph. Model.3
2025 Benchmarking neural radiance fields for autonomous robots: An overview
Yuhang Ming 0001, Xingrui Yang 0001, Zheng Chen 0016, Jinglun Feng, Yifan Xing, Guofeng Zhang 0001
Eng. Appl. Artif. Intell.6
2024 Learning for Transductive Threshold Calibration in Open-World Recognition
abstract
In deep metric learning for visual recognition, the calibration of distance thresholds is crucial for achieving desired model performance in the true positive rates (TPR) or true negative rates (TNR). However, calibrating this threshold presents challenges in open-world scenarios, where the test classes can be entirely disjoint from those encountered during training. We define the problem of finding distance thresholds for a trained embedding model to achieve target performance metrics over unseen open-world test classes as open-world threshold calibration. Existing posthoc threshold calibration methods, reliant on inductive inference and requiring a calibration dataset with a similar distance distribution as the test data, often prove ineffective in open-world scenarios. To address this, we introduce OpenGCN, a Graph Neural Network-based transductive threshold calibration method with enhanced adaptability and robustness. OpenGCN learns to predict pairwise connectivity for the unlabeled test instances embedded in a graph to determine its TPR and TNR at various distance thresholds, allowing for transductive inference of the distance thresholds which also incorporates test-time information. Extensive experiments across open-world visual recognition benchmarks validate OpenGCN's superiority over existing posthoc calibration methods for open-world threshold calibration.
Dongsheng An, Tianjun Xiao, Tong He 0002, Qingming Tang, Ying Nian Wu, Joseph Tighe, Yifan Xing
CVPR8
2024 Open-World Dynamic Prompt and Continual Visual Representation Learning
Youngeun Kim, Zhaowei Cai, Yantao Shen 0002, Rahul Duggal, Dripta S. Raychaudhuri, Zhuowen Tu, Yifan Xing, Onkar Dabeer
ECCV (49)9
2024 Threshold-Consistent Margin Loss for Open-World Deep Metric Learning
abstract
Existing losses used in deep metric learning (DML) for image retrieval often lead to highly non-uniform intra-class and inter-class representation structures across test classes and data distributions. When combined with the common practice of using a fixed threshold to declare a match, this gives rise to significant performance variations in terms of false accept rate (FAR) and false reject rate (FRR) across test classes and data distributions. We define this issue in DML as threshold inconsistency. In real-world applications, such inconsistency often complicates the threshold selection process when deploying large-scale image retrieval systems. To measure this inconsistency, we propose a novel variance-based metric called Operating-Point-Inconsistency-Score (OPIS) that quantifies the variance in the operating characteristics across classes. Using the OPIS metric, we find that achieving high accuracy levels in a DML model does not automatically guarantee threshold consistency. In fact, our investigation reveals a Pareto frontier in the high-accuracy regime, where existing methods to improve accuracy often lead to degradation in threshold consistency. To address this trade-off, we introduce the Threshold-Consistent Margin (TCM) loss, a simple yet effective regularization technique that promotes uniformity in representation structures across classes by selectively penalizing hard sample pairs. Large-scale experiments demonstrate TCM's effectiveness in enhancing threshold consistency while preserving accuracy, simplifying the threshold selection process in practical DML settings.
Linghan Xu, Qingming Tang, Ying Nian Wu, Joseph Tighe, Yifan Xing
ICLR7
2024 Object-based SLAM Using Superquadrics
abstract
Visual SLAM uses visual information, typically point features, to localise a camera and, at the same time, map the environment. In recent years, there has been interest in using scene-understanding capabilities to enhance the mapping process and object-level SLAM systems have appeared in response. However, most of the previous work is limited to prestored object models or pre-trained networks to represent the objects, which limits working scenarios or uses representations with limited scope, such as cubes or quadrics. To address this, we propose to use superquadrics as the object representation and, in this paper, present a proof of principle SLAM system in which object-based mapping is fully integrated with camera tracking via keyframe optimisation. The system was tested on simulated and real datasets, and the results show that the system can achieve lightweight and comparatively good object representation whilst also giving good camera trajectories estimates under certain scenarios.
Yifan Xing, Noe Samano, Wen Fan 0001, Andrew Calway
IROS1
2023 Tac-VGNN: A Voronoi Graph Neural Network for Pose-Based Tactile Servoing
abstract
Tactile pose estimation and tactile servoing are fundamental capabilities of robot touch. Reliable and precise pose estimation can be provided by applying deep learning models to high-resolution optical tactile sensors. Given the recent successes of Graph Neural Network (GNN) and the effectiveness of Voronoi features, we developed a Tactile Voronoi Graph Neural Network (Tac-VGNN) to achieve reliable pose-based tactile servoing relying on a biomimetic optical tactile sensor (TacTip). The GNN is well suited to modeling the distribution relationship between shear motions of the tactile markers, while the Voronoi diagram supplements this with area-based tactile features related to contact depth. The experiment results showed that the Tac-VGNN model can help enhance data interpretability during graph generation and model training efficiency significantly than CNN-based methods. It also improved pose estimation accuracy along vertical depth by 28.57% over vanilla GNN without Voronoi features and achieved better performance on the real surface following tasks with smoother robot control trajectories. For more project details, please view our website: https://sites.google.com/view/tac-vgnn/home
Wen Fan 0001, Max Yang, Yifan Xing, Nathan F. Lepora, Dandan Zhang 0001
ICRA3
2022 Knowledge Graph Enhanced Web API Recommendation via Neighbor Information Propagation for Multi-service Application Development
Zhen Chen 0007, Yifan Xing, Dianlong You
CollaborateCom (1)5
2022 PSS: Progressive Sample Selection for Open-World Visual Representation Learning
Yifan Xing, Tianjun Xiao, Tong He 0002, Zheng Zhang 0023, Hao Zhou 0045, Joseph Tighe
ECCV (31)3
2021 Learning Hierarchical Graph Neural Networks for Image Clustering
abstract
We propose a hierarchical graph neural network (GNN) model that learns how to cluster a set of images into an unknown number of identities using a training set of images annotated with labels belonging to a disjoint set of identities. Our hierarchical GNN uses a novel approach to merge connected components predicted at each level of the hierarchy to form a new graph at the next level. Unlike fully unsupervised hierarchical clustering, the choice of grouping and complexity criteria stems naturally from supervision in the training set. The resulting method, Hi-LANDER, achieves an average of 49% improvement in F-score and 7% increase in Normalized Mutual Information (NMI) relative to current GNN-based clustering algorithms. Additionally, state-of-the-art GNN-based methods rely on separate models to predict linkage probabilities and node densities as intermediate steps of the clustering process. In contrast, our unified framework achieves a three-fold decrease in computational cost. Our training and inference code are released1.
Yifan Xing, Tong He 0002, Tianjun Xiao, Yuanjun Xiong, Wei Xia 0009, David P. Wipf, Zheng Zhang 0001, Stefano Soatto
ICCV1
2020 A Real Time Simulation Optimization Framework for Vessel Collision Avoidance and the Case of Singapore Strait
abstract
Safety is a primary concern for the various transport means. For sea transport, this includes various aspects like human safety at sea and at port, and also environmental safety and sustainability. In heavy-traffic regions where the waters are congested and vessels sail very closely together, ensuring these safety needs can be challenging. In this paper, we leverage on the rich information transmitted through the automatic identification system (AIS) and propose, for the first time, an integrated simulation-optimization approach for real time collision avoidance. This enables capturing of stochastic dynamic behavior of vessels for better prediction and fast trajectory optimization for application in real time. Specifically, a realistic agent-based model is developed based on behavioral learning in a real-environment, and incorporated into a fast collision avoidance optimization model in real time to provide robust collision avoidance that is able to account for future stochastic consequences of the actions taken. To achieve this, we develop: 1) a vessel pattern recognition method that mines the rich AIS data to produce realistic trajectory models; 2) an agent-based simulation model to enhance future trajectory prediction; and 3) a fast surrogate-based sampling technique to generate collision avoidance maneuvers for vessel captains in real time. To illustrate the feasibility of the approach, we use the case of the Singapore strait, one of the busiest straits in the world.
Giulia Pedrielli, Yifan Xing, Jia Hao Peh, Kim Wee Koh, Szu Hui Ng
IEEE Trans. Intell. Transp. Syst.2
2019 A Self-Supervised Bootstrap Method for Single-Image 3D Face Reconstruction
abstract
State-of-the-art methods for 3D reconstruction of faces from a single image require 2D-3D pairs of ground-truth data for supervision. Such data is costly to acquire, and most datasets available in the literature are restricted to pairs for which the input 2D images depict faces in a near fronto-parallel pose. Therefore, many data-driven methods for single-image 3D facial reconstruction perform poorly on profile and near-profile faces. We propose a method to improve the performance of single-image 3D facial reconstruction networks by utilizing the network to synthesize its own training data for fine-tuning, comprising: (i) single-image 3D reconstruction of faces in near-frontal images without ground-truth 3D shape; (ii) application of a rigid-body transformation to the reconstructed face model; (iii) rendering of the face model from new viewpoints; and (iv) use of the rendered image and corresponding 3D reconstruction as additional data for supervised fine-tuning. The new 2D-3D pairs thus produced have the same high-quality observed for near fronto-parallel reconstructions, thereby nudging the network towards more uniform performance as a function of the viewing angle of input faces. Application of the proposed technique to the fine-tuning of a state-of-the-art single-image 3D-reconstruction network for faces demonstrates the usefulness of the method, with particularly significant gains for profile or near-profile views.
Yifan Xing, Rahul Tewari, Paulo R. S. Mendonça
WACV1