Yifang Xu

dblp:17/11453 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
abstract
The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some recent studies introduce a zero-shot setting to avoid fine-tuning, they overlook inherent language bias in the query, leading to erroneous localization. To tackle the aforementioned challenges, this paper proposes Moment-GPT, a tuning-free pipeline for zero-shot VMR utilizing frozen MLLMs. Specifically, we first employ LLaMA-3 to correct and rephrase the query to mitigate language bias. Subsequently, we design a span generator combined with MiniGPT-v2 to produce candidate spans adaptively. Finally, to leverage the video comprehension capabilities of MLLMs, we apply Video-ChatGPT and span scorer to select the most appropriate spans. Our proposed method substantially outperforms the state-of-the-art MLLM-based and zero-shot models on several public datasets, including QVHighlights, ActivityNet-Captions, and Charades-STA.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 0069, Wenxin Liang, Yang Li 0063, Sidan Du
AAAI1
2025 HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
abstract
Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works.
Yifang Xu, Benxiang Zhai, Yunzhuo Sun, Ming Li 0069, Yang Li 0063, Sidan Du
CVPR1
2025 FaceSnap: Enhanced ID-Fidelity Network for Tuning-Free Portrait Customization
Benxiang Zhai, Yifang Xu, Guofeng Zhang 0026, Yang Li 0063, Sidan Du
ICANN (2)2
2024 Multi-Modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
abstract
Given a video and a linguistic query, video moment retrieval and highlight detection (MR&HD) aim to locate all the relevant spans, while simultaneously predicting saliency scores. Most existing methods utilize RGB images as input, overlooking the inherent multi-modal visual signals like optical flow and depth. In this paper, we propose a Multi-modal Fusion and Query Refinement Network (MRNet) to learn complementary information from multi-modal cues. Specifically, we design a multi-modal fusion module to dynamically combine RGB, optical flow, and depth map. Furthermore, to simulate human understanding of sentences, we introduce a query refinement module that merges text at different granularities, containing word-, phrase-, and sentence-wise levels. Comprehensive experiments on QVHighlights and Charades datasets indicate that MRNet outperforms current SOTA methods, achieving notable improvements in MR-mAP@Avg (+3.41) and HD-HIT@1 (+3.46) on QVHighlights.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Zien Xie, Youyao Jia, Sidan Du
ICME1
2024 MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer
abstract
With the increasing demand for video understanding, video moment and highlight detection (MHD) has emerged as a critical research topic. MHD aims to localize all moments and predict clip-wise saliency scores simultaneously. Despite progress made by existing DETR-based methods, we observe that these methods coarsely fuse features from different modalities, which weakens the temporal intra-modal context and results in insufficient cross-modal interaction. To address this issue, we propose MH-DETR (Moment and Highlight DEtection TRansformer) tailored for MHD. Specifically, we introduce a simple yet efficient pooling operator within the uni-modal encoder to capture global intra-modal context. Moreover, to obtain temporally aligned cross-modal features, we design a plug-and-play cross-modal interaction module between the encoder and decoder, seamlessly integrating visual and textual features. Comprehensive experiments on QVHighlights, Charades-STA, Activity-Net, and TVSum datasets show that MH-DETR outperforms existing state-of-the-art methods, demonstrating its effectiveness and superiority. Our code is available at https://github.com/YoucanBaby/MH-DETR.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Youyao Jia, Sidan Du
IJCNN1
2024 GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
abstract
Moment retrieval (MR) and highlight detection (HD) aim to identify relevant moments and highlights in video from corresponding natural language query. Large language models (LLMs) have demonstrated proficiency in various computer vision tasks. However, existing methods for MR&HD have not yet been integrated with LLMs. In this letter, we propose a novel two-stage model that takes the output of LLMs as the input to the second-stage transformer encoder-decoder. First, MiniGPT-4 is employed to generate the detailed description of the video frame and rewrite the query statement, fed into the encoder as new features. Then, semantic similarity is computed between the generated description and the rewritten queries. Finally, continuous high-similarity video frames are converted into span anchors, serving as prior position information for the decoder. Experiments demonstrate that our approach achieves a state-of-the-art result, and by using only span anchors and similarity scores as outputs, positioning accuracy outperforms traditional methods, like Moment-DETR.
Yunzhuo Sun, Yifang Xu, Zien Xie, Yukun Shu, Sidan Du
IEEE Signal Process. Lett.2
2023 AGAA: An Android GUI Accessibility Adapter for Low Vision Users
abstract
The graphical user interface (GUI) is crucial for users to interact with mobile devices. However, accessibility issues in the GUI, such as undersized text and redundant information, lead to understanding and operating obstacles for billions of low vision users in our society. To alleviate this situation, academia and industry have proposed various accessibility-related methods. Still, their over-dependence on specific detection rules and their inability to automatically repair the GUI source code limit them in helping developers resolve these issues. In this paper, we propose a novel method, named AGAA, for capturing and repairing undersized text and redundant information issues in the GUI. The evaluation on 12 real-world apps and the user study on 36 low vision users demonstrate that AGAA is effective in resolving these issues and is useful in improving the mobile device experience for low vision users, respectively.
Yifang Xu, Zhuopeng Li, Huaxiao Liu, Yuzhou Liu 0001
COMPSAC1
2023 Query-Guided Refinement and Dynamic Spans Network for Video Highlight Detection and Temporal Grounding in Online Information Systems
abstract
With the surge in online video content, finding highlights and key video segments have garnered widespread attention. Given a textual query, video highlight detection (HD) and temporal grounding (TG) aim to predict frame-wise saliency scores from a video while concurrently locating all relevant spans. Despite recent progress in DETR-based works, these methods crudely fuse different inputs in the encoder, which limits effective cross-modal interaction. To solve this challenge, the authors design QD-Net (query-guided refinement and dynamic spans network) tailored for HD&TG. Specifically, they propose a query-guided refinement module to decouple the feature encoding from the interaction process. Furthermore, they present a dynamic span decoder that leverages learnable 2D spans as decoder queries, which accelerates training convergence for TG. On QVHighlights dataset, the proposed QD-Net achieves 61.87 HD-HIT@1 and 61.88 [email protected], yielding a significant improvement of +1.88 and +8.05, respectively, compared to the state-of-the-art method.
Yifang Xu, Yunzhuo Sun, Zien Xie, Benxiang Zhai, Youyao Jia, Sidan Du
Int. J. Semantic Web Inf. Syst.1
2023 Improved GEDI Canopy Height Extraction Based on a Simulated Ground Echo in Topographically Undulating Areas
abstract
The Global Ecosystem Dynamics Investigation (GEDI) instrument, which represents a new generation of spaceborne full-waveform LiDAR systems, is also the first spaceborne LiDAR system specifically designed to monitor the vertical structure of vegetation. In the four years of operation to date, the GEDI instrument has provided unprecedented observations for global forest height and forest above-ground biomass estimation studies. However, it remains a daunting challenge to obtain accurate canopy height estimates based on GEDI observations in topographically undulating areas. Studies on how to make full use of the actual topographic conditions within each GEDI footprint, to aid in canopy height extraction, are still lacking. In this paper, we propose a method based on the Shuttle Radar Topography Mission (SRTM) digital elevation model (DEM)-simulated ground echo for assisting in ground waveform identification and canopy top elevation correction, which in turn improves the accuracy of the maximum canopy height extraction in topographically undulating areas. The validation results show that the accuracy of the ground elevation, canopy top elevation, and maximum canopy height obtained by the proposed method is improved by 9.2%, 15%, and 14%, respectively, compared with the GEDI L2A product. Since the only auxiliary data used in this approach are the publicly available SRTM DEM product data, the proposed method has the potential to be applied to large regions or even globally. This paper will promote the application of GEDI observations in monitoring the vertical structure of vegetation in areas of topographic relief.
Hailong Tang, Huabing Huang, Peng Qin 0004, Yifang Xu
IEEE Trans. Geosci. Remote. Sens.5
2021 Pyramid Feature Attention Network for Monocular Depth Prediction
abstract
Deep convolutional neural networks (DCNNs) have achieved great success in monocular depth estimation (MDE). However, few existing works take the contributions for MDE of different levels feature maps into account, leading to inaccurate spatial layout, ambiguous boundaries and discontinuous object surface in the prediction. To better tackle these problems, we propose a Pyramid Feature Attention Network (PFANet) to improve the high-level context features and lowlevel spatial features. In the proposed PFANet, we design a Dual-scale Channel Attention Module (DCAM) to employ channel attention in different scales, which aggregate global context and local information from the high-level feature maps. To exploit the spatial relationship of visual features, we design a Spatial Pyramid Attention Module (SPAM) which can guide the network attention to multi-scale detailed information in the low-level feature maps. Finally, we introduce scale-invariant gradient loss to increase the penalty on errors in depth-wise discontinuous regions. Experimental results show that our method outperforms state-of-the-art methods on the KITTI dataset.
Yifang Xu, Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du
ICME1
2021 Generative Adversarial Domain Generalization via Cross-Task Feature Attention Learning for Prostate Segmentation
Yifang Xu, Ye Luo 0004, Enbei Zhu
ICONIP (2)1
2018 Quasi-Homography Warps in Image Stitching
abstract
The naturalness of warps is gaining extensive attention in image stitching. Recent warps, such as SPHP and AANAP, use global similarity warps to mitigate projective distortion (which enlarges regions); however, they necessarily bring in perspective distortion (which generates inconsistencies). In this paper, we propose a novel quasi-homography warp, which effectively balances the perspective distortion against the projective distortion in the non-overlapping region to create a more natural-looking panorama. Our approach formulates the warp as the solution of a bivariate system, where perspective distortion and projective distortion are characterized as slope preservation and scale linearization, respectively. Because our proposed warp only relies on a global homography, it is thus totally parameter free. A comprehensive experiment shows that a quasi-homography warp outperforms some state-of-the-art warps in urban scenes, including homography, AutoStitch and SPHP. A user study demonstrates that it wins most users' favor, compared to homography and SPHP.
Nan Li 0013, Yifang Xu, Chao Wang 0020
IEEE Trans. Multim.2
2004 Theoretical prediction of dynamic range and intensity discrimination for electrical noise-modulated pulse-train stimuli
abstract
We investigate dynamic range and intensity discrimination for electrical noise-modulated pulse-train stimuli using a stochastic auditory nerve model (Bruce, I.C. et al., IEEE Trans. Biomed Eng., vol.46, no.6, p.617-29, 1999). Based on a hypothesized monotonic relationship between loudness and the number of spikes, the theoretical prediction of the most uncomfortable level was determined by comparing spike counts to a fixed threshold (Bruce et al., IEEE Trans. Biomed Eng., vol.46, no.6, p.1393-1404, 1999). However, no specific rule for determining this fixed number has previously been suggested. We determine the most uncomfortable level based on the excitation pattern of the basilar membrane in a normal ear. The number of fibers corresponding to the portion of the basilar membrane driven at an uncomfortable stimulus level in a normal ear is related to the most uncomfortable spiking number. The intensity discrimination limens are predicted using signal detection theory via the probability mass function (PMF) of the neural response and via experimental simulations. The results show that the uncomfortable level for a pulse-train stimulus increases slightly as noise level increases. Combining this with our previous threshold predictions (Xu and Collins, IEEE Trans. Biomed. Eng.), we hypothesize that the dynamic range for noise-modulated pulse-train stimuli increases with additive noise. However, since our predictions indicate that intensity discrimination under noise degrades, the overall intensity coding performance does not improve significantly.
Yifang Xu, Leslie M. Collins
ICASSP (4)1
2002 Threshold prediction for noise-modulated electrical stimuli using a stochastic auditory nerve model: Implications for cochlear implants
abstract
The effect of a low level additive noise process on the input and output characteristics and threshold behavior of auditory nerves (ANs) is studied by means of a stochastic computational model. This paper derives the stochastic properties of the model input and output for adaptive threshold procedures. A closed form solution for the input, or amplitude, probability distribution is obtained via Markov models. The output statistics are derived by integrating over the noise-free probability mass function (PMF). All theoretical PMFs are verified by simulations. Theoretical threshold predictions as a function of noise level are made based on these PMFs and the results indicate that threshold is adversely affected by the presence of low levels of noise.
Yifang Xu, Leslie M. Collins
ICASSP1