Sifeng He

dblp:210/9816 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Acoustic Reverse Nonlinearity Against Speech Forgery in Real-Time Voice Applications
Ming Gao 0023, Lingfeng Zhang 0004, Yike Chen, Sifeng He, Feng Qian 0006, Lei Yang 0061, Fu Xiao 0001, Jinsong Han
INFOCOM4
2025 Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
abstract
Despite Contrastive Language–Image Pre-training (CLIP)'s remarkable capability to retrieve content across modalities, a substantial modality gap persists in its feature space. Intriguingly, we discover that off-the-shelf MLLMs (Multimodal Large Language Models) demonstrate powerful inherent modality alignment properties. While recent MLLM-based retrievers with unified architectures partially mitigate this gap, their reliance on coarse modality alignment mechanisms fundamentally limits their potential. In this work, We introduce MAPLE (Modality-Aligned Preference Learning for Embeddings), a novel framework that leverages the fine-grained alignment priors inherent in MLLM to guide cross-modal representation learning. MAPLE formulates the learning process as reinforcement learning with two key components: (1) Automatic preference data construction using off-the-shelf MLLM, and (2) a new Relative Preference Alignment (RPA) loss, which adapts Direct Preference Optimization (DPO) to the embedding learning setting. Experimental results show that our preference-guided alignment achieves substantial gains in fine-grained cross-modal retrieval, underscoring its effectiveness in handling nuanced semantic distinctions.
Rongbo Luan, Sifeng He
NeurIPS5
2024 Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual Retrieval
abstract
Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this paper we propose a multi-teacher distillation framework Whiten-MTD, which is able to transfer knowledge from off-the-shelf pre-trained retrieval models to a lightweight student model for efficient visual retrieval. Furthermore, we discover that the similarities obtained by different retrieval models are diversified and incommensurable, which makes it challenging to jointly distill knowledge from multiple models. Therefore, we propose to whiten the output of teacher models before fusion, which enables effective multi-teacher distillation for retrieval models. Whiten-MTD is conceptually simple and practically effective. Extensive experiments on two landmark image retrieval datasets and one video retrieval dataset demonstrate the effectiveness of our proposed method, and its good balance of retrieval performance and efficiency. Our source code is released at https://github.com/Maryeon/whiten_mtd.
Zhe Ma 0002, Jianfeng Dong, Shouling Ji, Zhenguang Liu, Xuhong Zhang 0002, Zonghui Wang, Sifeng He, Feng Qian 0006, Lei Yang 0061
AAAI7
2023 TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision
abstract
Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, and then detect and refine the boundaries of copied segments on similarity matrix under temporal constraints. In this paper, we propose TransVCL: an attention-enhanced video copy localization network, which is optimized directly from initial frame-level features and trained end-to-end with three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for similarity matrix generation, and a temporal alignment module for copied segments localization. In contrast to previous methods demanding the handcrafted similarity matrix, TransVCL incorporates long-range temporal information between feature sequence pair using self- and cross- attention layers. With the joint design and optimization of three components, the similarity matrix can be learned to present more discriminative copied patterns, leading to significant improvements over previous methods on segment-level labeled datasets (VCSL and VCDB). Besides the state-of-the-art performance in fully supervised setting, the attention architecture facilitates TransVCL to further exploit unlabeled or simply video-level labeled data. Additional experiments of supplementing video-level labeled datasets including SVD and FIVR reveal the high flexibility of TransVCL from full supervision to semi-supervision (with or without video-level annotation). Code is publicly available at https://github.com/transvcl/TransVCL.
Sifeng He, Minlong Lu, Chen Jiang 0006, Feng Qian 0006, Lei Yang 0061
AAAI1
2023 Boundary-aware Backward-Compatible Representation via Adversarial Learning in Image Retrieval
abstract
Image retrieval plays an important role in the Internet world. Usually, the core parts of mainstream visual retrieval systems include an online service of the embedding model and a large-scale vector database. For traditional model upgrades, the old model will not be replaced by the new one until the embeddings of all the images in the database are re-computed by the new model, which takes days or weeks for a large amount of data. Recently, backward-compatible training (BCT) enables the new model to be immediately deployed online by making the new embeddings directly comparable to the old ones. For BCT, improving the compatibility of two models with less negative impact on retrieval performance is the key challenge. In this paper, we introduce AdvBCT, an Adversarial Backward-Compatible Training method with an elastic boundary constraint that takes both compatibility and discrimination into consideration. We first employ adversarial learning to minimize the distribution disparity between embeddings of the new model and the old model. Meanwhile, we add an elastic boundary constraint during training to improve compatibility and discrimination efficiently. Extensive experiments on GLDv2, Revisited Oxford (ROxford), and Revisited Paris (RParis) demonstrate that our method outperforms other BCT methods on both compatibility and discrimination. The implementation of AdvBCT will be publicly available at https://github.com/Ashespt/AdvBCT.
Tan Pan, Furong Xu, Sifeng He, Chen Jiang 0006, Qingpei Guo, Feng Qian 0006, Lei Yang 0061
CVPR4
2023 Video Infringement Detection via Feature Disentanglement and Mutual Information Maximization
abstract
The self-media era provides us tremendous high quality videos. Unfortunately, frequent video copyright infringements are now seriously damaging the interests and enthusiasm of video creators. Identifying infringing videos is therefore a compelling task. Current state-of-the-art methods tend to simply feed high-dimensional mixed video features into deep neural networks and count on the networks to extract useful representations. Despite its simplicity, this paradigm heavily relies on the original entangled features and lacks constraints guaranteeing that useful task-relevant semantics are extracted from the features.
Zhenguang Liu, Xinyang Yu, Ruili Wang 0001, Zhe Ma 0002, Jianfeng Dong, Sifeng He, Feng Qian 0006, Roger Zimmermann, Lei Yang 0061
ACM Multimedia7
2023 Web Photo Source Identification based on Neural Enhanced Camera Fingerprint
abstract
With the growing popularity of smartphone photography in recent years, web photos play an increasingly important role in all walks of life. Source camera identification of web photos aims to establish a reliable linkage from the captured images to their source cameras, and has a broad range of applications, such as image copyright protection, user authentication, investigated evidence verification, etc. This paper presents an innovative and practical source identification framework that employs neural-network enhanced sensor pattern noise to trace back web photos efficiently while ensuring security. Our proposed framework consists of three main stages: initial device fingerprint registration, fingerprint extraction and cryptographic connection establishment while taking photos, and connection verification between photos and source devices. By incorporating metric learning and frequency consistency into the deep network design, our proposed fingerprint extraction algorithm achieves state-of-the-art performance on modern smartphone photos for reliable source identification. Meanwhile, we also propose several optimization sub-modules to prevent fingerprint leakage and improve accuracy and efficiency. Finally for practical system design, two cryptographic schemes are introduced to reliably identify the correlation between registered fingerprint and verified photo fingerprint, i.e. fuzzy extractor and zero-knowledge proof (ZKP). The codes for fingerprint extraction network and benchmark dataset with modern smartphone cameras photos are all publicly available at https://github.com/PhotoNecf/PhotoNecf 1.
Feng Qian 0006, Sifeng He, Honghao Huang, Huanyu Ma, Lei Yang 0061
WWW2
2022 A Large-scale Comprehensive Dataset and Copy-overlap Aware Evaluation Protocol for Segment-level Video Copy Detection
abstract
In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level labelled data, with 160k realistic video copy pairs containing more than 280k localized copied segment pairs, but also covers a variety of video categories and a wide range of video duration. All the copied segments inside each collected video pair are manually extracted and accompanied by precisely annotated starting and ending timestamps. Alongside the dataset, we also propose a novel evaluation protocol that better measures the prediction accuracy of copy overlapping segments between a video pair and shows improved adaptability in different scenarios. By benchmarking several baseline and state-of-the-art segment-level video copy detection methods with the proposed dataset and evaluation metric, we provide a comprehensive analysis that uncovers the strengths and weaknesses of current approaches, hoping to open up promising directions for future works. The VCSL dataset, metric and benchmark codes are all publicly available at https://github.com/alipay/vCSL.
Sifeng He, Chen Jiang 0006, Gang Liang, Tan Pan, Qing Wang 0068, Furong Xu, Jingxiong Liu, Kaiming Huang, Feng Qian 0006, Lei Yang 0061
CVPR1
2021 Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
abstract
With the explosive growth of web videos in recent years, large-scale Content-Based Video Retrieval (CBVR) becomes increasingly essential in video filtering, recommendation, and copyright protection. Segment-level CBVR (S-CBVR) locates the start and end time of similar segments in finer granularity, which is beneficial for user browsing efficiency and infringement detection especially in long video scenarios. The challenge of S-CBVR task is how to achieve high temporal alignment accuracy with efficient computation and low storage consumption. In this paper, we propose a Segment Similarity and Alignment Network (SSAN) in dealing with the challenge which is firstly trained end-to-end in S-CBVR. SSAN is based on two newly proposed modules in video retrieval: (1) An efficient Self-supervised Keyframe Extraction (SKE) module to reduce redundant frame features, (2) A robust Similarity Pattern Detection (SPD) module for temporal alignment. In comparison with uniform frame extraction, SKE not only saves feature storage and search time, but also introduces comparable accuracy and limited extra computation time. In terms of temporal alignment, SPD localizes similar segments with higher accuracy and efficiency than existing deep learning methods. Furthermore, we jointly train SSAN with SKE and SPD and achieve an end-to-end improvement. Meanwhile, the two key modules SKE and SPD can also be effectively inserted into other video retrieval pipelines and gain considerable performance improvements. Experimental results on public datasets show that SSAN can obtain higher alignment accuracy while saving storage and online query computational cost compared to existing methods.
Chen Jiang 0006, Kaiming Huang, Sifeng He, Lei Yang 0061, Qing Wang 0068, Furong Xu, Tan Pan
ACM Multimedia3
2020 A Novel Self-Feedback Intelligent Vision Measure for Fast and Accurate Alignment in Flip-Chip Packaging
abstract
A template matching (TM) algorithm has been widely employed in a visual inspection process of waferlevel flip-chip packaging, but the structures of traditional TM algorithms are always direct feedforward, which leads to the difficulty in achieving fast speed and high accuracy at the same time. The motivation of this article is to combine the ability to enable the chip visual measurement running in a fast-speed and high-accuracy manner. First, a novel selffeedback intelligent template matching (SFI-TM) structure is proposed, which can enable the intermediate information in the matching process to be fully utilized. The intelligent speed and resolution regulation rules are combined to construct the SFI-TM algorithm. Then, the reliability and robustness of the SFI-TM algorithm are theoretically analyzed to make sure it works in a stable manner. Finally, a series of practical chip alignment visual inspection experiments, including parameter testing, comparing experiments with the other five proposed visual detection algorithms, and robustness testing, are carried out in detail, respectively. The experimental results indicate that the SFI-TM algorithm can achieve the highest average measurement accuracy (1.08 μm) with almost the same measurement speed as fast as Tiny YOLOv2, and that it can resist the brightness variation from -20 to 30 gray value, the pepper-salt noise with a density of 0.5/pixel, and the Gaussian noise with a large variance of 2.5. In addition, it has good robustness against blurring and distortion of images under the dynamic speed motion processes.
Zelong Wu, Hui Tang 0003, Zhaoyang Feng, Sifeng He, Jian Gao 0002, Xin Chen 0005, Yunbo He, Xun Chen 0002
IEEE Trans. Ind. Informatics5
2019 Fast Super-Resolution in MRI Images Using Phase Stretch Transform, Anchored Point Regression and Zero-Data Learning
abstract
Medical imaging is fundamentally challenging due to absorption and scattering in tissues and by the need to minimize illumination of the patient with harmful radiation. Common problems are low spatial resolution, limited dynamic range and low contrast. These predicaments have fueled interest in enhancing medical images using digital post processing. In this paper, we propose and demonstrate an algorithm for real-time inference that is suitable for edge computing. Our locally adaptive learned filtering technique named Phase Stretch Anchored Regression (PhSAR) combines the Phase Stretch Transform for local features extraction in visually impaired images with clustered anchored points to represent image feature space and fast regression based learning. In contrast with the recent widely-used deep neural network for image super-resolution, our algorithm achieves significantly faster inference and less hallucination on image details and is interpretable. Tests on brain MRI images using zero-data learning reveal its robustness with explicit PSNR improvement and lower latency compared to relevant benchmarks.
Sifeng He, Bahram Jalali
ICIP1
2017 A regularized on-line sequential extreme learning machine with forgetting property for fast dynamic hysteresis modeling
abstract
Piezoelectric ceramics(PZT)actuator has been widely used in flexure-guided nanopositioning stage because of their high resolution. However, it is quite hard to achieve high-rate precision positioning control because of the complex hysteresis nonlinearity effect of PZT actuator. Thus, an online RELM algorithm with forgetting property(FReOS-ELM) is proposed to handle this issue. Firstly, we adopt regularized extreme learning machine(RELM)to build an intelligent hysteresis model. The training of the algorithm is completed only in one step, which avoids the shortcomings of the traditional hysteresis model based on artificial neural network(ANN) that slow training speed and easy to fall into the local minimum. Then, based on the regularized on-line sequential extreme learning machine(ReOS-ELM), an on-line RELM algorithm with forgetting property(FReOS-ELM) is designed, which can avoid the computational load of ReOS-ELM in the process of adding new data for learning on-line. In the experiment, a real-time voltage signal with varying frequencies and amplitudes is adopted, and the output displacement data of the nanopositioning stage is also acquired and analyzed. The results powerfully verify that the performance of the established hysteresis model based on the proposed FReOS-ELM is satisfactory, which can be used to improve the practical positioning performance for flexure nanopositioning stage.
Zelong Wu, Hui Tang 0003, Sifeng He, Jian Gao 0002, Xin Chen 0005, Chengqiang Cui, Yunbo He, Yangmin Li 0001
IROS3