Yongfei Liu

dblp:121/0754 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A multi-scale fusion attention network for small-object defect detection in power transmission lines
Yongfei Liu, Shuaina Huang, Zongqi You
Eng. Appl. Artif. Intell.1
2025 CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
abstract
Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, Zhi Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kechi Zhang, Ge Li 0001, Yihong Dong, Yongfei Liu, Zhi Jin 0001
ACL (1)7
2025 AEFA: An Ensemble Framework for Fraud Detection in the Forex Market
Weiyuan Wang, Jianke Yu, Zhengyi Yang 0001, Mingchen Ju, Shuyue Yu, Jinglin Wu, Lifan Liu, Yongfei Liu, John Shepherd 0001, Wenjie Zhang 0001
ADMA (3)8
2025 Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
abstract
Preference alignment in Large Language Models (LLMs) has significantly improved their ability to adhere to human instructions and intentions. However, existing direct alignment algorithms primarily focus on relative preferences and often overlook the qualitative aspects of responses, despite having access to preference data that includes reward scores from judge models during AI feedback. Striving to maximize the implicit reward gap between the chosen and the slightly inferior rejected responses can cause overfitting and unnecessary unlearning of the high-quality rejected responses. The unawareness of the reward scores also drives the LLM to indiscriminately favor the low-quality chosen responses and fail to generalize to optimal responses that are sparse in data. To overcome these shortcomings, our study introduces reward-conditioned LLM policies that discern and learn from the entire spectrum of response quality within the dataset, helping extrapolate to more optimal regions. We propose an effective yet simple data relabeling method that conditions the preference pairs on quality scores to construct a reward-augmented dataset. The experiments across various benchmarks and diverse models demonstrate that our approach consistently boosts DPO by a considerable margin. Through comprehensive ablation studies, we demonstrate that our method not only maximizes the utility of preference data but also mitigates the issue of unlearning, demonstrating its broad effectiveness beyond mere data expansion. Our code is available at https://github.com/shenao-zhang/reward-augmented-preference.
Shenao Zhang, Boyi Liu 0001, Yufeng Zhang 0007, Yingxiang Yang, Yongfei Liu, Liyu Chen, Zhaoran Wang 0001
ICML6
2024 Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
abstract
In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-language connector has been relatively less explored. In this study, we aim to propose a strong vision-language connector that enables MLLM to simultaneously achieve high accuracy and low computation cost. We first reveal the existence of the visual anchors in Vision Transformer and propose a cost-effective search algorithm to progressively extract them. Building on these findings, we introduce the Anchor Former (AcFormer), a novel vision-language connector designed to leverage the rich prior knowledge obtained from these visual anchors during pretraining, guiding the aggregation of information. Through extensive experimentation, we demonstrate that the proposed method significantly reduces computational costs by nearly two-thirds, while simultaneously outperforming baseline methods. This highlights the effectiveness and efficiency of AcFormer.
Haogeng Liu, Quanzeng You, Yongfei Liu, Huaibo Huang, Ran He 0001, Hongxia Yang
NeurIPS4
2023 Cascade Sparse Feature Propagation Network for Interactive Segmentation
Chuyu Zhang, Hui Ren 0003, Chuanyang Hu, Yongfei Liu, Xuming He 0001
BMVC4
2023 Intrusion Detection Based on Sampling and Improved OVA Technique on Imbalanced Data
abstract
Network-based Intrusion Detection(NID) is an effective means to deal with network attacks. NID is able to detect different types of network attacks by analyzing network traffic. However, in the real world, network traffic contains majority and minority class attacks as well as a large number of normal traffic samples. The imbalance in the number of training samples of various types of network traffic makes network intrusion detection very poor. Due to the lack of training samples, traditional NID can’t learn the characteristics of minority class attacks, which leads to the failure of NID to detect minority class attacks. Therefore, in order to solve the problem brought by imbalanced data, we propose a network intrusion detection algorithm based on the sampling and improved One-vs-All(OVA) technique. The dataset is balanced by downsampling the majority class data based on K-means clustering and oversampling the minority class data based on Auxiliary Classifier Generative Adversarial Network(ACGAN), improve classification accuracy through OVA-based model training and testing. We conduct validation experiments on the NSL-KDD dataset, and the experimental results show that the proposed method achieves excellent results in terms of Accuracy, Precision, Recall and F1-score. Compared with existing state-of-the-art methods, the proposed method not only achieves excellent detection performance with low false positive rate, but also addresses the learning problem of imbalanced data more effectively.
Yongfei Liu, Hong Li 0004, Wenyuan Zhang 0002, Fei Lyu 0001, Shuaizong Si
CSCWD1
2023 HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models
abstract
Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Recently, Contrastive Language-Image Pre-training (CLIP) has shown great potential in providing interaction prior for HOI detectors via knowledge distillation. However, such approaches often rely on large-scale training data and suffer from inferior performance under few/zero-shot scenarios. In this paper, we propose a novel HOI detection framework that efficiently extracts prior knowledge from CLIP and achieves better generalization. In detail, we first introduce a novel interaction decoder to extract informative regions in the visual feature map of CLIP via a cross-attention mechanism, which is then fused with the detection backbone by a knowledge integration block for more accurate human- object pair detection. In addition, prior knowledge in CLIP text encoder is leveraged to generate a classifier by embedding HOI descriptions. To distinguish fine-grained interactions, we build a verb classifier from training data via visual semantic arithmetic and a lightweight verb representation adapter. Furthermore, we propose a training-free enhancement to exploit global HOI predictions from CLIP. Extensive experiments demonstrate that our method outperforms the state of the art by a large margin on various settings, e.g. +4.04 mAP on HICO-Det. The source code is available in https://github.com/Artanic30/HOICLIP.
Shan Ning, Longtian Qiu, Yongfei Liu, Xuming He 0001
CVPR3
2023 Grounded Image Text Matching with Mismatched Relation Reasoning
abstract
This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating vision-language (VL) models on this task, with a focus on the challenging settings of limited training data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained VL models often lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. Our RCRN can he interpreted as a modular program and delivers strong performance in terms of both length generalization and data efficiency. The code and data are available on https://githuh.coin/SHTUPLUS/GITM-MR.
Yu Wu 0014, Yana Wei, Haozhe Wang 0002, Yongfei Liu, Sibei Yang, Xuming He 0001
ICCV4
2023 Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning
Yongfei Liu, Desen Zhou, Tinne Tuytelaars, Xuming He 0001
ICLR2
2023 A New Federated Learning Model for Host Intrusion Detection System Under Non-IID Data
abstract
Host Intrusion Detection System (HIDS) is an important research topic in the field of cyberspace security. With the explosion in the number of malicious attacks in recent years, machine learning-based detection method is now the most common and efficient approach. While traditional centralized machine learning needs to transmit data to the central server for training, which not only requires the central server to have large computing resources, but also causes problems such as sensitive data leakage and communication overhead. As a distributed machine learning paradigm, Federated Learning (FL) can achieve multi-party collaborative training and aggregate a unified global model without data sharing, which can well alleviate these problems. It is worth noting that existing studies on the use of FL in HIDS are all conducted in the scenario where the data is independent and identically distributed (IID). However, due to the different context of hosts, the data generated by hosts is usually non-independent and identically distributed (Non-IID) in reality. Therefore, We investigate the impact of Non-IID data with different skew levels on FL in HIDS. On this basis, we propose a data augmentation FL algorithm based on Synthetic Minority Over-Sampling Technique (SMOTE) to reduce the impact of Non-IID data. We also develop a data collection module using extended Berkeley Packet Filter (eBPF) technology to collect a dataset for experiments. Experimental results show that our proposed FL algorithm can effectively improve the performance of HIDS under Non-IID data.
Yongfei Liu, Lanxue Zhang, Liangxiong Li, Tong Li 0012, Bingzhen Wu
SMC3
2022 Edge Federated Learning for Social Profit Optimality: A Cooperative Game Approach
Wenyuan Zhang 0002, Guangjun Wu, Yongfei Liu, Binbin Li 0003
CollaborateCom (1)3
2022 VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers
abstract
Breakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems. However, although visualization and interpretability tools have become available for NLP models, internal mechanisms of vision and multimodal transformers remain largely opaque. With the success of these transformers, it is increasingly critical to understand their inner workings, as unraveling these black-boxes will lead to more capable and trustworthy models. To contribute to this quest, we propose VL-InterpreT, which provides novel interactive visualizations for interpreting the attentions and hidden representations in multimodal transformers. VL-InterpreT is a task agnostic and integrated tool that (1) tracks a variety of statistics in attention heads throughout all layers for both vision and language components, (2) visualizes cross-modal and intra-modal attentions through easily readable heatmaps, and (3) plots the hidden representations of vision and language tokens as they pass through the transformer layers. In this paper, we demonstrate the functionalities of VL-InterpreT through the analysis of KD-VLP, an end-to-end pretraining vision-language multimodal transformer-based model, in the tasks of Visual Commonsense Reasoning (VCR) and WebQA, two visual question answering benchmarks. Furthermore, we also present a few interesting findings about multimodal transformer behaviors that were learned through our tool.
Estelle Aflalo, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan 0001, Vasudev Lal
CVPR4
2022 Federated Learning-Based Intrusion Detection on Non-IID Data
Yongfei Liu, Guangjun Wu, Wenyuan Zhang 0002, Jun Li 0085
ICA3PP1
2021 Relation-aware Instance Refinement for Weakly Supervised Visual Grounding
abstract
Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize weak supervision from only image-caption pairs. Previous methods typically rely on matching query phrases directly to a precomputed, fixed object candidate pool, which leads to inaccurate localization and ambiguous matching due to lack of semantic relation constraints. In our paper, we propose a novel context-aware weakly-supervised learning method that incorporates coarse-to-fine object refinement and entity relation modeling into a two-stage deep network, capable of producing more accurate object representation and matching. To effectively train our network, we introduce a self-taught regression loss for the proposal locations and a classification loss based on parsed entity relations. Extensive experiments on two public benchmarks Flickr30K Entities and ReferItGame demonstrate the efficacy of our weakly grounding framework. The results show that we outperform the previous methods by a considerable margin, achieving 59.27% top-1 accuracy in Flickr30K Entities and 37.68% in the ReferItGame dataset respectively1.
Yongfei Liu, Lin Ma 0002, Xuming He 0001
CVPR1
2020 Learning Cross-Modal Context Graph for Visual Grounding
abstract
Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic ambiguities. Prior works typically focus on learning representations of individual phrases with limited context information. To address their limitations, this paper proposes a language-guided graph representation to capture the global context of grounding entities and their relations, and develop a cross-modal graph matching strategy for the multiple-phrase visual grounding task. In particular, we introduce a modular graph neural network to compute context-aware representations of phrases and object proposals respectively via message propagation, followed by a graph-based matching module to generate globally consistent localization of grounding phrases. We train the entire graph neural network jointly in a two-stage strategy and evaluate it on the Flickr30K Entities benchmark. Extensive experiments show that our method outperforms the prior state of the arts by a sizable margin, evidencing the efficacy of our grounding framework. Code is available at https://github.com/youngfly11/LCMCG-PyTorch.
Yongfei Liu, Xiaodan Zhu 0001, Xuming He 0001
AAAI1
2020 Part-Aware Prototype Network for Few-Shot Semantic Segmentation
Yongfei Liu, Xiangyi Zhang, Songyang Zhang 0001, Xuming He 0001
ECCV (9)1
2019 Pose-Aware Multi-Level Feature Network for Human Object Interaction Detection
abstract
Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring relation instances and subtle visual difference between relation categories. To address those challenges, we propose a multi-level relation detection strategy that utilizes human pose cues to capture global spatial configurations of relations and as an attention mechanism to dynamically zoom into relevant regions at human part level. We develop a multi-branch deep network to learn a pose-augmented relation representation at three semantic levels, incorporating interaction context, object features and detailed semantic part cues. As a result, our approach is capable of generating robust predictions on fine-grained human object interactions with interpretable outputs. Extensive experimental evaluations on public benchmarks show that our model outperforms prior methods by a considerable margin, demonstrating its efficacy in handling complex scenes.
Desen Zhou, Yongfei Liu, Rongjie Li, Xuming He 0001
ICCV3
2016 Parallel Implementation of the Range-Doppler Radar Processing on a GPU Architecture
abstract
Graphic processing units (GPUs) is widely used to accelerate the processing speed of the radar detection procedure, including the range compression, coherent integration and constant false alarm rate. Specifically, detailed parallel design of the radar algorithm and the thread programming are shown. The experimental results show that, by engaging the parallel technology into the radar processing procedure, much high speedup ratio can be obtained. Furthermore, precise target detection can be guaranteed.
Guanghui Zhao 0003, Yongfei Liu, Shuping Zhang, Fangfang Shen, Yaohai Lin, Guangming Shi
ISPDC2