Yibing Liu

dblp:63/1829 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
abstract
Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on the critical challenge of how to restore surrogate-driven protected data in diverse MLLM scenarios. We first bridge this research gap by contributing the SPPE (Surrogate Privacy Protected Editable) dataset, which includes a wide range of privacy categories and user instructions to simulate real MLLM applications. This dataset offers protected surrogates alongside their various MLLM-edited versions, thus enabling the direct assessment of privacy recovery quality. By formulating privacy recovery as a guided generation task conditioned on complementary multimodal signals, we further introduce a unified approach that reliably reconstructs private content while preserving the fidelity of MLLM-generated edits. The experiments on both SPPE and InstructPix2Pix further show that our approach generalizes well across diverse visual content and editing tasks, achieving a strong balance between privacy protection and MLLM usability.
Yibing Liu, Peilin Chen 0001, Yung-Hui Li, Shiqi Wang 0001, Sam Kwong
AAAI2
2026 Robust and flexible multi-view subspace clustering with nuclear norm
Shaojun Shi, Yibing Liu, Canyu Zhang 0001, Feiping Nie 0001
Pattern Recognit.2
2025 Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
abstract
We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e.g.,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs.
Kecheng Chen, Hui Liu 0036, Jie Liu 0044, Yibing Liu, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li
NeurIPS5
2025 Trend-constrained pairing based incremental transfer learning for remaining useful life prediction of bearings in wind turbines
Xilin Li, Wei Teng, Luo Wang, Jingpeng Hu, Dikang Peng, Yibing Liu
Expert Syst. Appl.7
2024 Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization
abstract
The out-of-distribution (OOD) problem generally arises when neural networks encounter data that significantly deviates from the training data distribution, i.e., in-distribution (InD). In this paper, we study the OOD problem from a neuron activation view. We first formulate neuron activation states by considering both the neuron output and its influence on model decisions. Then, to characterize the relationship between neurons and OOD issues, we introduce the *neuron activation coverage* (NAC) -- a simple measure for neuron behaviors under InD data. Leveraging our NAC, we show that 1) InD and OOD inputs can be largely separated based on the neuron behavior, which significantly eases the OOD detection problem and beats the 21 previous methods over three benchmarks (CIFAR-10, CIFAR-100, and ImageNet-1K). 2) a positive correlation between NAC and model generalization ability consistently holds across architectures and datasets, which enables a NAC-based criterion for evaluating model robustness. Compared to prevalent InD validation criteria, we show that NAC not only can select more robust models, but also has a stronger correlation with OOD test performance.
Yibing Liu, Chris Xing Tian, Haoliang Li, Lei Ma 0003, Shiqi Wang 0001
ICLR1
2024 Multi-view Spectral Clustering Based on Topological Manifold Learning
Shaojun Shi, Yibing Liu, Xueling Chen
PRCV (1)2
2024 M$^{3}$3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
abstract
Face presentation attacks (FPA), also known as face spoofing, have brought increasing concerns to the public through various malicious applications, such as financial fraud and privacy leakage. Therefore, safeguarding face recognition systems against FPA is of utmost importance. Although existing learning-based face anti-spoofing (FAS) models can achieve outstanding detection performance, they lack generalization capability and suffer significant performance drops in unforeseen environments. Many methodologies seek to use auxiliary modality data (e.g., depth and infrared maps) during the presentation attack detection (PAD) to address this limitation. However, these methods can be limited since (1) they require specific sensors such as depth and infrared cameras for data capture, which are rarely available on commodity mobile devices, and (2) they cannot work properly in practical scenarios when either modality is missing or of poor quality. In this paper, we devise an accurate and robustMultiModalMobileFaceAnti-Spoofing system namedM$^{3}$FASto overcome the issues above. The primary innovation of this work lies in the following aspects: (1) To achieve robust PAD, our system combines visual and auditory modalities using three commonly available sensors: camera, speaker, and microphone; (2) We design a novel two-branch neural network with three hierarchical feature aggregation modules to perform cross-modal feature fusion; (3). We propose a multi-head training strategy, allowing the model to output predictions from the vision, acoustic, and fusion heads, resulting in a more flexible PAD. Extensive experiments have demonstrated the accuracy, robustness, and flexibility of M$^{3}$FAS under various challenging experimental settings. The source code and dataset are available at:https://github.com/ChenqiKONG/M3FAS/.
Chenqi Kong, Kexin Zheng, Yibing Liu, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li
IEEE Trans. Dependable Secur. Comput.3
2024 Generalization Beyond Feature Alignment: Concept Activation-Guided Contrastive Learning
abstract
Learning invariant representations via contrastive learning has seen state-of-the-art performance in domain generalization (DG). Despite such success, in this paper, we find that its core learning strategy - feature alignment - could heavily hinder model generalization. Drawing insights in neuron interpretability, we characterize this problem from a neuron activation view. Specifically, by treating feature elements as neuron activation states, we show that conventional alignment methods tend to deteriorate the diversity of learned invariant features, as they indiscriminately minimize all neuron activation differences. This instead ignores rich relations among neurons - many of them often identify the same visual concepts despite differing activation patterns. With this finding, we present a simple yet effective approach, Concept Contrast (CoCo), which relaxes element-wise feature alignments by contrasting high-level concepts encoded in neurons. Our CoCo performs in a plug-and-play fashion, thus it can be integrated into any contrastive method in DG. We evaluate CoCo over four canonical contrastive methods, showing that CoCo promotes the diversity of feature representations and consistently improves model generalization capability. By decoupling this success through neuron coverage analysis, we further find that CoCo potentially invokes more meaningful neurons during training, thereby improving model learning.
Yibing Liu, Chris Xing Tian, Haoliang Li, Shiqi Wang 0001
IEEE Trans. Image Process.1
2024 A Multi-Objective Resource Pre-Allocation Scheme Using SDN for Intelligent Transportation System
abstract
As 5-th Generation (5G) mobile communication and edge computing technologies mature, Intelligent Transportation System (ITS) are gradually becoming a reality. In the 5G heterogeneous network, resources such as computing, storage, and communication are allocated to each Road Side Unit (RSU) to provide intelligent services for vehicles. However, the existing average allocation method based on historical experience can easily lead to over-concentration or insufficient resources, which causes waste and reduces the Quality of Service (QoS). To solve this problem, this paper proposes a Multi-Objective Neural Time-series Prediction (M-ONTP) scheme for resource pre-allocation scenario in ITS. The scheme takes into account the complexity and diversity of service resource, innovatively treats the number of vehicles and communication power as joint optimization metrics, and proposes a multi-objective learning model. Benefiting from the vehicle data collected by RSUs in real time, we utilize historical traffic information to predict future road load and rely on Software Defined Network (SDN) to design a flexible resource pre-allocated architecture for ITS. To enhance the effectiveness of feature capture, M-ONTP also organically integrates various neural networks, which can appropriately handle large-scale time-series traffic flow. And we choose two layers of road data for fitting, which ensures that the model has a wide horizon to receive sufficient information. SUMO-based simulation experiments show that our scheme accurately realizes the prediction of joint objective and has significant performance advantage over other models. Meanwhile, our pre-allocation strategy reduces the total resource consumption by about 7%, increases the sufficiency rate by about 7%, and decreases the redundancy by about 12% while ensuring enough service resource to maintain normal QoS, which validates the effectiveness of M-ONTP.
Yibing Liu, Lijun Huo, Xiongtao Zhang, Jun Wu 0001
IEEE Trans. Intell. Transp. Syst.1
2024 Swarm Learning and Knowledge Distillation Empowered Self-Driving Detection Against Threat Behavior for Intelligent IoT
abstract
The combination of mobile communication and the Internet of Things (IoT) has made physical devices more intelligent, bringing great convenience to our lives. However, the deep integration of personal information and the Internet increases the risk of data leakage and is easily exploited maliciously. In addition, due to limited system resources, smart devices with lightweight design are required. Therefore, it is necessary to realize low-energy and effective abnormal behavior detection of IoT devices, but the existing detection methods have disadvantages such as leakage of user privacy, low accuracy, and difficulty in dynamically improving the effect. To address these issues, this paper proposes a dynamic interactive minor anomaly detection scheme called ADONIS based on Swarm Learning (SL). The scheme combines the concept of swarm defense and utilizes SL to achieve local data fusion, which improves the detection effect and protects user privacy. Moreover, the decentralized structure of SL can cope with the impact of single node damage to enhance the robustness of IoT services. Furthermore, we propose training and detection decoupling framework to achieve high accuracy, low energy consumption, and low latency. It improves the performance by fitting the training model with full data, and simplifies the complexity of the detection model using knowledge distillation. We also design a self-enhancing dynamic strategy based on the decoupling framework to maintain powerful detection capability through human-computer interaction (HCI) and continuous learning. The framework relies on traffic data to keep the model sensitive to new behavior through iterative training without disturbing the user. Finally, simulation experiments show that our proposed scheme can achieve 82.2% accuracy, reduce the average detection time to 8.22$ ms$, and simplify the model complexity by 15.9%. Compared with existing methods, ADONIS can provide lighter, safer and more accurate anomaly detection.
Yibing Liu, Xiongtao Zhang, Lijun Huo, Jun Wu 0001, Mohsen Guizani
IEEE Trans. Mob. Comput.1
2023 Green Floating Blockchain-Empowered Co-Trust Security Mechanism with Energy Efficiency Against Attack Threat for 6G-IoV
abstract
The Internet of Vehicles (IoV) based on the 6-th Generation (6G) communication brings convenience, but also raises anxiety about information security. Researchers have developed static security schemes based on the blockchain, but it results in excessive resource occupation. And it is difficult to resist some attacks such as desynchronization, Denial of Service (DoS), and covert intrusion. In response to the above problems, this paper proposes the Floating Blockchain Consensus Security (FBCS) scheme. It models the attack risk of Road Side Unit (RSU) through the traffic flow prediction to present the credibility evaluation and constructs the floating blockchain based on the results. And the security capability is adjusted through the dynamic joining and exit mode of trusted nodes and untrusted nodes. FBCS also establishes a relationship between blockchain size and energy consumption. Once the system resources are found to be insufficient, it applies the cloud center-based supplementary certification mechanism to provide authentication to untrusted nodes to enhance the stability of the IoV.Theoretical analysis and simulation experiments prove that the FBCS can afford data privacy, maintain moderate security capabilities, and adapt defense capabilities according to the attack environment, reduce resource occupation to save cost.
Yibing Liu, Lijun Huo, Hansong Xu, Jun Wu 0001
GLOBECOM1
2023 MRSA: Mask Random Array Protocol for Efficient Secure Handover Authentication in 5G HetNets
abstract
The emergence of new communication applications adds high heterogeneity to 5G-networks. With the increase of heterogeneity, handover of user equipment between different service HetNets is frequent. It must smoothly realize user-free switching to provide services continuously. Although the 3 rd Generation Partnership Project (3GPP) has proposed a standard protocol for this scenario, it is found that these protocols cannot satisfy key forward/backward secrecy, lacks mutual authentication, etc. Further, it can be subjected to replay, DoS and other attacks. To alleviate these problems, we propose a mask random array protocol, MRSA. For efficient, secure handover authentication in 5G HetNets, we first design a verification mechanism called mask array, which depends on a random number self-circulating encryption structure. The mechanism can not only check the identity of the communication entity but also evaluate the freshness of the message. Second, we devise the mask array-based key derivation method to ensure the whole mechanism's key security. Third, formal proof and automated analysis are established to verify the efficiency and safety of the proposed MRSA protocol. Finally, function and robustness analysis illustrate the ability to resist attacks, while the simulation base station communication analysis shows the efficiency of the protocol from three aspects of data, time and energy. MRSA has significant performance advantages compared to existing schemes in 5G HetNets.
Yibing Liu, Lijun Huo, Jun Wu 0001, Mohsen Guizani
IEEE Trans. Dependable Secur. Comput.1
2023 Swarm Learning-Based Dynamic Optimal Management for Traffic Congestion in 6G-Driven Intelligent Transportation System
abstract
As city boundaries expand and the vehicles continues to proliferate, the transportation system is increasingly overloaded, greatly increasing people’s commuting burden and extending the resulting negative effects to all areas of work and life. It is a big issue that needs to be solved urgently. However, due to the development of infrastructure and technologies in 6G-driven Intelligent Transportation Systems (ITS), it becomes possible to alleviate urban congestion. Existing solutions either optimize the path planning of each vehicle, or only focus on solving the problem of resource allocation of a single road, neither can take advantage of self-organizing networks and easily fall into local optimum. Combining the above reasons, we propose the Direction Decide as a Service (DDaaS) scheme. First, it contains a novel three-layer service architecture based on Swarm Learning (SL), which enables orderly transmission of traffic data and control instructions and protects user privacy. Second, an improved local model and aggregation method is incorporated into DDaaS, which enables to make accurate predictions when the road resources at a single intersection are insufficient. Third, we propose a dynamic traffic control algorithm to provide signal light switching decisions for rapidly changing ITS. Finally, constructing an urban road simulation experiment combined with SUMO, we prove that DDaaS can reduce traffic congestion effectively and has significant advantages compared to other schemes.
Yibing Liu, Lijun Huo, Jun Wu 0001, Ali Kashif Bashir
IEEE Trans. Intell. Transp. Syst.1
2022 Rethinking Attention-Model Explainability through Faithfulness Violation Test
abstract
Attention mechanisms are dominating the explainability of deep models. They produce probability distributions over the input, which are widely deemed as feature-importance indicators. However, in this paper, we find one critical limitation in attention explanations: weakness in identifying the polarity of feature impact. This would be somehow misleading – features with higher attention weights may not faithfully contribute to model predictions; instead, they can impose suppression effects. With this finding, we reflect on the explainability of current attention-based techniques, such as Attention $\bigodot$ Gradient and LRP-based attention explanations. We first propose an actionable diagnostic methodology (henceforth faithfulness violation test) to measure the consistency between explanation weights and the impact polarity. Through the extensive experiments, we then show that most tested explanation methods are unexpectedly hindered by the faithfulness violation issue, especially the raw attention. Empirical analyses on the factors affecting violation issues further provide useful observations for adopting explanation methods in attention models.
Yibing Liu, Haoliang Li, Chenqi Kong, Jing Li 0049, Shiqi Wang 0001
ICML1
2022 A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQA
abstract
Knowledge-based Visual Question Answering (VQA) expects models to rely on external knowledge for robust answer prediction. Though significant it is, this paper discovers several leading factors impeding the advancement of current state-of-the-art methods. On the one hand, methods which exploit the explicit knowledge take the knowledge as a complement for the coarsely trained VQA model. Despite their effectiveness, these approaches often suffer from noise incorporation and error propagation. On the other hand, pertaining to the implicit knowledge, the multi-modal implicit knowledge for knowledge-based VQA still remains largely unexplored. This work presents a unified end-to-end retriever-reader framework towards knowledge-based VQA. In particular, we shed light on the multi-modal implicit knowledge from vision-language pre-training models to mine its potential in knowledge reasoning. As for the noise problem encountered by the retrieval operation on explicit knowledge, we design a novel scheme to create pseudo labels for effective knowledge supervision. This scheme is able to not only provide guidance for knowledge retrieval, but also drop these instances potentially error-prone towards question answering. To validate the effectiveness of the proposed method, we conduct extensive experiments on the benchmark dataset. The experimental results reveal that our method outperforms existing baselines by a noticeable margin. Beyond the reported numbers, this paper further spawns several insights on knowledge utilization for future research with some empirical findings.
Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng 0001, Mohan Kankanhalli
ACM Multimedia4
2022 TR-AKA: A two-phased, registered authentication and key agreement protocol for 5G mobile networks
abstract
Abstract With the development of mobile communications, the authors step into the 5G era. 5G‐AKA is intended to serve as a certification standard for users to safely and stably enjoy 5G mobile services. However, recent research proves that some attacks may happened based on information deciphering and message replay. In addition, the authors found that some variants of 5G‐AKA are also at risk of linkability attacks. Therefore, in order to solve the above‐mentioned risks, the authors propose a two‐phased, registered protocol TR‐AKA. This protocol finishes two‐way authentication between the user and the home network through only two message‐sending rounds, and a temporary identity variable group is applied to replace the real information for transmission. A newly designed variable is also used to strictly control the increase of the sequence number to promote the sensitivity of freshness. Finally, the authors choose the Tamarin tool to prove that TR‐AKA has achieved authentication and privacy‐protection objectives, and further discussions have also verified the rationality of the protocol design.
Yibing Liu, Lijun Huo
IET Inf. Secur.1
2022 Answer Questions with Right Image Regions: A Visual Attention Regularization Approach
abstract
Visual attention in Visual Question Answering (VQA) targets at locating the right image regions regarding the answer prediction, offering a powerful technique to promote multi-modal understanding. However, recent studies have pointed out that the highlighted image regions from the visual attention are often irrelevant to the given question and answer, leading to model confusion for correct visual reasoning. To tackle this problem, existing methods mostly resort to aligning the visual attention weights with human attentions. Nevertheless, gathering such human data is laborious and expensive, making it burdensome to adapt well-developed models across datasets. To address this issue, in this article, we devise a novel visual attention regularization approach, namely, AttReg, for better visual grounding in VQA. Specifically, AttReg first identifies the image regions that are essential for question answering yet unexpectedly ignored (i.e., assigned with low attention weights) by the backbone model. And then a mask-guided learning scheme is leveraged to regularize the visual attention to focus more on these ignored key regions. The proposed method is very flexible and model-agnostic, which can be integrated into most visual attention-based VQA models and require no human attention supervision. Extensive experiments over three benchmark datasets, i.e., VQA-CP v2, VQA-CP v1, and VQA v2, have been conducted to evaluate the effectiveness of AttReg. As a by-product, when incorporating AttReg into the strong baseline LMH, our approach can achieve a new state-of-the-art accuracy of 60.00% with an absolute performance gain of 7.01% on the VQA-CP v2 benchmark dataset. In addition to the effectiveness validation, we recognize that the faithfulness of the visual attention in VQA has not been well explored in literature. In the light of this, we propose to empirically validate such property of visual attention and compare it with the prevalent gradient-based approaches.
Yibing Liu, Jianhua Yin 0001, Xuemeng Song, Weifeng Liu 0001, Liqiang Nie, Min Zhang 0005
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
abstract
Benefiting from the advancement of computer vision, natural language processing and information retrieval techniques, visual question answering (VQA), which aims to answer questions about an image or a video, has received lots of attentions over the past few years. Although some progress has been achieved so far, several studies have pointed out that current VQA models are heavily affected by the language prior problem, which means they tend to answer questions based on the co-occurrence patterns of question keywords (e.g., how many) and answers (e.g., 2) instead of understanding images and questions. Existing methods attempt to solve this problem by either balancing the biased datasets or forcing models to better understand images. However, only marginal effects and even performance deterioration are observed for the first and second solution, respectively. In addition, another important issue is the lack of measurement to quantitatively measure the extent of the language prior effect, which severely hinders the advancement of related techniques.
Zhiyong Cheng 0001, Liqiang Nie, Yibing Liu, Yinglong Wang 0001, Mohan Kankanhalli
SIGIR4
2006 Soft Sensor Based on Support Vector Machine for Effective Wind Speed in Large Variable Wind
abstract
Accurate estimation of wind speed can improve control capabilities of wind turbines. The wind speed is generally different at every point on the surface covered by the blades due to three dimensional and time-varied wind fields; effective wind speed cannot be measured by an anemometer. In this paper the problem of effective wind speed estimation is considered as soft sensor. Soft sensor modeling utilizes technique of support vector machine, which is efficient for the problem characterized by small sample, nonlinearity, high dimension, local minima and has high generalization. Comparing with Kalman filter, simulation results show support vector machine is an effective method for soft sensor modeling. Effective wind speed estimation can exactly track the trend of wind speed and has high estimation precision.
Xiyun Yang, Xiaojuan Han, Lingfeng Xu, Yibing Liu
ICARCV4