Ruibin Li

dblp:259/8982 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A data-driven approach for sustainable maintenance: Wind-induced spalling evaluation of cement-based structures
Xiaoning Cui, Xueqing Xu, Ruibin Li, Fu Yu
Eng. Appl. Artif. Intell.4
2026 Channel-Aware Optimal Transport: A Theoretical Framework for Generative Communication
Xiqiang Qu, Ruibin Li, Jun Chen 0005, Lei Yu 0003, Xinbing Wang
IEEE Trans. Inf. Theory2
2025 Mjölnir: Breaking the Shield of Perturbation-Protected Gradients via Adaptive Diffusion
abstract
Perturbation-based mechanisms, such as differential privacy, mitigate gradient leakage attacks by introducing noise into the gradients, thereby preventing attackers from reconstructing clients' private data from the leaked gradients. However, can gradient perturbation protection mechanisms truly defend against all gradient leakage attacks? In this paper, we present the first attempt to break the shield of gradient perturbation protection in Federated Learning for the extraction of private information. We focus on common noise distributions, specifically Gaussian and Laplace, and apply our approach to DNN and CNN models. We introduce Mjölnir, a perturbation-resilient gradient leakage attack that is capable of removing perturbations from gradients without requiring additional access to the original model structure or external data. Specifically, we leverage the inherent diffusion properties of gradient perturbation protection to develop a novel diffusion-based gradient denoising model for Mjölnir. By constructing a surrogate client model that captures the structure of perturbed gradients, we obtain crucial gradient data for training the diffusion model. We further utilize the insight that monitoring disturbance levels during the reverse diffusion process can enhance gradient denoising capabilities, allowing Mjölnir to generate gradients that closely approximate the original, unperturbed versions through adaptive sampling steps. Extensive experiments demonstrate that Mjölnir effectively recovers the protected gradients and exposes the Federated Learning process to the threat of gradient leakage, achieving superior performance in gradient denoising and private data recovery.
Xuan Liu 0001, Siqi Cai 0001, Qihua Zhou, Song Guo 0001, Ruibin Li, Kaiwei Lin
AAAI5
2025 DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement
abstract
Recent years have witnessed the rise of Neural-enhanced Video Streaming (NeVS), which integrates neural restoration models into video codecs for higher compression-restoration performance. Despite its benefit, existing work has not well explored the full potential of NeVS paradigm, due to: (1) post-streaming restoration by decoder while lacking the proactive collaboration of encoder, (2) end-to-end optimization based on conventional rate-distortion theory, which has been verified that low distortion is not always a synonym for high perceptual quality, and (3) coupled design for domain-specific tasks that cannot generalize to various video codecs. Observing these limitations, our objective is not to incrementally present an improved restoration model. Instead, we focus on the encoder-decoder synergy, i.e., the codec, which is non-trivial since it inherently strikes the rate-distortion-perception trade-off of NeVS. Aiming at this target, we propose the Diffusion-enhanced Neural Codec (DeNC), a plug-and-play module for current NeVS paradigm, to significantly reduce the required bitrates while preserving high perceptual quality of restored videos. Our key design is twofold. First, DeNC improves the encoder's compression efficiency by simultaneously reducing the resolution and color bit-depth of frame referencing. Second, DeNC empowers the decoder with perception-oriented restoration capability by making its diffusion-based restoration process aware of the encoder's compression conditions. Real-world evaluations show that DeNC improves compression ratios with nearly an order of magnitude and achieves much higher restoration quality (e.g., 93+ VMAF and 23% higher MOS) over the latest baselines.
Qihua Zhou, Ruibin Li, Jingcai Guo, Yaodong Huang, Zhenda Xu, Laizhong Cui, Song Guo 0001
AAAI2
2025 RORem: Training a Robust Object Remover with Human-in-the-Loop
abstract
Despite the significant advancements, existing object removal methods struggle with incomplete removal, incorrect content synthesis and blurry synthesized regions, resulting in low success rates. Such issues are mainly caused by the lack of high-quality paired training data, as well as the self-supervised training paradigm adopted in these methods, which forces the model to in-paint the masked regions, leading to ambiguity between synthesizing the masked objects and restoring the background. To address these issues, we propose a semi-supervised learning strategy with human-in-the-loop to create high-quality paired training data, aiming to train a Robust Object Remover (RORem). We first collect 60K training pairs from open-source datasets to train an initial object removal model for generating removal samples, and then utilize human feedback to select a set of high-quality object removal pairs, with which we train a discriminator to automate the following training data generation process. By iterating this process for several rounds, we finally obtain a substantial object removal dataset with over 200K pairs. Fine-tuning the pre-trained stable diffusion model with this dataset, we obtain our RORem, which demonstrates state-of-the-art object removal performance in terms of both reliability and image quality. Particularly, RORem improves the object removal success rate over previous methods by more than 18%. The dataset, source code and trained model are available at https://github.com/leeruibin/RORem.
Ruibin Li, Tao Yang 0042, Song Guo 0001, Lei Zhang 0006
CVPR1
2025 InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset Construction
abstract
Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly consist of low-resolution, short duration, and limited amount of source videos with unsatisfactory editing quality, limiting the performance of trained editing models. In this work, we present a high-quality Instruction-based Video Editing dataset with 1M triplets, namely InsViE-1M. We first curate high-resolution and high-quality source videos and images, then design an effective editing-filtering pipeline to construct high-quality editing triplets for model training. For a source video, we generate multiple edited samples of its first frame with different intensities of classifier-free guidance, which are automatically filtered by GPT-4o with carefully crafted guidelines. The edited first frame is propagated to subsequent frames to produce the edited video, followed by another round of filtering for frame quality and motion evaluation. We also generate and filter a variety of video editing triplets from high-quality images. With the InsViE-1M dataset, we propose a multi-stage learning strategy to train our InsViE model, progressively enhancing its instruction following and editing ability. Extensive experiments demonstrate the advantages of our InsViE-1M dataset and the trained model over state-of-the-art works. Codes are available at \href{https://github.com/langmanbusi/InsViE}{InsViE}.
Yuhui Wu 0001, Liyi Chen 0002, Ruibin Li, Lei Zhang 0006
ICCV3
2025 Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
abstract
Vision-Language Models (VLMs) are essential for multimodal tasks, especially compositional reasoning (CR) tasks, which require distinguishing fine-grained semantic differences between visual and textual embeddings. However, existing methods primarily fine-tune the model by generating text-based hard negative samples, neglecting the importance of image-based negative samples, which results in insufficient training of the visual encoder and ultimately impacts the overall performance of the model. Moreover, negative samples are typically treated uniformly, without considering their difficulty levels, and the alignment of positive samples is insufficient, which leads to challenges in aligning difficult sample pairs. To address these issues, we propose Adaptive Hard Negative Perturbation Learning (AHNPL). AHNPL translates text-based hard negatives into the visual domain to generate semantically disturbed image-based negatives for training the model, thereby enhancing its overall performance. AHNPL also introduces a contrastive learning approach using a multimodal hard negative loss to improve the model's discrimination of hard negatives within each modality and a dynamic margin loss that adjusts the contrastive margin according to sample difficulty to enhance the distinction of challenging sample pairs. Experiments on three public datasets demonstrate that our method effectively boosts VLMs' performance on complex CR tasks. The source code is available at https://github.com/nynu-BDAI/AHNPL.
Ruibin Li
IJCAI2
2024 On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks
abstract
The explosive growth of video traffic on today's Internet promotes the rise of Neural-enhanced Video Streaming (NeVS), which effectively improves the rate-distortion trade-off by employing a cheap neural super-resolution model for quality enhancement on the receiver side. Missing by existing work, we reveal that the NeVS pipeline may suffer from a practical threat, where the crucial codec component (i.e., encoder for compression and decoder for restoration) can trigger adversarial attacks in a man-in-the-middle manner to significantly destroy video recovery performance and finally incurs the malfunction of downstream video perception tasks. In this paper, we are the first attempt to inspect the vulnerability of NeVS and discover a novel adversarial attack, called codec hijacking, where the injected invisible perturbation conspires with the malicious encoding matrix by reorganizing the spatial-temporal bit allocation within the bitstream size budget. Such a zero-day vulnerability makes our attack hard to defend because there is no visual distortion on the recovered videos until the attack happens. More seriously, this attack can be extended to diverse enhancement models, thus exposing a wide range of video perception tasks under threat. Evaluation based on state-of-the-art video codec benchmark illustrates that our attack significantly degrades the recovery performance of NeVS over previous attack methods. The damaged video quality finally leads to obvious malfunction of downstream tasks with over 75% success rate. We hope to arouse public attention on codec hijacking and its defence.
Qihua Zhou, Jingcai Guo, Song Guo 0001, Ruibin Li, Jie Zhang 0076, Zhenda Xu
AAAI4
2024 Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
Ruibin Li, Ruihuang Li, Song Guo 0001, Lei Zhang 0006
ECCV (26)1
2024 ParsNets: A Parsimonious Composition of Orthogonal and Low-Rank Linear Networks for Zero-Shot Learning
Jingcai Guo, Qihua Zhou, Xiaocheng Lu, Ruibin Li, Jie Zhang 0076, Junyang Chen 0001, Xin Xie 0001, Song Guo 0001
IJCAI4
2024 FreePIH: Training-Free Painterly Image Harmonization with Diffusion Model
abstract
This paper provides an efficient training-free painterly image harmonization (PIH) method, dubbed FreePIH, that leverages only a pre-trained diffusion model to achieve state-of-the-art harmonization results. Unlike existing methods that require either training auxiliary networks or fine-tuning a large pre-trained backbone, or both, to harmonize a foreground object with a painterly-style background image, our FreePIH tames the denoising process as a plug-in module for foreground image style transfer. Specifically, we find that the very last few steps of the denoising (i.e., generation) process strongly correspond to the stylistic information of images, and based on this, we propose to augment the latent features of both the foreground and background images with Gaussians for a direct denoising-based harmonization. To guarantee the fidelity of the harmonized image, we make use of latent features to enforce the consistency of the content and stability of the foreground objects in the latent space, and meanwhile, aligning both fore-/back-grounds with the same style. Moreover, to accommodate the generation with more structural and textural details, we further integrate text prompts to attend to the latent features, hence improving the generation quality. Quantitative and qualitative evaluations on COCO and LAION 5B datasets demonstrate that our method can surpass representative baselines by large margins.
Ruibin Li, Jingcai Guo, Qihua Zhou, Song Guo 0001
ACM Multimedia1
2022 ctFS: Replacing File Indexing with Hardware Memory Translation through Contiguous File Allocation for Persistent Memory
Ruibin Li, Xiang Ren 0003, Xu Zhao 0004, Siwei He, Michael Stumm, Ding Yuan 0004
FAST1
2022 Cluster-based content caching driven by popularity prediction
Bosen Jia, Ruibin Li, Chenyang Wang 0001, Chao Qiu, Xiaofei Wang 0001
CCF Trans. High Perform. Comput.2
2022 ctFS: Replacing File Indexing with Hardware Memory Translation through Contiguous File Allocation for Persistent Memory
abstract
Persistent byte-addressable memory (PM) is poised to become prevalent in future computer systems. PMs are significantly faster than disk storage, and accesses to PMs are governed by the Memory Management Unit (MMU) just as accesses with volatile RAM. These unique characteristics shift the bottleneck from I/O to operations such as block address lookup—for example, in write workloads, up to 45% of the overhead in ext4-DAX is due to building and searching extent trees to translate file offsets to addresses on persistent memory. We propose a novel contiguous file system, ctFS, that eliminates most of the overhead associated with indexing structures such as extent trees in the file system. ctFS represents each file as a contiguous region of virtual memory, hence a lookup from the file offset to the address is simply an offset operation, which can be efficiently performed by the hardware MMU at a fraction of the cost of software-maintained indexes. Evaluating ctFS on real-world workloads such as LevelDB shows it outperforms ext4-DAX and SplitFS by 3.6× and 1.8×, respectively.
Ruibin Li, Xiang Ren 0003, Xu Zhao 0004, Siwei He, Michael Stumm, Ding Yuan 0004
ACM Trans. Storage1
2021 Neighboring-Aware Caching in Heterogeneous Edge Networks by Actor-Attention-Critic Learning
abstract
With the development of network technology and the surge in demand, the speed and throughput of data and applications are leading to the skyrocketing increase in traffic. The communication and collaboration between heterogeneous edge servers are indispensable. In this scenario with heterogeneous edges, there is a common understanding on the fact that an effective edge caching algorithm could play the role of enabler to reduce the network resource consumption and content fetch delay. However, most of the existing studies on multi-agent caching methods focus more on the overall situation, while ignoring the mutual influence between different agents. In this context, we model the edge caching content replacement problem as a Markov process and deploy attention mechanism based on the Actor-Attention-Critic algorithm to realize a neighboring-aware edge caching (NAEC) strategy. The proposed method makes full use of the communication between base stations to exchange neighboring information, so that we can reduce the pressure on the backbone and further improve user satisfaction. The simulation results have verified the feasibility and effectiveness of the proposed algorithm.
Ruibin Li, Chenyang Wang 0001, Xiaofei Wang 0001, Victor C. M. Leung
ICC2
2021 Anchored User Selection for Traffic Offloading Optimization in D2D-Aided Mobile-Edge Computing
abstract
Recently, integrated with the advanced communication technologies (e.g., 5G) and artificial intelligence (AI), mobile-edge intelligence (MEI) is regarded as the promising method to deal with the emerging challenges. Specifically, Device-to-Device (D2D) communications have been put forward to reduce the traffic pressure while extending cellular network capacity. However, the stability of the social network is important for the design of efficient and reliable traffic offloading strategy, which is often absent from the related work. Besides, most existing studies merely model the relation between a node pair as a binary or continuous value, neglecting the rich information between users. Moreover, many traditional models are conducted based on small-scale data sets or online Internet services, severely confining their applications in the D2D scenario. Thus, it is necessary to understand the network structure and select the key users to address the aforementioned challenges. In this article, we first propose a network representation model, named MPPT, to regard the multidimensional relations as a probability in a third-order (3-D) tensor space. Then, a mobile D2D social community is derived by integrating an edge base station (BS) and the nearby D2D users, and develop an anchored user selection algorithm to maintain the stability of multiple D2D social communities by choosing and retaining critical users adaptively under the limited network resources. Finally, we devise a probability-based onion layers anchored$(k,r)$-core (P-OLAK) algorithm to identify the anchor users. The large-scale data sets-based experimental results show the superiorities of the proposed methods.
Chenyang Wang 0001, Ruibin Li, Zheng Di, Chao Qiu, Xiaofei Wang 0001
IEEE Internet Things J.2
2021 SimEdgeIntel: A open-source simulation platform for resource management in edge intelligence
Chenyang Wang 0001, Ruibin Li, Chao Qiu, Xiaofei Wang 0001
J. Syst. Archit.2
2021 Attention-Weighted Federated Deep Reinforcement Learning for Device-to-Device Assisted Heterogeneous Collaborative Edge Caching
abstract
In order to meet the growing demands for multimedia service access and release the pressure of the core network, edge caching and device-to-device (D2D) communication have been regarded as two promising techniques in next generation mobile networks and beyond. However, most existing related studies lack consideration of effective cooperation and adaptability to the dynamic network environments. In this article, based on the flexible trilateral cooperation among user equipment, edge base stations and a cloud server, we propose a D2D-assisted heterogeneous collaborative edge caching framework by jointly optimizing the node selection and cache replacement in mobile networks. We formulate the joint optimization problem as a Markov decision process, and use a deep Q-learning network to solve the long-term mixed integer linear programming problem. We further design an attention-weighted federated deep reinforcement learning (AWFDRL) model that uses federated learning to improve the training efficiency of the Q-learning network by considering the limited computing and storage capacity, and incorporates an attention mechanism to optimize the aggregation weights to avoid the imbalance of local model quality. We prove the convergence of the corresponding algorithm, and present simulation results to show the effectiveness of the proposed AWFDRL framework in reducing average delay of content access, improving hit rate and offloading traffic.
Xiaofei Wang 0001, Ruibin Li, Chenyang Wang 0001, Xiuhua Li 0001, Tarik Taleb, Victor C. M. Leung
IEEE J. Sel. Areas Commun.2
2020 Edge Caching Replacement Optimization for D2D Wireless Networks via Weighted Distributed DQN
abstract
Duplicated download has been a big problem that affects the users' quality of service/experience (QoS/QoE) of current mobile networks. Edge caching and Device-to-Device communication are two promising technologies to release the pressure of repeated traffic downloading from the cloud. There are many researches about the edge caching policy. However, these researches have some limitations in the real scenarios. Traditional methods are lacking the self-adaptive ability in the dynamic environment and privacy issues will occur in centralized learning methods. In this paper, based on the virtue of Deep Q-Network (DQN), we propose a weighted distributed DQN model (WDDQN) to solve the cache replacement problem. Our model enables collaboratively to learn a shared predictive model. Trace-driven simulation results show that our proposed model outperforms some classical and state-of-the-art schemes.
Ruibin Li, Chenyang Wang 0001, Xiaofei Wang 0001, Victor C. M. Leung, Xiuhua Li 0001, Tarik Taleb
WCNC1