Ruoxi Chen

dblp:246/3090 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Security and privacy · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
abstract
Personalized diffusion models (PDMs) have become prominent for adapting pre-trained text-to-image models to generate images of specific subjects using minimal training data. However, PDMs are susceptible to minor adversarial perturbations, leading to significant degradation when fine-tuned on corrupted datasets. These vulnerabilities are exploited to create protective perturbations that prevent unauthorized image generation. Existing purification methods attempt to red-team the protective perturbation to break the protection but often over-purify images, resulting in information loss. In this work, we conduct an in-depth analysis of the fine-tuning process of PDMs through the lens of shortcut learning. We hypothesize and empirically demonstrate that adversarial perturbations induce a latent-space misalignment between images and their text prompts in the CLIP embedding space. This misalignment causes the model to erroneously associate noisy patterns with unique identifiers during fine-tuning, resulting in poor generalization. Based on these insights, we propose a systematic red-teaming framework that includes data purification and contrastive decoupling learning. We first employ off-the-shelf image restoration techniques to realign images with their original semantic content in latent space. Then, we introduce contrastive decoupling learning with noise tokens to decouple the learning of personalized concepts from spurious noise patterns. Our study not only uncovers shortcut learning vulnerabilities in PDMs but also provides a thorough evaluation framework for developing stronger protection. Our extensive evaluation demonstrates its advantages over existing purification methods and its robustness against adaptive perturbations.
Yixin Liu 0002, Ruoxi Chen, Lichao Sun 0001
KDD (1)2
2026 BGF-DR: bidirectional greybox fuzzing for DNS resolver vulnerability discovery
Ruoxi Chen, Hongxin Su, Tiantian Zhu 0001
Comput. Secur.3
2025 Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
abstract
Many real-world user queries (e.g. *"How do to make egg fried rice?"*) could benefit from systems capable of generating responses with both textual steps with accompanying images, similar to a cookbook. Models designed to generate interleaved text and images face challenges in ensuring consistency within and across these modalities. To address these challenges, we present ISG, a comprehensive evaluation framework for interleaved text-and-image generation. ISG leverages a scene graph structure to capture relationships between text and image blocks, evaluating responses on four levels of granularity: holistic, structural, block-level, and image-specific. This multi-tiered evaluation allows for a nuanced assessment of consistency, coherence, and accuracy, and provides interpretable question-answer feedback. In conjunction with ISG, we introduce a benchmark, ISG-Bench, encompassing 1,150 samples across 8 categories and 21 subcategories. This benchmark dataset includes complex language-vision dependencies and golden answers to evaluate models effectively on vision-centric tasks such as style transfer, a challenging area for current models. Using ISG-Bench, we demonstrate that recent unified vision-language models perform poorly on generating interleaved content. While compositional approaches that combine separate language and image models show a 111% improvement over unified models at the holistic level, their performance remains suboptimal at both block and image levels. To facilitate future work, we develop ISG-Agent, a baseline agent employing a *"plan-execute-refine"* pipeline to invoke tools, achieving a 122% performance improvement.
Dongping Chen, Ruoxi Chen, Shu Pu, Yanru Wu, Caixi Chen, Benlin Liu, Yue Huang 0001, Yao Wan 0001, Pan Zhou 0001, Ranjay Krishna
ICLR2
2025 MultiRef: Controllable Image Generation with Multiple Visual References
abstract
Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs - either text prompts or individual reference images. In this paper, we focus on the task of controllable image generation using multiple visual references. We introduce MultiRef-bench, a rigorous evaluation framework comprising 990 synthetic and 1,000 real-world samples that require incorporating visual content from multiple reference images. The synthetic samples are synthetically generated through our data engine RefBlend, with 10 reference types and 33 reference combinations. Based on RefBlend, we further construct a dataset MultiRef containing 38k high-quality images to facilitate further research. Our experiments across three interleaved image-text models (i.e., OmniGen, ACE, and Show-o) and six agentic frameworks (e.g., ChatDiT and LLM + SD) reveal that even state-of-the-art systems struggle with multi-reference conditioning, with the best model OmniGen achieving only 66.6% in synthetic samples and 79.0% in real-world cases on average compared to the golden answer. These findings provide valuable directions for developing more flexible and human-like creative tools that can effectively integrate multiple sources of visual inspiration. The dataset is publicly available at: https://multiref.github.io/.
Ruoxi Chen, Dongping Chen, Siyuan Wu 0001, Shiyun Lang, Peter Sushko, Gaoyang Jiang, Yao Wan 0001, Ranjay Krishna
ACM Multimedia1
2025 Fight Perturbations With Perturbations: Defending Adversarial Attacks via Neuron Influence
abstract
The vulnerabilities of deep learning models towards adversarial attacks have attracted increasing attention, especially when models are deployed in security-critical domains. Numerous defense methods, including reactive and proactive ones, have been proposed for model robustness improvement. Reactive defenses, such as conducting transformations to remove perturbations, usually fail to handle large perturbations. The proactive defenses that involve retraining, suffer from the attack dependency and high computation cost. In this article, we consider defense methods from the general effect of adversarial attacks that take on neurons inside the model. We introduce the concept of neuron influence, which can quantitatively measure neurons’ contribution to correct classification. Then, we observe that almost all attacks fool the model by suppressing neurons with larger influence and enhancing those with smaller influence. Based on this, we proposeNeuron-level Inverse Perturbation(NIP), a novel defense against general adversarial attacks. It calculates neuron influence from benign examples and then modifies input examples by generating inverse perturbations that can in turn strengthen neurons with larger influence and weaken those with smaller influence. Extensive experiments on benchmark datasets and models show that NIP outperforms the state-of-the-art methods in terms of i)effective- it shows better defense success rate ($\sim \!\!\times 1.45$) against 13 adversarial attacks; ii)elastic- it maintains better defense ($\sim \!\!\times 3.4$in the worst case) on large perturbations; iii)efficient- it runs with only$\sim \!\!1/6$time cost; iv)extensible- it can be applied to speaker recognition models and Baidu online image platforms. We further evaluate NIP against potential adaptive attacks and provide interpretable analysis for its effectiveness.
Ruoxi Chen, Haibo Jin, Haibin Zheng, Jinyin Chen, Zhenguang Liu
IEEE Trans. Dependable Secur. Comput.1
2024 EditShield: Protecting Unauthorized Image Editing by Instruction-Guided Diffusion Models
Ruoxi Chen, Haibo Jin, Yixin Liu 0002, Jinyin Chen, Haohan Wang, Lichao Sun 0001
ECCV (63)1
2024 CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
Haibo Jin, Ruoxi Chen, Jinyin Chen, Haibin Zheng, Haohan Wang
ECCV (47)2
2024 MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
abstract
Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to the absence multimodal benchmarks that align with human preferences. Drawing inspiration from the concept of LLM-as-a-Judge within LLMs, this paper introduces a novel benchmark, termed MLLM-as-a-Judge, to assess the ability of MLLMs in assisting judges across diverse modalities, encompassing three distinct tasks: Scoring Evaluation, Pair Comparison, and Batch Ranking. Our study reveals that, while MLLMs demonstrate remarkable human-like discernment in Pair Comparisons, there is a significant divergence from human preferences in Scoring Evaluation and Batch Ranking tasks. Furthermore, a closer examination reveals persistent challenges in the evaluative capacities of LLMs, including diverse biases, hallucinatory responses, and inconsistencies in judgment, even in advanced models such as GPT-4V. These findings emphasize the pressing need for enhancements and further research efforts to be undertaken before regarding MLLMs as fully reliable evaluators. In light of this, we advocate for additional efforts dedicated to supporting the continuous development within the domain of MLLM functioning as judges. The code and dataset are publicly available at our project homepage: https://mllm-judge.github.io/.
Dongping Chen, Ruoxi Chen, Yaochen Wang 0001, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan 0001, Pan Zhou 0001, Lichao Sun 0001
ICML2
2024 An Online Rcm Adjusting System for Robot-Assisted Retinal Surgeries
abstract
In robot-assisted retinal surgery, a Remote Center of Motion (Rcm) allows the surgical instrument to rotate around a distal fixed point without any lateral translations. The Rcm point should be perfectly aligned inside the trocar. Otherwise, unexpected tool translations at the expected remote center will enlarge the force applied to the trocar and consequently result in post-operative complications. Due to the narrow size of the trocar and the lack of real-time detection equipment, the Rcm point is hard to be perfectly located inside the trocar. Even if the Rcm is perfectly aligned, the movement of the tissue around the eyeball could make it inappropriate again. In this paper, inspired by the control strategy of surgeons, an online Rcm adjusting strategy is proposed. Instead of only using one fixed Rcm point, to restrict the force between the surgical tool and the trocar, the proposed strategy adjusts the position of the Rcm point during the motion. The results show our approach significantly reduces the force between the robot end-effector and surgical port by 64.2%. In addition, the results also demonstrate that our approach complies the Rcm trajectories without deforming or spoiling the working space, which is significantly important for obeying surgeon’s instructions in practice.
Ting Wang 0028, Huanqi Ni, Yanlin Li 0006, Ruoxi Chen, M. Ali Nasseri, Haotian Lin 0001, Kai Huang 0001
IROS5
2024 Like teacher, like pupil: Transferring backdoors via feature-based knowledge distillation
Jinyin Chen, Zhiqi Cao, Ruoxi Chen, Haibin Zheng, Qi Xuan 0001, Xing Yang 0004
Comput. Secur.3
2024 AdvCheck: Characterizing adversarial examples via local gradient checking
Ruoxi Chen, Haibo Jin, Jinyin Chen, Haibin Zheng, Shilian Zheng, Xiaoniu Yang, Xing Yang 0004
Comput. Secur.1
2023 Unleashing the Potential of Adjacent Snippets for Weakly-supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) intends to detect action instances with only weak supervision, e.g., video-level labels. The current de facto pipeline locates action instances by thresholding and grouping continuous high-score regions on temporal class activation sequences. In this route, the capacity of the model to recognize the relationships between adjacent snippets is of vital importance which determines the quality of the action boundaries. However, it is error-prone since the variations between adjacent snippets are typically subtle, and unfortunately this is overlooked in the literature. To tackle the issue, we propose a novel WTAL approach named Convex Combination Consistency between Neighbors (C3BN). C3BN consists of two key ingredients: a micro data augmentation strategy that increases the diversity in-between adjacent snippets by convex combination of adjacent snippets, and a macro-micro consistency regularization that enforces the model to be invariant to the transformations w.r.t. video semantics, snippet predictions, and snippet representations. Consequently, fine-grained patterns in-between adjacent snippets are enforced to be explored, thereby resulting in a more robust action boundary localization. Experimental results demonstrate the effectiveness of C3BN on top of various baselines for WTAL with video-level and point-level supervision. Code is at: https://github.com/canbaoburen/C3BN.
Qinying Liu, Zilei Wang, Ruoxi Chen
ICME3
2023 Excitement surfeited turns to errors: Deep learning testing framework based on excitable neurons
Haibo Jin, Ruoxi Chen, Haibin Zheng, Jinyin Chen, Yao Cheng 0002, Yue Yu 0001, Tieming Chen, Xianglong Liu 0001
Inf. Sci.2
2022 Salient feature extractor for adversarial defense on deep neural networks
Ruoxi Chen, Jinyin Chen, Haibin Zheng, Qi Xuan 0001, Zhaoyan Ming, Wenrong Jiang
Inf. Sci.1
2021 FineFool: A novel DNN object contour attack on image recognition based on the attention perturbation adversarial technique
Jinyin Chen, Haibin Zheng, Hui Xiong 0005, Ruoxi Chen, Tianyu Du, Zhen Hong, Shouling Ji
Comput. Secur.4
2020 RCA-SOC: A novel adversarial defense by refocusing on critical areas and strengthening object contours
Jinyin Chen, Haibin Zheng, Ruoxi Chen, Hui Xiong 0005
Comput. Secur.3