Ruoxi Cheng

dblp:300/6824 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0001-3261-642XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 100%
Artificial intelligence
1 paper
Vision and language · 77% Efficient and distributed learning · 23%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › image captioning
image caption evaluation
1.012026
L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training · IEEE Trans. Multim. 2026
Computer vision › Vision and language
image captioning
1.012026
L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training · IEEE Trans. Multim. 2026
Security and privacy of machine learning
adversarial attack
0.912025
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization · EMNLP 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization · EMNLP 2025
Security and privacy of machine learning
toxic content generation
0.912025
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312026
L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training · IEEE Trans. Multim. 2026
Machine learning › Efficient and distributed learning
model compression
0.312026
L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training · IEEE Trans. Multim. 2026

Methods — techniques the papers use, named apart from their topics

weight multiplexing · 1.0similarity regulator loss · 1.0matrix decomposition · 1.0CLIP · 1.0multimodal interaction · 0.9black-box attack · 0.9
YearPublicationVenuePosition
2026 L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training
abstract
We propose a novel embedding-based captioning metric termed asL-CLIPScorethat can be used for efficiently evaluating caption quality and training captioning model. L-CLIPScore is calculated from a lightweight CLIP (L-CLIP), which is a dual-encoder architecture compressed and distilled from CLIP. To compress, we apply two powerful techniques which are weight multiplexing and matrix decomposition for reducing the parameters of encoders and word embedding matrix, respectively. To distill, we design a novel multi-modal Similarity Regulator (SR) loss to transfer more vision-language alignment knowledge. Specifically, SR loss amplifies the multi-modal embedding similarity if the given image-text pair is matched and diminishes the similarity if the pair is non-matched. By compressing and distilling by this novel SR loss, our L-CLIP achieves comparable multi-modal alignment ability to the original CLIP while it requires fewer computation resources and running time. We carry out exhaustive experiments to validate the efficiency and effectiveness of L-CLIPScore when using it as the judge to evaluate caption quality. We also discover that when using L-CLIPScore as the supervisor to train the captioning model, it should be mixed up by an n-gram-based metric and meanwhile analyze why using L-CLIPScore only will cause fail training
Yingzhe Peng, Xu Yang 0021, Ruoxi Cheng, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002
IEEE Trans. Multim.4
2025 SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
abstract
Traditional methods for evaluating the robustness of large language models (LLMs) often rely on standardized benchmarks, which can escalate costs and limit evaluations across varied domains. This paper introduces a novel framework designed to autonomously evaluate the robustness of LLMs by incorporating refined adversarial prompts and domain-constrained knowledge guidelines in the form of knowledge graphs. Our method systematically generates descriptive sentences from domain-constrained knowledge graph triplets to formulate adversarial prompts, enhancing the relevance and challenge of the evaluation. These prompts, generated by the LLM itself and tailored to evaluate its own robustness, undergo a rigorous filtering and refinement process, ensuring that only those with high textual fluency and semantic fidelity are used. This self-evaluation mechanism allows the LLM to evaluate its robustness without the need for external benchmarks. We assess the effectiveness of our framework through extensive testing on both proprietary models like ChatGPT and open-source models such as Llama-3.1, Phi-3, and Mistral. Results confirm that our approach not only reduces dependency on conventional data but also provides a targeted and efficient means of evaluating LLM robustness in constrained domains.
Aihua Pei, Zehua Yang, Shunan Zhu, Ruoxi Cheng, Ju Jia
COLING4
2025 PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
abstract
Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Zhiqiang Wang, Xiaojun Jia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Xiaojun Jia
EMNLP1
2025 AGR: Age Group fairness Reward for Bias Mitigation in LLMs
abstract
LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little explored. The scarcity of instruction-tuning and preference datasets for age bias hampers its detection and measurement, and existing fine-tuning methods seldom address age-related fairness. In this paper, we construct age bias preference datasets and instruction-tuning datasets for RLHF. We introduce AGR, an age fairness reward to reduce differences in the response quality of LLMs across different age groups. Extensive experiments demonstrate that this reward significantly improves response accuracy and reduces performance disparities across age groups. Our source code and datasets are available at the anonymous link.
Shuirong Cao, Ruoxi Cheng, Zhiqiang Wang 0006
ICASSP2
2025 Talk the Talk, Debate the Bias: LLM Alignment via Role-Play Rumble
Ruoxi Cheng, Shaowei Yuan, Yizhong Ding
ICIC (24)1
2025 Introspective Reward Modeling via Inverse Reinforcement Learning for LLM Alignment
Ruoxi Cheng, Shaowei Yuan, Yizhong Ding
ICIC (23)2
2025 Speaker Inference Detection Using Only Text
Ruoxi Cheng, Yizhong Ding, Shaowei Yuan
ICICS (3)1
2025 Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
abstract
Audio can disclose PII, particularly when combined with related text data. Therefore, it is essential to develop tools to detect privacy leakage in Contrastive Language-Audio Pretraining(CLAP). Existing MIAs need audio as input, risking exposure of voiceprint and requiring costly shadow models. We first propose PRMID, a membership inference detector based probability ranking given by CLAP, which does not require training shadow models but still requires both audio and text of the individual as input. To address these limitations, we then propose USMID, a textual unimodal speaker-level membership inference detector, querying the target model using only text data. We randomly generate textual gibberish that are clearly not in training dataset. Then we extract feature vectors from these texts using the CLAP model and train a set of anomaly detectors on them. During inference, the feature vector of each test text is input into the anomaly detector to determine if the speaker is in the training set (anomalous) or not (normal). If available, USMID can further enhance detection by integrating real audio of the tested speaker. Extensive experiments on various CLAP model architectures and datasets demonstrate that USMID outperforms baseline methods using only text data.
Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Zhiqiang Wang 0006
ICMR1