Xiaomeng Hu

dblp:319/7072 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 45% Trustworthy machine learning · 28% Transfer learning and domain adaptation · 9%
Network and information security
3 papers
Security and privacy of machine learning · 89% Digital forensics and information hiding · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning › large language model safety
jailbreak defense
1.622025
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models · AAAI 2025
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model safety
decoding-time safety alignment
0.912025
CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention · NeurIPS 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency · EMNLP 2025
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
jailbreak defense
0.912025
CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention · NeurIPS 2025
Machine learning › Reinforcement learning
reward design
0.912025
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization · EMNLP 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.912025
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency · EMNLP 2025
Security and privacy of machine learning › large language model security
adversarial attacks on language models
0.912025
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models · AAAI 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models · AAAI 2025
Natural language and speech › Language models and text generation
hallucination detection
0.812024
Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection · EMNLP 2024
Security and privacy of machine learning › large language model safety
jailbreak detection
0.812024
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes · NeurIPS 2024
Security and privacy of machine learning
large language model safety
0.812024
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes · NeurIPS 2024
Security and privacy of machine learning
adversarial learning
0.712023
RADAR: Robust AI-Text Detection via Adversarial Learning · NeurIPS 2023
Digital forensics and information hiding › synthetic media detection
machine-generated text detection
0.712023
RADAR: Robust AI-Text Detection via Adversarial Learning · NeurIPS 2023
Natural language and speech › Language models and text generation
pre-trained language model
0.612022
P3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning · SIGIR 2022
Computer vision › Vision and language › vision-language model
prompt learning
0.612022
P3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning · SIGIR 2022
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.612022
P3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning · SIGIR 2022
Information retrieval
retrieval models
0.612022
P3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning · SIGIR 2022
Machine learning › Representation and self-supervised learning
cycle consistency
0.312025
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency · EMNLP 2025
Natural language and speech › Language models and text generation
decoding
0.312025
CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

soft removal · 0.9self-training · 0.9rollback mechanism · 0.9reward hybridization · 0.9reinforcement learning · 0.9introspection-based intervention · 0.9instruction tuning · 0.9guard model · 0.9gradient-based token attribution · 0.9cycle consistency · 0.9affirmation loss · 0.9refusal loss · 0.8gradient analysis · 0.8first-order gradient · 0.8detection threshold · 0.8autoregressive language model · 0.8paraphrasing · 0.7large language model · 0.7
YearPublicationVenuePosition
2025 Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
abstract
Large Language Models (LLMs) are increasingly being integrated into services such as ChatGPT to provide responses to user queries. To mitigate potential harm and prevent misuse, there have been concerted efforts to align the LLMs with human values and legal compliance by incorporating various techniques, such as Reinforcement Learning from Human Feedback (RLHF), into the training of the LLMs. However, recent research has exposed that even aligned LLMs are susceptible to adversarial manipulations known as Jailbreak Attacks. To address this challenge, this paper proposes a method called Token Highlighter to inspect and mitigate the potential jailbreak threats in the user query. Token Highlighter introduced a concept called Affirmation Loss to measure the LLM's willingness to answer the user query. It then uses the gradient of Affirmation Loss for each token in the user query to locate the jailbreak-critical tokens. Further, Token Highlighter exploits our proposed Soft Removal technique to mitigate the jailbreak effects of critical tokens via shrinking their token embeddings. Experimental results on two aligned LLMs (LLaMA-2 and Vicuna-V1.5) demonstrate that the proposed method can effectively defend against a variety of Jailbreak Attacks while maintaining competent performance on benign questions of the AlpacaEval benchmark. In addition, Token Highlighter is a cost-effective and interpretable defense because it only needs to query the protected LLM once to compute the Affirmation Loss and can highlight the critical tokens upon refusal.
Xiaomeng Hu, Tsung-Yi Ho
AAAI1
2025 CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
abstract
Zhanming Shen, Hao Chen, Yulei Tang, Shaolin Zhu, Wentao Ye, Xiaomeng Hu, Haobo Wang, Gang Chen, Junbo Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhanming Shen, Hao Chen 0081, Yulei Tang, Shaolin Zhu, Wentao Ye, Xiaomeng Hu, Haobo Wang 0001, Gang Chen 0001, Junbo Zhao 0002
EMNLP6
2025 LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
abstract
Qi Zhang, Shouqing Yang, Lirong Gao, Hao Chen, Xiaomeng Hu, Jinglei Chen, Jiexiang Wang, Sheng Guo, Bo Zheng, Haobo Wang, Junbo Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Qi Zhang 0077, Shouqing Yang, Lirong Gao, Hao Chen 0081, Xiaomeng Hu, Jinglei Chen, Jiexiang Wang, Haobo Wang 0001, Junbo Zhao 0002
EMNLP5
2025 CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
abstract
As large language models (LLMs) are increasingly deployed in real-world applications, ensuring the safety of their outputs during decoding has become a critical challenge. However, existing decoding-time interventions, such as Contrastive Decoding, often force a severe trade-off between safety and response quality. In this work, we propose **CARE**, a novel framework for decoding-time safety alignment that integrates three key components: (1) a guard model for real-time safety monitoring, enabling detection of potentially unsafe content; (2) a rollback mechanism with a token buffer to correct unsafe outputs efficiently at an earlier stage without disrupting the user experience; and (3) a novel introspection-based intervention strategy, where the model generates self-reflective critiques of its previous outputs and incorporates these reflections into the context to guide subsequent decoding steps. The framework achieves a superior safety-quality trade-off by using its guard model for precise interventions, its rollback mechanism for timely corrections, and our novel introspection method for effective self-correction. Experimental results demonstrate that our framework achieves a superior balance of safety, quality, and efficiency, attaining a **low harmful response rate** and **minimal disruption to the user experience** while **maintaining high response quality**.
Xiaomeng Hu, Fei Huang 0002, Chenhan Yuan, Junyang Lin, Tsung-Yi Ho
NeurIPS1
2024 Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection
abstract
In recent years, large language models (LLMs) have achieved remarkable success in the field of natural language generation.Compared to previous small-scale models, they are capable of generating fluent output based on the provided prefix or prompt.However, one critical challenge -the hallucination problem -remains to be resolved.Generally, the community refers to the undetected hallucination scenario where the LLMs generate text unrelated to the input text or facts.In this study, we intend to model the distributional distance between the regular conditional output and the unconditional output, which is generated without a given input text.Based upon Taylor Expansion for this distance at the output probability space, our approach manages to leverage the embedding and first-order gradient information.The resulting approach is plug-and-play that can be easily adapted to any autoregressive LLM.On the hallucination benchmarks HADES and other datasets, our approach achieves state-of-the-art performance.
Xiaomeng Hu, Yiming Zhang 0023, Ru Peng, Chenwei Wu 0010, Gang Chen 0001, Junbo Zhao 0002
EMNLP1
2024 Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
abstract
Large Language Models (LLMs) are becoming a prominent generative AI tool, where the user enters a query and the LLM generates an answer. To reduce harm and misuse, efforts have been made to align these LLMs to human values using advanced training techniques such as Reinforcement Learning from Human Feedback (RLHF). However, recent studies have highlighted the vulnerability of LLMs to adversarial jailbreak attempts aiming at subverting the embedded safety guardrails. To address this challenge, this paper defines and investigates the **Refusal Loss** of LLMs and then proposes a method called **Gradient Cuff** to detect jailbreak attempts. Gradient Cuff exploits the unique properties observed in the refusal loss landscape, including functional values and its smoothness, to design an effective two-step detection strategy. Experimental results on two aligned LLMs (LLaMA-2-7B-Chat and Vicuna-7B-V1.5) and six types of jailbreak attacks (GCG, AutoDAN, PAIR, TAP, Base64, and LRL) show that Gradient Cuff can significantly improve the LLM's rejection capability for malicious jailbreak queries, while maintaining the model's performance for benign user queries by adjusting the detection threshold.
Xiaomeng Hu, Tsung-Yi Ho
NeurIPS1
2023 RADAR: Robust AI-Text Detection via Adversarial Learning
abstract
Recent advances in large language models (LLMs) and the intensifying popularity of ChatGPT-like applications have blurred the boundary of high-quality text generation between humans and machines. However, in addition to the anticipated revolutionary changes to our technology and society, the difficulty of distinguishing LLM-generated texts (AI-text) from human-generated texts poses new challenges of misuse and fairness, such as fake content generation, plagiarism, and false accusations of innocent writers. While existing works show that current AI-text detectors are not robust to LLM-based paraphrasing, this paper aims to bridge this gap by proposing a new framework called RADAR, which jointly trains a $\underline{r}$obust $\underline{A}$I-text $\underline{d}$etector via $\underline{a}$dversarial lea$\underline{r}$ning. RADAR is based on adversarial training of a paraphraser and a detector. The paraphraser's goal is to generate realistic content to evade AI-text detection. RADAR uses the feedback from the detector to update the paraphraser, and vice versa. Evaluated with 8 different LLMs (Pythia, Dolly 2.0, Palmyra, Camel, GPT-J, Dolly 1.0, LLaMA, and Vicuna) across 4 datasets, experimental results show that RADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in place. We also identify the strong transferability of RADAR from instruction-tuned LLMs to other LLMs, and evaluate the improved capability of RADAR via GPT-3.5-Turbo.
Xiaomeng Hu, Tsung-Yi Ho
NeurIPS1
2022 Multi-Objective Geometric Optimization of A Multi-Link Manipulator Using Parameterized Design Method
abstract
The performance of a robot is closely related to its structure. From the initial design of link lengths to structural optimization, it is still the research hotspot in recent years. To make the manipulator lightweight and ensure its working range and flexibility, researchers have proposed many optimization methods, most of which are for specific working scenarios, requirements, and robot structures, therefore their generality is limited. The optimization of the manipulator should be a comprehensive method. That is, we should pay attention to the joint configuration and each link length at the beginning of the design. Particularly, the geometric parameters of each link, which not only affect the range of the workspace but also have a direct impact on the working space, working efficiency, and flexibility of the manipulator. In this paper, a generalized optimization framework is proposed for multi-link manipulators. Starting from the optimization of manipulator link lengths, firstly, the geometry of the manipulator and workspace is parameterized; then the performance indicators are established; lastly, the geometric size of the manipulator is optimized according to the workspace limits and task requirements. Besides, we verified its feasibility and generality by applying this method to different TBM scenarios.
Xiaomeng Hu, Weiwei Wan, Liang Du 0002, Jianjun Yuan 0003, Shugen Ma
IROS1
2022 P3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning
abstract
Compared to other language tasks, applying pre-trained language models (PLMs) for search ranking often requires more nuances and training signals. In this paper, we identify and study the two mismatches between pre-training and ranking fine-tuning: the training schema gap regarding the differences in training objectives and model architectures, and the task knowledge gap considering the discrepancy between the knowledge needed in ranking and that learned during pre-training. To mitigate these gaps, we propose Pre-trained, Prompt-learned and Pre-finetuned Neural Ranker (P3 Ranker). P3 Ranker leverages prompt-based learning to convert the ranking task into a pre-training like schema and uses pre-finetuning to initialize the model on intermediate supervised tasks. Experiments on MS MARCO and Robust04 show the superior performances of P3 Ranker in few-shot ranking. Analyses reveal that P3 Ranker is able to better accustom to the ranking task through prompt-based learning and retrieve necessary ranking-oriented knowledge gleaned in pre-finetuning, resulting in data-efficient PLM adaptation. Our code is available at https://github.com/NEUIR/P3Ranker.
Xiaomeng Hu, Shi Yu 0001, Chenyan Xiong, Zhenghao Liu 0001, Zhiyuan Liu 0001, Ge Yu 0001
SIGIR1