VLDB 2026 Research / reviewers in the wild / expert
Ming Xu 0006
dblp:43/3362-6
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-1061-819XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EmoRAG: Evaluating RAG Robustness to Symbolic PerturbationsabstractRetrieval-Augmented Generation (RAG) systems are increasingly central to robust AI, enhancing large language model (LLM) faithfulness by incorporating external knowledge. However, our study unveils a critical, overlooked vulnerability: their profound susceptibility to subtle symbolic perturbations, particularly through near-imperceptible emotional icons (e.g., "(@_@)") that can catastrophically mislead retrieval, termed EmoRAG. We demonstrate that injecting a single emoticon into a query makes it nearly 100% likely to retrieve semantically unrelated texts, which contain a matching emoticon. Our extensive experiment across general question-answering and code domains, using a range of state-of-the-art retrievers and generators, reveals three key findings: (I) Single-Emoticon Disaster: Minimal emoticon injections cause maximal disruptions, with a single emoticon almost 100% dominating RAG output. (II) Positional Sensitivity: Placing an emoticon at the beginning of a query can cause severe perturbation, with F1-Scores exceeding 0.92 across all datasets. (III) Parameter-Scale Vulnerability: Counterintuitively, models with larger parameters exhibit greater vulnerability to the interference. We provide an in-depth analysis to uncover the underlying mechanisms of these phenomena. Furthermore, we raise a critical concern regarding the robustness assumption of current RAG systems, envisioning a threat scenario where an adversary exploits this vulnerability to manipulate the RAG system. We evaluate standard defenses and find them insufficient against EmoRAG. To address this, we propose targeted defenses, analyzing their strengths and limitations in mitigating emoticon-based perturbations. Finally, we outline future directions for building robust RAG systems. Xinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu 0006, Xuanwang Zhang, Yidong Wang 0003, Xiaojun Jia, Kun Wang 0056, Qingsong Wen, XiaoFeng Wang 0001, Wei Dong 0007 |
KDD (1) | 4 |
| 2026 | ARuleCon: Agentic Security Rule ConversionabstractThe real-time demand for web security makes Security Information and Event Management (SIEM) platforms and their applied security rule an integral part of the intrusion detection life-cycle. However, the heterogeneity of vendor-specific rules (e.g., Splunk SPL, Microsoft KQL, IBM AQL, Google YARA-L, and RSA ESA) makes cross-platform rule reuse extremely difficult, requiring deep domain knowledge for reliable conversion. As a result, an autonomous and accurate rule conversion framework can significantly lead to effort savings, preserving the value of existing rules. In this paper, we propose ARuleCon, an agentic SIEM-rule conversion approach. Using ARuleCon, the security professionals do not need to distill the source rules' logic and re-map it to target vendors, instead, they provide the source rules, the documentation of the target rules and ARuleCon can purposely convert to the target vendors without more intervention. To achieve this, ARuleCon is equipped with intermediate representation (IR) that aligns core detection logic into vendor-neutral layer, agentic RAG pipeline that retrieves authoritative official vendor documentation to address the convension/schema mismatches, and Python-based consistency check that running both source and target rules in controlled test environments to mitigate subtle semantic drifts. We present a comprehensive evaluation of ARuleCon ranging from textual alignment between the source and target rules, and the execution success of target rules, showcasing ARuleCon can convert rules with higher fidelity, outperforming the baseline LLM models by 15% averagely. Finally, we perform a case study and interview with our industry collaborators 1, which showcases that ARuleCon can significantly save the expert's time on understanding the cross-SIEM's documentation and remapping the logic. Ming Xu 0006, Hongtai Wang, Yanpei Guo, Zhengmin Yu, Weili Han, Hoon Wei Lim, Jin Song Dong 0001, Jiaheng Zhang |
WWW | 1 |
| 2025 | On the Account Security Risks Posed by Password Strength Meters
Ming Xu 0006, Weili Han, Jitao Yu, Yun Lin 0001, Jin Song Dong 0001 |
AsiaCCS | 1 |
| 2025 | Using Parallel Techniques to Accelerate PCFG-Based Password Cracking AttacksabstractTextual passwords play an important role among access-control mechanisms and are usually stored as ciphertext in the server. However, an attacker may attempt to hash a large number of candidate passwords to find the match of the target hash of a password database. To crack the password database, attackers in industry usually use the cracking software like Hashcat. Academic researchers recently proposed many data-driven probabilistic models, in which the Probabilistic Context-free Grammars (PCFG, for short) stand out. Despite the great cracking efficiency, the data-driven models are seldom used by industrial practice due to the significant slow generation speed of password candidates. To bridge the gap and promote the efficient data-driven models being practically used in industry, we propose that using parallel techniques to accelerate the candidate password generation, enabling the integration of PCFG models into the practically-used Hashcat tool. To this end, we mainly propose two algorithms to accelerate the password generation for PCFG-based models: first, we design a storage structure with the memory load balance strategy to more evenly store the data structures used to generate passwords; second, we design an algorithm to produce candidate passwords in parallel by different threads. Based on the two algorithms, we proposeParallel_PCFG, and implementParallel_PCFGupon Hashcat based on its built-in GPU kernel. We comprehensively evaluateParallel_PCFGagainst state-of-the-art data-driven models, and find thatParallel_PCFGonly takes 14.33% of time to achieve the same cracking rates compared with the best-performing models, paving a way about the integration between PCFG-based models and Hashcat. Ming Xu 0006, Kai Zhang 0006, Jitao Yu, Luwei Cheng, Weili Han |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Detecting and Explaining Anomalies Caused by Web Tamper Attacks via Building Consistency-based NormalityabstractWeb applications are crucial infrastructures in the modern society, which have high demand of reliability and security. However, their frontend can be manipulable by the clients (e.g., the frontend code can be modified to bypass some validation steps), which incurs the runtime anomaly when operating the web service. Existing state-of-the-art anomaly detectors largely learn a deep learning model from the collected logs to predict abnormal logs with a probability. While effective in general, those approaches can suffer from (1) inaccuracy caused by subtle difference between the normal and abnormal/attack logs and (2) additional efforts for root cause analysis. Ming Xu 0006, Yun Lin 0001, Xiwen Teoh, Xiaofei Xie, Frank Liaw, Hongyu Zhang 0002, Jin Song Dong 0001 |
ASE | 2 |
| 2023 | Improving Real-world Password Guessing Attacks via Bi-directional Transformers
Ming Xu 0006, Jitao Yu, Chuanwang Wang, Haoqi Wu, Weili Han |
USENIX Security Symposium | 1 |
| 2022 | #Segments: A Dominant Factor of Password Security to Resist against Data-driven Guessing
Chuanwang Wang, Ming Xu 0006, Weili Han |
Comput. Secur. | 3 |
| 2021 | Digit Semantics based Optimization for Practical Password Cracking ToolsabstractUsers usually create their passwords with meaningful digits, i.e. digit semantics, which can be partially exploited by probabilistic password guessing models with a data-driven methodology for better efficiency. However, these semantics are largely ignored by current practical password cracking tools, like John the Ripper (JtR) and Hashcat. Chuanwang Wang, Wenqiang Ruan, Ming Xu 0006, Weili Han |
ACSAC | 5 |
| 2021 | Chunk-Level Password Guessing: Towards Modeling Refined Password Composition RepresentationsabstractTextual password security hinges on the guessing models adopted by attackers, in which a suitable password composition representation is an influential factor. Unfortunately, the conventional models roughly regard a password as a sequence of characters, or natural-language-based words, which are password-irrelevant. Experience shows that passwords exhibit internal and refined patterns, e.g., "4ever, ing or 2015", varying significantly among periods and regions. However, the refined representations and their security impacts could not be automatically understood by state-of-the-art guessing models (e.g., Markov). Ming Xu 0006, Chuanwang Wang, Jitao Yu, Kai Zhang 0006, Weili Han |
CCS | 1 |
| 2021 | TransPCFG: Transferring the Grammars From Short Passwords to Guess Long Passwords EffectivelyabstractLong passwords are gaining popularity in password policy recommendations; however, data-driven guessing studies are woefully inadequate in adapting to long passwords, lacking in both guessing efficiency and their composition guidelines. For state-of-the-art data-driven password guessing methods such as PCFGs (Probabilistic Context-free Grammars), their guessing efficiency is limited by the presence of a large scale training data, or the lack thereof. Given that long passwords leaked in the real world are typically scarce, coupled with the fact that the data-driven methods’ performance depends on training data, obtaining good performance on long passwords has become a key challenge. To overcome the dataset limitation, we propose a frameworkTransPCFG, that transfers the knowledge, (i.e., grammars in PCFGs), from short passwords to facilitate long password guessing. We further perform an empirical evaluation based on three real-world datasets and the results demonstrate superior performance over the state-of-the-art data-driven guessing methods under${10}^{14}$offline guesses. For passwords with 16 characters,TransPCFGcan compromise an average of 23.30% of the passwords, outperforming PCFG_v4.1 by 56.10%. Additionally,for better password-composition guidelines, we find that long password-composition policies requiring more segments are more resistant to guessing attacks. For the segment, the password12zxcvbnword1997has four segments since it follows the template${Digit}_{2}{Keyboard}_{6}{Letter}_{4}{Year}_{4}$. We thus recommend users to create long passwords with four or more segments instead of the widely recommended more character classes for security. Weili Han, Ming Xu 0006, Chuanwang Wang, Kai Zhang 0006, Xiaoyang Sean Wang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | An Explainable Password Strength Meter Addon via Textual Pattern RecognitionabstractTextual passwords are still dominating the authentication of remote file sharing and website logins, although researchers recently showed several vulnerabilities about this authentication mechanism. When a user creates or changes a password, a website usually leverages a password strength meter (PSM for short) to show the strength of the password. When the password is evaluated as a weak one, the user may replace the password with a stronger or securer one. However, the user is usually confused when the password, especially a frequently used password, is shown as a weak one. We argue that an explainable password strength meter addon, which could show the reasons of weak, may help users to more effectively create a secure password. Unfortunately, we find few sites in Alexa global top 100 showing these details. Motivated to help users with an explainable PSM, this paper proposes an addon to PSMs providing feedbacks in the form of pattern passwords explaining why a password is weak. This PSM addon can detect twelve types of patterns, which cover a very large proportion among 70 million of leaked real passwords from high-profile websites. According to our evaluation and user study, our PSM addon, which leverages textual pattern passwords, can effectively detect these popular patterns and effectively help users create securer passwords. Ming Xu 0006, Weili Han |
Secur. Commun. Networks | 1 |