Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jonathan Rusert

dblp:242/4812 · also Jon Rusert · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Security and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Security and privacy of machine learning · 78% Digital forensics and information hiding · 17% Privacy and data protection · 5%
Artificial intelligence
2 papers
Trustworthy machine learning · 58% Information extraction and text analysis · 42%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
adversarial attack
0.912025
RedHerring Attack: Testing the Reliability of Attack Detection · EMNLP 2025
Security and privacy of machine learning › adversarial attack
evasion attack
0.912025
RedHerring Attack: Testing the Reliability of Attack Detection · EMNLP 2025
Security and privacy of machine learning › adversarial attack
textual adversarial attack
0.912025
RedHerring Attack: Testing the Reliability of Attack Detection · EMNLP 2025
Natural language and speech › Information extraction and text analysis
abusive language detection
0.612022
On the Robustness of Offensive Language Classifiers · ACL (1) 2022
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.612022
On the Robustness of Offensive Language Classifiers · ACL (1) 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
On the Robustness of Offensive Language Classifiers · ACL (1) 2022
Digital forensics and information hiding
authorship attribution
0.612022
Adversarial Authorship Attribution for Deobfuscation · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
text classification
0.312025
RedHerring Attack: Testing the Reliability of Attack Detection · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

contrastive reasoning · 1.7confidence check · 1.7word embeddings · 0.6attention · 0.6adversarial training · 0.6adversarial attack · 0.6
YearPublicationVenuePosition
2025 BinarySelect to Improve Accessibility of Black-Box Attack Research
abstract
Adversarial text attack research is useful for testing the robustness of NLP models, however, the rise of transformers has greatly increased the time required to test attacks. Especially when researchers do not have access to adequate resources (e.g. GPUs). This can hinder attack research, as modifying one example for an attack can require hundreds of queries to a model, especially for black-box attacks. Often these attacks remove one token at a time to find the ideal one to change, requiring n queries (the length of the text) right away. We propose a more efficient selection method called BinarySelect which combines binary search and attack selection methods to greatly reduce the number of queries needed to find a token. We find that BinarySelect only needs log_2(n) * 2 queries to find the first token compared to n queries. We also test BinarySelect in an attack setting against 5 classifiers across 3 datasets and find a viable tradeoff between number of queries saved and attack effectiveness. For example, on the Yelp dataset, the number of queries is reduced by 32% (72 less) with a drop in attack effectiveness of only 5 points. We believe that BinarySelect can help future researchers study adversarial attacks and black-box problems more efficiently and opens the door for researchers with access to less resources.
Shatarupa Ghosh, Jonathan Rusert
COLING2
2025 RedHerring Attack: Testing the Reliability of Attack Detection
abstract
In response to adversarial text attacks, attack detection models have been proposed and shown to successfully identify text modified by adversaries.Attack detection models can be leveraged to provide an additional check for NLP models and give signals for human input.However, the reliability of these models has not yet been thoroughly explored.Thus, we propose and test a novel attack setting and attack, Red-Herring.RedHerring aims to make attack detection models unreliable by modifying a text to cause the detection model to predict an attack, while keeping the classifier correct.This creates a tension between the classifier and detector.If a human sees that the detector is giving an "incorrect" prediction, but the classifier a correct one, then the human will see the detector as unreliable.We test this novel threat model on 4 datasets against 3 detectors defending 4 classifiers.We find that RedHerring is able to drop detection accuracy between 20 -71 points, while maintaining (or improving) classifier accuracy.As an initial defense, we propose a simple confidence check which requires no retraining of the classifier or detector and increases detection accuracy greatly.This novel threat model offers new insights into how adversaries may target detection models.
Jonathan Rusert
EMNLP1
2024 VertAttack: Taking Advantage of Text Classifiers' Horizontal Vision
abstract
Text classification systems have continuously improved in performance over the years.However, nearly all current SOTA classifiers have a similar shortcoming, they process text in a horizontal manner.Vertically written words will not be recognized by a classifier.In contrast, humans are easily able to recognize and read words written both horizontally and vertically.Hence, a human adversary could write problematic words vertically and the meaning would still be preserved to other humans.We simulate such an attack, VertAttack.VertAttack identifies which words a classifier is reliant on and then rewrites those words vertically.We find that VertAttack is able to greatly drop the accuracy of 4 different transformer models on 5 datasets.For example, on the SST2 dataset, VertAttack is able to drop RoBERTa's accuracy from 94 to 13%.Furthermore, since VertAttack does not replace the word, meaning is easily preserved.We verify this via a human study and find that crowdworkers are able to correctly label 77% perturbed texts perturbed, compared to 81% of the original texts.We believe VertAttack offers a look into how humans might circumvent classifiers in the future and thus inspire a look into more robust algorithms.042 ual characters, by flipping character, introducing 043 or removing whitespace (Gröndahl et al., 2018), or 044 replacing characters with visually similar charac-045 ters (Eger et al., 2019).Word-based attacks re-046 place words with similar words which are less 047 known to the target classifier (Li et al., 2020; Wang 048 et al., 2022).One weakness of current SOTA at-049 tacks is that they constrain themselves to horizontal 050 changes.That is, the final result is still read in a 051 left-to-right (English) manner.This is a disadvan-052 tage because the attacker restricts themselves to the 053 same domain as the classifier which is also only 054 able to read text horizontally.055 Humans have the ability to read text in multiple 056 directions, not just horizontally.Thus, a human 057 attacker who wants to communicate a message to 058 others, while avoiding a website automatically clas-059 sifying that text, could write the words vertically 060 and the meaning would still be preserved.We sim-061 1 719 ulate this with VertAttack.VertAttack exploits the current limitation of classifiers' inability to read text vertically.Specifically, VertAttack perturbs input text by changing information rich words from horizontally to vertically written.Our research makes the following contributions
Jonathan Rusert
NAACL-HLT1
2022 On the Robustness of Offensive Language Classifiers
abstract
Social media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale.However, despite their real-world deployment, we do not yet comprehensively understand the extent to which offensive language classifiers are robust against adversarial attacks.Prior work in this space is limited to studying robustness of offensive language classifiers against primitive attacks such as misspellings and extraneous spaces.To address this gap, we systematically analyze the robustness of state-of-theart offensive language classifiers against more crafty adversarial attacks that leverage greedyand attention-based word selection and contextaware embeddings for word replacement.Our results on multiple datasets show that these crafty adversarial attacks can degrade the accuracy of offensive language classifiers by more than 50% while also being able to preserve the readability and meaning of the modified text.
Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan
ACL (1)1
2022 Adversarial Authorship Attribution for Deobfuscation
abstract
Recent advances in natural language processing have enabled powerful privacy-invasive authorship attribution.To counter authorship attribution, researchers have proposed a variety of rule-based and learning-based text obfuscation approaches.However, existing authorship obfuscation approaches do not consider the adversarial threat model.Specifically, they are not evaluated against adversarially trained authorship attributors that are aware of potential obfuscation.To fill this gap, we investigate the problem of adversarial authorship attribution for deobfuscation.We show that adversarially trained authorship attributors are able to degrade the effectiveness of existing obfuscators from 20-30% to 5-10%.We also evaluate the effectiveness of adversarial training when the attributor makes incorrect assumptions about whether and which obfuscator was used.While there is a a clear degradation in attribution accuracy, it is noteworthy that this degradation is still at or above the attribution accuracy of the attributor that is not adversarially trained at all.Our results underline the need for stronger obfuscation approaches that are resistant to deobfuscation.* This paper is third in the series.See (Mahmood et al., 2019) and (Mahmood et al., 2020) for the first two papers.
Wanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan
ACL (1)2
2022 Don't sweat the small stuff, classify the rest: Sample Shielding to protect text classifiers against adversarial attacks
abstract
Deep learning (DL) is being used extensively for text classification.However, researchers have demonstrated the vulnerability of such classifiers to adversarial attacks.Attackers modify the text in a way which misleads the classifier while keeping the original meaning close to intact.State-of-the-art (SOTA) attack algorithms follow the general principle of making minimal changes to the text so as to not jeopardize semantics.Taking advantage of this we propose a novel and intuitive defense strategy called Sample Shielding.It is attacker and classifier agnostic, does not require any reconfiguration of the classifier or external resources and is simple to implement.Essentially, we sample subsets of the input text, classify them and summarize these into a final decision.We shield three popular DL text classifiers with Sample Shielding, test their resilience against four SOTA attackers across three datasets in a realistic threat setting.Even when given the advantage of knowing about our shielding strategy the adversary's attack success rate is <= 10% with only one exception and often < 5%.Additionally, Sample Shielding maintains near original accuracy when applied to original texts.Crucially, we show that the 'make minimal changes' approach of SOTA attackers leads to critical vulnerabilities that can be defended against with an intuitive sampling strategy.1
Jonathan Rusert, Padmini Srinivasan
NAACL-HLT1
2019 No Place to Hide: Inadvertent Location Privacy Leaks on Twitter
abstract
Abstract There is a natural tension between the desire to share information and keep sensitive information private on online social media. Privacy seeking social media users may seek to keep their location private by avoiding the mentions of location revealing words such as points of interest (POIs), believing this to be enough. In this paper, we show that it is possible to uncover the location of a social media user’s post even when it is not geotagged and does not contain any POI information. Our proposed approach Jasoosachieves this by exploiting the shared vocabulary between users who reveal their location and those who do not. To this end, Jasoosuses a variant of the Naive Bayes algorithm to identify location revealing words or hashtags based on both temporal and atemporal perspectives. Our evaluation using tweets collected from four different states in the United States shows that Jasooscan accurately infer the locations of close to half a million tweets corresponding to more than 20,000 distinct users (i.e., more than 50% of the test users) from the four states. Our work demonstrates that location privacy leaks do occur despite due precautions by a privacy conscious user. We design and evaluate countermeasures based Jasoosto mitigate location privacy leaks.
Jonathan Rusert, Osama Khalid, Dat Hong, Zubair Shafiq, Padmini Srinivasan
Proc. Priv. Enhancing Technol.1