VLDB 2026 Research / reviewers in the wild / expert
Eugene Jang
dblp:294/0262
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew 0002, Marek Suppa, Hila Gonen, Shachar Mirkin, Vasile Florian Pais, Kaja Dobrovoljc, Voula Giouli, Jun Kevin, Eugene Jang, Eungseo Kim, Jeongyeon Seo, Xenophon Gialis, Yuval Pinter |
LREC | 10 |
| 2025 | Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level TokenizersabstractTokenization is a crucial step that bridges human-readable text with model-readable discrete tokens.However, recent studies have revealed that tokenizers can be exploited to elicit unwanted model behaviors.In this work, we investigate incomplete tokens, i.e., undecodable tokens with stray bytes resulting from bytelevel byte-pair encoding (BPE) tokenization.We hypothesize that such tokens are heavily reliant on their adjacent tokens and are fragile when paired with unfamiliar tokens.To demonstrate this vulnerability, we introduce improbable bigrams: out-of-distribution combinations of incomplete tokens designed to exploit their dependency.Our experiments show that improbable bigrams are significantly prone to hallucinatory behaviors.Surprisingly, the same phrases have drastically lower rates of hallucination (90% reduction in Llama3.1)when an alternative tokenization is used.We caution against the potential vulnerabilities introduced by byte-level BPE tokenizers, which may introduce blind spots to language models. Eugene Jang, Kimin Lee, Jin-Woo Chung, Keuntae Park, Seungwon Shin 0001 |
EMNLP | 1 |
| 2025 | Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionabstractIn light of rising drug-related concerns and the increasing role of social media, sales and discussions of illicit drugs have become commonplace online. Social media platforms hosting user-generated content must therefore perform content moderation, which is a difficult task due to the vast amount of jargon used in drug discussions. Previous works on drug jargon detection were limited to extracting a list of terms, but these approaches have fundamental problems in practical application. First, they are trivially evaded using word substitutions. Second, they cannot distinguish whether euphemistic terms (pot, crack) are being used as drugs or as their benign meanings. We argue that drug content moderation should be done using contexts, rather than relying on a banlist. However, manually annotated datasets for training such a task are not only expensive but also prone to becoming obsolete. We present JEDIS, a framework for detecting illicit drug jargon terms by analyzing their contexts. JEDIS utilizes a novel approach that combines distant supervision and delexicalization, which allows JEDIS to be trained without human-labeled data while being robust to new terms and euphemisms. Experiments on two manually annotated datasets show JEDIS significantly outperforms state-of-the-art word-based baselines in terms of F1-score and detection coverage in drug jargon detection. We also conduct qualitative analysis that demonstrates JEDIS is robust against pitfalls faced by existing approaches. Minkyoo Song, Eugene Jang, Jaehan Kim, Seungwon Shin 0001 |
KDD (1) | 2 |
| 2025 | Tweezers: A Framework for Security Event Detection via Event Attribution-centric Tweet Embedding
Hanna Kim, Eugene Jang, Dayeon Yim, Kicheol Kim, Jin-Woo Chung, Seungwon Shin 0001, Xiaojing Liao |
NDSS | 3 |
| 2024 | DRAINCLoG: Detecting Rogue Accounts with Illegally-obtained NFTs using Classifiers Learned on Graphs
Hanna Kim, Eugene Jang, Jin-Woo Chung, Seungwon Shin 0001 |
NDSS | 3 |
| 2023 | WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language ModelsabstractContent Warning: This paper contains examples of homophobic and transphobic stereotypes.We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community.The benchmark is community-sourced, via application of a novel method that generates a bias benchmark from a community survey.We apply our benchmark to several popular LLMs and find that off-the-shelf models generally do exhibit considerable anti-queer bias.Finally, we show that LLM bias against a marginalized community can be somewhat mitigated by finetuning on data written about or by members of that community, and that social media text written by community members is more effective than news text written about the community by non-members.Our method for community-in-the-loop benchmark development provides a blueprint for future researchers to develop community-driven, harms-grounded LLM benchmarks for other marginalized communities. Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan May |
ACL (1) | 3 |
| 2023 | DarkBERT: A Language Model for the Dark Side of the InternetabstractRecent research has suggested that there are clear differences in the language used in the Dark Web compared to that of the Surface Web.As studies on the Dark Web commonly require textual analysis of the domain, language models specific to the Dark Web may provide valuable insights to researchers.In this work, we introduce DarkBERT, a language model pretrained on Dark Web data.We describe the steps taken to filter and compile the text data used to train DarkBERT to combat the extreme lexical and structural diversity of the Dark Web that may be detrimental to building a proper representation of the domain.We evaluate Dark-BERT and its vanilla counterpart along with other widely used language models to validate the benefits that a Dark Web domain specific model offers in various use cases.Our evaluations show that DarkBERT outperforms current language models and may serve as a valuable resource for future research on the Dark Web. Youngjin Jin, Eugene Jang, Jin-Woo Chung, Seungwon Shin 0001 |
ACL (1) | 2 |
| 2023 | Measuring Online Emotional Reactions to EventsabstractThe rich and dynamic information environment of social media provides researchers, policy makers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using this data to understand social behavior is difficult due heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect, and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method on a corpus of tweets from a large US metropolitan area between January and August, 2020, covering a period of great social change. We demonstrate that our method is able to disaggregate topics to measure population's emotional and moral reactions. This capability allows for better monitoring of population's reactions during crises using online data. Siyi Guo, Ashwin Rao, Eugene Jang, Yuanfeixue Nan, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ASONAM | 4 |
| 2022 | Shedding New Light on the Language of the Dark WebabstractYoungjin Jin, Eugene Jang, Yongjae Lee, Seungwon Shin, Jin-Woo Chung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Youngjin Jin, Eugene Jang, Seungwon Shin 0001, Jin-Woo Chung |
NAACL-HLT | 2 |
| 2021 | Generating Negative Samples by Manipulating Golden Responses for Unsupervised Learning of a Response Evaluation ModelabstractEvaluating the quality of responses generated by open-domain conversation systems is a challenging task.This is partly because there can be multiple appropriate responses to a given dialogue history.Reference-based metrics that rely on comparisons to a set of known correct responses often fail to account for this variety, and consequently correlate poorly with human judgment.To address this problem, researchers have investigated the possibility of assessing response quality without using a set of known correct responses.Tao et al. (2018) demonstrated that an automatic response evaluation model could be made using unsupervised learning for the next-utterance prediction (NUP) task.For unsupervised learning of such a model, we propose a method of manipulating a golden response to create a new negative response that is designed to be inappropriate within the context while maintaining high similarity with the original golden response.We find, from our experiments on English datasets, that using the negative samples generated by our method alongside random negative samples can increase the model's correlation with human evaluations.The process of generating such negative samples is automated and does not rely on human annotation. 1 Chaehun Park, Eugene Jang, Wonsuk Yang, Jong Park |
NAACL-HLT | 2 |
| 2021 | Optimizing Domain Specificity of Transformer-based Language Models for Extractive Summarization of Financial News Articles in Korean
Huije Lee, Wonsuk Yang, Chaehun Park, Hoyun Song, Eugene Jang, Jong C. Park |
PACLIC | 5 |