EDBT 2026 Demo / reviewers in the wild / expert
Yugo Murawaki
dblp:94/3673
· DBLP profile ↗
29ranked-venue papers
13as first author
8since 2021 · last 2026
0000-0002-0863-1507ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 13 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
3 papers |
Digital forensics and information hiding · 100% | |
| Theoretical computer science
1 paper |
Coding theory · 100% | |
| Artificial intelligence
6 papers |
Information extraction and text analysis · 30% Generative modeling · 22% Knowledge representation and reasoning · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 62% Bioinformatics and computational biology · 38% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Digital forensics and information hiding › steganography
linguistic steganography |
2.0 | 2 | 2026 | Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography · ACL (1) 2026 Efficient Provably Secure Linguistic Steganography via Range Coding · ACL (1) 2026 |
Digital forensics and information hiding › steganography › secure steganography
provably secure steganography |
1.0 | 1 | 2026 | Efficient Provably Secure Linguistic Steganography via Range Coding · ACL (1) 2026 |
Coding theory › source coding
entropy coding |
1.0 | 1 | 2026 | Efficient Provably Secure Linguistic Steganography via Range Coding · ACL (1) 2026 |
Coding theory › source coding › entropy coding
range coding |
1.0 | 1 | 2026 | Efficient Provably Secure Linguistic Steganography via Range Coding · ACL (1) 2026 |
Digital forensics and information hiding
steganography |
0.9 | 1 | 2025 | Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models · EMNLP 2025 |
Digital forensics and information hiding › steganography
text steganography |
0.9 | 1 | 2025 | Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models · EMNLP 2025 |
Digital forensics and information hiding › watermarking
text watermarking |
0.9 | 1 | 2025 | Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models · EMNLP 2025 |
Digital forensics and information hiding
watermarking |
0.9 | 1 | 2025 | Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models · EMNLP 2025 |
Bioinformatics and computational biology › phylogenetics
phylogenetic inference |
0.5 | 2 | 2020 | Analyzing Correlated Evolution of Multiple Features Using Latent Representations · EMNLP 2018 Latent Geographical Factors for Analyzing the Evolution of Dialects in Contact · EMNLP (1) 2020 |
Machine learning › Generative modeling › generative model
probabilistic generative model |
0.4 | 1 | 2020 | Latent Geographical Factors for Analyzing the Evolution of Dialects in Contact · EMNLP (1) 2020 |
Computational social science and digital humanities
linguistic typology |
0.3 | 1 | 2018 | Analyzing Correlated Evolution of Multiple Features Using Latent Representations · EMNLP 2018 |
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
prompt distillation |
0.3 | 1 | 2026 | Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography · ACL (1) 2026 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.1 | 1 | 2011 | Non-parametric Bayesian Segmentation of Japanese Noun Phrases · EMNLP 2011 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 1 | 2011 | Non-parametric Bayesian Segmentation of Japanese Noun Phrases · EMNLP 2011 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2018 | Analyzing Correlated Evolution of Multiple Features Using Latent Representations · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
morphological analysis |
0.1 | 1 | 2008 | Online Acquisition of Japanese Unknown Morphemes using Morphological Constraints · EMNLP 2008 |
Natural language and speech › Language models and text generation
low-resource language processing |
0.0 | 1 | 2008 | Online Acquisition of Japanese Unknown Morphemes using Morphological Constraints · EMNLP 2008 |
Methods — techniques the papers use, named apart from their topics
self-distillation · 2.0range coding · 2.0prompt distillation · 2.0language model · 2.0anchored sliding window · 2.0tokenization · 0.9large language model · 0.9probabilistic generative model · 0.9admixture analysis · 0.9phylogenetic inference · 0.7latent representation · 0.7minimally supervised learning · 0.4distant supervision · 0.4discourse relations · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Provably Secure Linguistic Steganography via Range CodingabstractLinguistic steganography involves embedding secret messages within seemingly innocuous texts to enable covert communication.Provable security, which is a long-standing goal and key motivation, has been extended to languagemodel-based steganography.Previous provably secure approaches have achieved perfect imperceptibility, measured by zero Kullback-Leibler (KL) divergence, but at the expense of embedding capacity.In this paper, we attempt to directly use a classic entropy coding method (range coding) to achieve secure steganography, and then propose an efficient and provably secure linguistic steganographic method with a rotation mechanism.Experiments across various language models show that our method achieves around 100% entropy utilization (embedding efficiency) for embedding capacity, outperforming the existing baseline methods.Moreover, it achieves high embedding speeds (up to 1554.66 bits/s on GPT-2).The code is available at github.com/ryehr/RRC_steganography. Ruiyi Yan, Yugo Murawaki |
ACL (1) | 2 |
| 2026 | Anchored Sliding Window: Toward Robust and Imperceptible Linguistic SteganographyabstractLinguistic steganography based on language models typically assumes that steganographic texts are transmitted without alteration, making them fragile to even minor modifications.While previous work mitigates this fragility by limiting the context window, it significantly compromises text quality.In this paper, we propose the anchored sliding window (ASW) framework to improve imperceptibility and robustness.In addition to the latest tokens, the prompt and a bridge context are anchored within the context window, encouraging the model to compensate for the excluded tokens.We formulate the optimization of the bridge context as a variant of prompt distillation, which we further extend using self-distillation strategies.Experiments show that our ASW significantly and consistently outperforms the baseline method in text quality, imperceptibility, and robustness across diverse settings.The code is available at github.com/ryehr/ASW_steganography. Ruiyi Yan, Shiao Meng, Yugo Murawaki |
ACL (1) | 3 |
| 2026 | Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Keizo Kato, Chenhui Chu, Yugo Murawaki, Sadao Kurohashi |
LREC | 3 |
| 2025 | Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language ModelsabstractLarge language models have significantly enhanced the capacities and efficiency of text generation.On the one hand, they have improved the quality of text-based steganography.On the other hand, they have also underscored the importance of watermarking as a safeguard against malicious misuse.In this study, we focus on tokenization inconsistency (TI) between the sender and the receiver in steganography and watermarking, where TI can undermine robustness.Our investigation reveals that the problematic tokens responsible for TI exhibit two key characteristics: infrequency and temporariness.Based on these findings, we propose two tailored solutions for TI elimination: a stepwise verification method for steganography and a post-hoc rollback method for watermarking.Experiments show that (1) compared to traditional disambiguation methods in steganography, directly addressing TI leads to improvements in fluency, imperceptibility, and antisteganalysis capacity; (2) for watermarking, addressing TI enhances detectability and robustness against attacks.The code is available at https://github.com/ryehr/Consistency. Ruiyi Yan, Yugo Murawaki |
EMNLP | 2 |
| 2024 | Domain Transferable Semantic Frames for Expert Interview DialoguesabstractInterviews are an effective method to elicit critical skills to perform particular processes in various domains. In order to understand the knowledge structure of these domain-specific processes, we consider semantic role and predicate annotation based on Frame Semantics. We introduce a dataset of interview dialogues with experts in the culinary and gardening domains, each annotated with semantic frames. This dataset consists of (1) 308 interview dialogues related to the culinary domain, originally assembled by Okahisa et al. (2022), and (2) 100 interview dialogues associated with the gardening domain, which we newly acquired. The labeling specifications take into account the domain-transferability by adopting domain-agnostic labels for frame elements. In addition, we conducted domain transfer experiments from the culinary domain to the gardening domain to examine the domain transferability with our dataset. The experimental results showed the effectiveness of our domain-agnostic labeling scheme. Taishi Chika, Taro Okahisa, Takashi Kodama, Yin Jou Huang, Yugo Murawaki, Sadao Kurohashi |
LREC/COLING | 5 |
| 2024 | Principal Component Analysis as a Sanity Check for Bayesian Phylolinguistic ReconstructionabstractBayesian approaches to reconstructing the evolutionary history of languages rely on the tree model, which assumes that these languages descended from a common ancestor and underwent modifications over time. However, this assumption can be violated to different extents due to contact and other factors. Understanding the degree to which this assumption is violated is crucial for validating the accuracy of phylolinguistic inference. In this paper, we propose a simple sanity check: projecting a reconstructed tree onto a space generated by principal component analysis. By using both synthetic and real data, we demonstrate that our method effectively visualizes anomalies, particularly in the form of jogging. Yugo Murawaki |
LREC/COLING | 1 |
| 2024 | Identifying Source Language Expressions for Pre-editing in Machine TranslationabstractMachine translation-mediated communication can benefit from pre-editing source language texts to ensure accurate transmission of intended meaning in the target language. The primary challenge lies in identifying source language expressions that pose difficulties in translation. In this paper, we hypothesize that such expressions tend to be distinctive features of texts originally written in the source language (native language) rather than translations generated from the target language into the source language (machine translation). To identify such expressions, we train a neural classifier to distinguish native language from machine translation, and subsequently isolate the expressions that contribute to the model’s prediction of native language. Our manual evaluation revealed that our method successfully identified characteristic expressions of the native language, despite the noise and the inherent nuances of the task. We also present case studies where we edit the identified expressions to improve translation quality. Norizo Sakaguchi, Yugo Murawaki, Chenhui Chu, Sadao Kurohashi |
LREC/COLING | 2 |
| 2021 | Frustratingly Easy Edit-based Linguistic Steganography with a Masked Language ModelabstractWith advances in neural language models, the focus of linguistic steganography has shifted from edit-based approaches to generationbased ones.While the latter's payload capacity is impressive, generating genuine-looking texts remains challenging.In this paper, we revisit edit-based linguistic steganography, with the idea that a masked language model offers an off-the-shelf solution.The proposed method eliminates painstaking rule construction and has a high payload capacity for an edit-based model.It is also shown to be more secure against automatic detection than a generation-based method while offering better control of the security/payload capacity tradeoff. Honai Ueoka, Yugo Murawaki, Sadao Kurohashi |
NAACL-HLT | 2 |
| 2020 | Native-like Expression Identification by Contrasting Native and Proficient Second Language SpeakersabstractWe propose a novel task of native-like expression identification by contrasting texts written by native speakers and those by proficient second language speakers.This task is highly challenging mainly because 1) the combinatorial nature of expressions prevents us from choosing candidate expressions a priori and 2) the distributions of the two types of texts overlap considerably.Our solution to the first problem is to combine a powerful neural network-based classifier of sentencelevel nativeness with an explainability method that measures an approximate contribution of a given expression to the classifier's prediction.To address the second problem, we introduce a special label neutral and reformulate the classification task as complementary-label learning.Our crowdsourcing-based evaluation and in-depth analysis suggest that our method successfully uncovers linguistically interesting usages distinctive of native speech. Oleksandr Harust, Yugo Murawaki, Sadao Kurohashi |
COLING | 2 |
| 2020 | Latent Geographical Factors for Analyzing the Evolution of Dialects in ContactabstractAnalyzing the evolution of dialects remains a challenging problem because contact phenomena hinder the application of the standard tree model.Previous statistical approaches to this problem resort to admixture analysis, where each dialect is seen as a mixture of latent ancestral populations.However, such ancestral populations are hardly interpretable in the context of the tree model.In this paper, we propose a probabilistic generative model that represents latent factors as geographical distributions.We argue that the proposed model has higher affinity with the tree model because a tree can alternatively be represented as a set of geographical distributions.Experiments involving synthetic and real data suggest that the proposed method is both quantitatively and qualitatively superior to the admixture model. Yugo Murawaki |
EMNLP (1) | 1 |
| 2020 | Adapting BERT to Implicit Discourse Relation Classification with a Focus on Discourse ConnectivesabstractBERT, a neural network-based language model pre-trained on large corpora, is a breakthrough in natural language processing, significantly outperforming previous state-of-the-art models in numerous tasks. However, there have been few reports on its application to implicit discourse relation classification, and it is not clear how BERT is best adapted to the task. In this paper, we test three methods of adaptation. (1) We perform additional pre-training on text tailored to discourse classification. (2) In expectation of knowledge transfer from explicit discourse relations to implicit discourse relations, we add a task named explicit connective prediction at the additional pre-training step. (3) To exploit implicit connectives given by treebank annotators, we add a task named implicit connective prediction at the fine-tuning step. We demonstrate that these three techniques can be combined straightforwardly in a single training pipeline. Through comprehensive experiments, we found that the first and second techniques provide additional gain while the last one did not. Yudai Kishimoto, Yugo Murawaki, Sadao Kurohashi |
LREC | 2 |
| 2019 | A Hybrid Generative/Discriminative Model for Rapid Prototyping of Domain-Specific Named Entity Recognition
Suzushi Tomori, Yugo Murawaki, Shinsuke Mori |
CICLing (2) | 2 |
| 2019 | Minimally Supervised Learning of Affective Events Using Discourse RelationsabstractJun Saito, Yugo Murawaki, Sadao Kurohashi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jun Saito, Yugo Murawaki, Sadao Kurohashi |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Bayesian Learning of Latent Representations of Language StructuresabstractWe borrow the concept of representation learning from deep learning research, and we argue that the quest for Greenbergian implicational universals can be reformulated as the learning of good latent representations of languages, or sequences of surface typological features. By projecting languages into latent representations and performing inference in the latent space, we can handle complex dependencies among features in an implicit manner. The most challenging problem in turning the idea into a concrete computational model is the alarmingly large number of missing values in existing typological databases. To address this problem, we keep the number of model parameters relatively small to avoid overfitting, adopt the Bayesian learning framework for its robustness, and exploit phylogenetically and/or spatially related languages as additional clues. Experiments show that the proposed model recovers missing values more accurately than others and that some latent variables exhibit phylogenetic and spatial signals comparable to those of surface features. Yugo Murawaki |
Comput. Linguistics | 1 |
| 2018 | A Knowledge-Augmented Neural Network Model for Implicit Discourse Relation ClassificationabstractIdentifying discourse relations that are not overtly marked with discourse connectives remains a challenging problem. The absence of explicit clues indicates a need for the combination of world knowledge and weak contextual clues, which can hardly be learned from a small amount of manually annotated data. In this paper, we address this problem by augmenting the input text with external knowledge and context and by adopting a neural network model that can effectively handle the augmented text. Experiments show that external knowledge did improve the classification accuracy. Contextual information provided no significant gain for implicit discourse relations, but it did for explicit ones. Yudai Kishimoto, Yugo Murawaki, Sadao Kurohashi |
COLING | 2 |
| 2018 | Analyzing Correlated Evolution of Multiple Features Using Latent RepresentationsabstractStatistical phylogenetic models have allowed the quantitative analysis of the evolution of a single categorical feature and a pair of binary features, but correlated evolution involving multiple discrete features is yet to be explored.Here we propose latent representation-based analysis in which (1) a sequence of discrete surface features is projected to a sequence of independent binary variables and (2) phylogenetic inference is performed on the latent space.In the experiments, we analyze the features of linguistic typology, with a special focus on the order of subject, object and verb.Our analysis suggests that languages sharing the same word order are not necessarily a coherent group but exhibit varying degrees of diachronic stability depending on other features. Yugo Murawaki |
EMNLP | 1 |
| 2018 | Universal Dependencies Version 2 for Japanese
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto 0001, Mai Omura, Yugo Murawaki |
LREC | 9 |
| 2018 | Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations
Yudai Kishimoto, Shinnosuke Sawada, Yugo Murawaki, Daisuke Kawahara, Sadao Kurohashi |
LREC | 3 |
| 2018 | Annotating Modality Expressions and Event Factuality for a Japanese Chess Commentary Corpus
Suguru Matsuyoshi, Hirotaka Kameko, Yugo Murawaki, Shinsuke Mori |
LREC | 3 |
| 2017 | Diachrony-aware Induction of Binary Latent Representations from Typological FeaturesabstractAlthough features of linguistic typology are a promising alternative to lexical evidence for tracing evolutionary history of languages, a large number of missing values in the dataset pose serious difficulties for statistical modeling. In this paper, we combine two existing approaches to the problem: (1) the synchronic approach that focuses on interdependencies between features and (2) the diachronic approach that exploits phylogenetically- and/or spatially-related languages. Specifically, we propose a Bayesian model that (1) represents each language as a sequence of binary latent parameters encoding inter-feature dependencies and (2) relates a language’s parameters to those of its phylogenetic and spatial neighbors. Experiments show that the proposed model recovers missing values more accurately than others and that induced representations retain phylogenetic and spatial signals observed for surface features. Yugo Murawaki |
IJCNLP(1) | 1 |
| 2016 | Contrasting Vertical and Horizontal Transmission of Typological FeaturesabstractLinguistic typology provides features that have a potential of uncovering deep phylogenetic relations among the world’s languages. One of the key challenges in using typological features for phylogenetic inference is that horizontal (spatial) transmission obscures vertical (phylogenetic) signals. In this paper, we characterize typological features with respect to the relative strength of vertical and horizontal transmission. To do this, we first construct (1) a spatial neighbor graph of languages and (2) a phylogenetic neighbor graph by collapsing known language families. We then develop an autologistic model that predicts a feature’s distribution from these two graphs. In the experiments, we managed to separate vertically and/or horizontally stable features from unstable ones, and the results are largely consistent with previous findings. Kenji Yamauchi, Yugo Murawaki |
COLING | 2 |
| 2016 | Wikification for Scriptio Continua
Yugo Murawaki, Shinsuke Mori |
LREC | 1 |
| 2016 | Statistical Modeling of Creole GenesisabstractCreole languages do not fit into the traditional tree model of evolutionary history because multiple languages are involved in their formation.In this paper, we present several statistical models to explore the nature of creole genesis.After reviewing quantitative studies on creole genesis, we first tackle the question of whether creoles are typologically distinct from non-creoles.By formalizing this question as a binary classification problem, we demonstrate that a linear classifier fails to separate creoles from non-creoles although the two groups have substantially different distributions in the feature space.We then model a creole language as a mixture of source languages plus a special restructurer.We find a pervasive influence of the restructurer in creole genesis and some statistical universals in it, paving the way for more elaborate statistical models. Yugo Murawaki |
HLT-NAACL | 1 |
| 2015 | Continuous Space Representations of Linguistic Typology and their Application to Phylogenetic InferenceabstractFor phylogenetic inference, linguistic typology is a promising alternative to lexical evidence because it allows us to compare an arbitrary pair of languages.A challenging problem with typology-based phylogenetic inference is that the changes of typological features over time are less intuitive than those of lexical features.In this paper, we work on reconstructing typologically natural ancestors To do this, we leverage dependencies among typological features.We first represent each language by continuous latent components that capture feature dependencies.We then combine them with a typology evaluator that distinguishes typologically natural languages from other possible combinations of features.We perform phylogenetic inference in the continuous space and use the evaluator to ensure the typological naturalness of inferred ancestors.We show that the proposed method reconstructs known language families more accurately than baseline methods.Lastly, assuming the monogenesis hypothesis, we attempt to reconstruct a common ancestor of the world's languages. Yugo Murawaki |
HLT-NAACL | 1 |
| 2013 | Global Model for Hierarchical Multi-Label Text Classification
Yugo Murawaki |
IJCNLP | 1 |
| 2012 | Semi-Supervised Noun Compound Analysis with Edge and Span Features
Yugo Murawaki, Sadao Kurohashi |
COLING | 1 |
| 2011 | Non-parametric Bayesian Segmentation of Japanese Noun Phrases
Yugo Murawaki, Sadao Kurohashi |
EMNLP | 1 |
| 2010 | Online Japanese Unknown Morpheme Detection using Orthographic Variation
Yugo Murawaki, Sadao Kurohashi |
LREC | 1 |
| 2008 | Online Acquisition of Japanese Unknown Morphemes using Morphological Constraints
Yugo Murawaki, Sadao Kurohashi |
EMNLP | 1 |