EDBT 2026 Demo / reviewers in the wild / expert
Fei Cheng 0002
dblp:06/5591-2
· DBLP profile ↗
23ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0001-5161-0544ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMsabstractYihua Zhu, Qianying Liu, Jiaxin Wang, Fei Cheng, Chaoran Liu, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yihua Zhu 0002, Qianying Liu, Fei Cheng 0002, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira |
ACL (1) | 4 |
| 2026 | Building Effective Japanese Medical LLMs with an Open Recipe for Domain Adaptation through Continued Pre-training
Akiko Aizawa, Yuki Arase, Fei Cheng 0002, Teruhito Kanazawa, Daisuke Kawahara, Kazuma Kobayashi, Takashi Kodama, Sadao Kurohashi, Yusuke Oda, Tsuta Yuma, Zhishen Yang, Rio Yokota |
LREC | 3 |
| 2026 | Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language ModelsabstractModern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but revealing challenge: judging the arrow of time (AoT)-whether a short clip is played forward or backward. We introduce AoT-PsyPhyBENCH, a psychophysically validated benchmark that tests whether VLMs can infer temporal direction in natural videos using the same stimuli and behavioral baselines established for humans. Our comprehensive evaluation of open-weight and proprietary, reasoning and non-reasoning VLMs reveals that most models perform near chance, and even the best model lags far behind human accuracy on physically irreversible processes (e.g., free fall, diffusion/explosion) and causal manual actions (division/addition) that humans recognize almost instantly. These results highlight a fundamental gap in current multimodal systems: while they capture rich visual-semantic correlations, they lack the inductive biases required for temporal continuity and causal understanding. We release the code and data for AoT-PsyPhyBENCH to encourage further progress in the physical and temporal reasoning capabilities of VLMs. Shiho Matta, Lis Pereira, Peitao Han, Shigeru Kitazawa, Fei Cheng 0002 |
LREC | 5 |
| 2026 | Biomedical concept recognition with error-aware negative-enhanced ranking frameworkabstractMOTIVATION: Mention-agnostic biomedical concept recognition (MA-BCR) requires inferring ontology concepts directly from passages, without relying on explicit mention spans. Prior work has mainly focused on generative and classification-based approaches. Ranking-based methods typically use a retrieve-rerank pipeline, and this paradigm has not been systematically studied for MA-BCR. Consequently, it remains unclear how ranking-based approaches compare with existing paradigms and what types of supervision are most beneficial for ranker training under limited annotation settings. RESULTS: Through a systematic comparison of ranking-, generative-, and classification-based paradigms, we show that a two-stage retrieve-rerank architecture is the most robust and scalable backbone for MA-BCR. Building on this finding, we propose ENR, an error-aware negative-enhanced ranking framework that augments training with false positives collected from heterogeneous recognizers, improving reranking performance without increasing inference-time cost. Experiments on MM-HPO and MM-GO (two datasets derived from MedMentions-ST21pv) demonstrate that ENR substantially outperforms prior approaches. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are available in Github at https://github.com/sl-633/enr-recognizer or in Zenodo at https://doi.org/10.5281/zenodo.20730803. Noriki Nishida, Fei Cheng 0002, Takehito Utsuro, Yuji Matsumoto 0001 |
Bioinform. | 3 |
| 2026 | Dissecting GraphRAG: A Modular Analysis of Knowledge Structuring for Factoid Question AnsweringabstractAbstract We present a systematic analysis of module-level design choices in GraphRAG, a retrieval-augmented generation framework that integrates structured knowledge graphs into question answering. Focusing on triple extraction, community clustering, and report generation, we evaluate multiple strategies across two knowledge-intensive benchmarks. Our results show that high-quality triple extraction is critical, as the accuracy and coverage of the resulting knowledge graph can become a bottleneck for downstream reasoning. We also find that the granularity of fundamental knowledge units, as determined by community clustering, has a significant impact on downstream performance: Achieving a balance between factual detail and topical coherence within each unit is important to enable precise and comprehensive retrieval and to facilitate effective multi-hop reasoning. In addition, simple template-based reporting outperforms LLM-based summarization in both accuracy and efficiency. These findings provide practical guidance for the structure- aware design of retrieval-augmented systems. Noriki Nishida, Rumana Ferdous Munne, Narumi Tokunaga, Yuki Yamagata, Fei Cheng 0002, Kouji Kozaki, Yuji Matsumoto 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsabstractZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li, Ke Hu, Zhehuai Chen, Shinji Watanabe, Fei Cheng, Chenhui Chu, Sadao Kurohashi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li 0010, Zhehuai Chen, Shinji Watanabe 0001, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ACL (1) | 9 |
| 2025 | Causal Tree Extraction from Medical Case Reports: A Novel Task for Experts-like Text ComprehensionabstractExtracting causal relationships from a medical case report is essential for comprehending the case, particularly its diagnostic process.Since the diagnostic process is regarded as a bottom-up inference, causal relationships in cases naturally form a multi-layered tree structure.The existing tasks, such as medical relation extraction, are insufficient for capturing the causal relationships of an entire case, as they treat all relations equally without considering the hierarchical structure inherent in the diagnostic process.Thus, we propose a novel task, Causal Tree Extraction (CTE), which receives a case report and generates a causal tree with the primary disease as the root, providing an intuitive understanding of a case's diagnostic process.Subsequently, we construct a Japanese case report CTE dataset, J-Casemap, propose a generation-based CTE method that outperforms the baseline by 20.2 points in the human evaluation, and introduce evaluation metrics that reflect clinician preferences.Further experiments also show that J-Casemap enhances the performance of solving other medical tasks, such as question answering. Sakiko Yahata, Fei Cheng 0002, Sadao Kurohashi, Hisahiko Sato, Ryozo Nagai |
EMNLP | 3 |
| 2024 | An Empirical Study of Synthetic Data Generation for Implicit Discourse Relation RecognitionabstractImplicit Discourse Relation Recognition (IDRR), which is the task of recognizing the semantic relation between given text spans that do not contain overt clues, is a long-standing and challenging problem. In particular, the paucity of training data for some error-prone discourse relations makes the problem even more challenging. To address this issue, we propose a method of generating synthetic data for IDRR using a large language model. The proposed method is summarized as two folds: extraction of confusing discourse relation pairs based on false negative rate and synthesis of data focused on the confusion. The key points of our proposed method are utilizing a confusion matrix and adopting two-stage prompting to obtain effective synthetic data. According to the proposed method, we generated synthetic data several times larger than training examples for some error-prone discourse relations and incorporated it into training. As a result of experiments, we achieved state-of-the-art macro-F1 performance thanks to the synthetic data without sacrificing micro-F1 performance and demonstrated its positive effects especially on recognizing some infrequent discourse relations. Kazumasa Omura, Fei Cheng 0002, Sadao Kurohashi |
LREC/COLING | 2 |
| 2024 | Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on JapaneseabstractThe creation of instruction data and evaluation benchmarks for serving Large language models often involves enormous human annotation. This issue becomes particularly pronounced when rapidly developing such resources for a non-English language like Japanese. Instead of following the popular practice of directly translating existing English resources into Japanese (e.g., Japanese-Alpaca), we propose an efficient self-instruct method based on GPT-4. We first translate a small amount of English instructions into Japanese and post-edit them to obtain native-level quality. GPT-4 then utilizes them as demonstrations to automatically generate Japanese instruction data. We also construct an evaluation benchmark containing 80 questions across 8 categories, using GPT-4 to automatically assess the response quality of LLMs without human references. The empirical results suggest that the models fine-tuned on our GPT-4 self-instruct data significantly outperformed the Japanese-Alpaca across all three base pre-trained models. Our GPT-4 self-instruct data allowed the LLaMA 13B model to defeat GPT-3.5 (Davinci-003) with a 54.37% win-rate. The human evaluation exhibits the consistency between GPT-4’s assessments and human preference. Our high-quality instruction data and evaluation benchmark are released here. Yikun Sun, Nobuhiro Ueda, Sakiko Yahata, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
LREC/COLING | 5 |
| 2023 | ComSearch: Equation Searching with Combinatorial Strategy for Solving Math Word Problems with Weak SupervisionabstractPrevious studies have introduced a weaklysupervised paradigm for solving math word problems requiring only the answer value annotation.While these methods search for correct value equation candidates as pseudo labels, they search among a narrow sub-space of the enormous equation space.To address this problem, we propose a novel search algorithm with combinatorial strategy ComSearch, which can compress the search space by excluding mathematically equivalent equations.The compression allows the searching algorithm to enumerate all possible equations and obtain high-quality data.We investigate the noise in the pseudo labels that hold wrong mathematical logic, which we refer to as the false-matching problem, and propose a ranking model to denoise the pseudo labels.Our approach holds a flexible framework to utilize two existing supervised math word problem solvers to train pseudo labels, and both achieve state-of-the-art performance in the weak supervision task. 1 Qianying Liu, Wenyu Guan, Jianhao Shen, Fei Cheng 0002, Sadao Kurohashi |
EACL | 4 |
| 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsabstractIn spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fullysupervised baselines (e.g., fine-tuned BERT) in relation extraction (RE).This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness.In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic.We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1.Specifically, GPT-RE achieves SOTA performances on the Semeval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets.Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other predefined labels, is substantially alleviated by our method.We show an empirical analysis.1 Fei Cheng 0002, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li 0001, Sadao Kurohashi |
EMNLP | 2 |
| 2023 | Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech RecognitionabstractLow-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit comparable term frequency distributions, which enables us to construct a Huffman tree for performing multilingual hierarchical Softmax decoding. This hierarchical structure enables cross-lingual knowledge sharing among similar tokens, thereby enhancing low-resource training outcomes. Empirical analyses demonstrate that our method is effective in improving the accuracy and efficiency of low-resource speech recognition. Qianying Liu, Zhuo Gong, Zhengdong Yang, Sheng Li 0010, Chenchen Ding, Nobuaki Minematsu, Hao Huang 0009, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ICASSP | 9 |
| 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation ExtractionabstractRelation extraction (RE) has achieved remarkable progress with the help of pre-trained language models.However, existing RE models are usually incapable of handling two situations: implicit expressions and long-tail relation types, caused by language complexity and data sparsity.In this paper, we introduce a simple enhancement of RE using k nearest neighbors (kNN-RE).kNN-RE allows the model to consult training relations at test time through a nearest-neighbor search and provides a simple yet effective means to tackle the two issues above.Additionally, we observe that kNN-RE serves as an effective way to leverage distant supervision (DS) data for RE.Experimental results show that the proposed kNN-RE achieves state-of-the-art performances on a variety of supervised RE datasets, i.e., ACE05, SciERC, and Wiki80, along with outperforming the best model to date on the i2b2 and Wiki80 datasets in the setting of allowing using DS.Our code and models are available at: https://github.com/YukinoWan/kNN-RE. Qianying Liu, Zhuoyuan Mao, Fei Cheng 0002, Sadao Kurohashi, Jiwei Li 0001 |
EMNLP | 4 |
| 2022 | JaMIE: A Pipeline Japanese Medical Information Extraction System with Novel Relation AnnotationabstractIn the field of Japanese medical information extraction, few analyzing tools are available and relation extraction is still an under-explored topic. In this paper, we first propose a novel relation annotation schema for investigating the medical and temporal relations between medical entities in Japanese medical reports. We experiment with the practical annotation scenarios by separately annotating two different types of reports. We design a pipeline system with three components for recognizing medical entities, classifying entity modalities, and extracting relations. The empirical results show accurate analyzing performance and suggest the satisfactory annotation quality, the superiority of the latest contextual embedding models. and the feasible annotation strategy for high-accuracy demand. Fei Cheng 0002, Shuntaro Yada, Ribeka Tanaka, Eiji Aramaki, Sadao Kurohashi |
LREC | 1 |
| 2022 | Improving Event Duration Question Answering by Leveraging Existing Temporal Information Extraction DataabstractUnderstanding event duration is essential for understanding natural language. However, the amount of training data for tasks like duration question answering, i.e., McTACO, is very limited, suggesting a need for external duration information to improve this task. The duration information can be obtained from existing temporal information extraction tasks, such as UDS-T and TimeBank, where more duration data is available. A straightforward two-stage fine-tuning approach might be less likely to succeed given the discrepancy between the target duration question answering task and the intermediary duration classification task. This paper resolves this discrepancy by automatically recasting an existing event duration classification task from UDS-T to a question answering task similar to the target McTACO. We investigate the transferability of duration information by comparing whether the original UDS-T duration classification or the recast UDS-T duration question answering can be transferred to the target task. Our proposed model achieves a 13% Exact Match score improvement over the baseline on the McTACO duration question answering task, showing that the two-stage fine-tuning approach succeeds when the discrepancy between the target and intermediary tasks are resolved. Felix Giovanni Virgo, Fei Cheng 0002, Sadao Kurohashi |
LREC | 2 |
| 2022 | RODA: Reverse Operation Based Data Augmentation for Solving Math Word ProblemsabstractAutomatically solving math word problems is a critical task in the field of natural language processing. Recent models have reached their performance bottleneck and require more high-quality data for training. We propose a novel data augmentation method that reverses the mathematical logic of math word problems to produce new high-quality math problems and introduce new knowledge points that can benefit learning the mathematical reasoning logic. We apply the augmented data on two SOTA math word problem solving models and compare our results with a strong data augmentation baseline. Experimental results show the effectiveness of our approach (we release our code and data athttps://github.com/yiyunya/RODA). Qianying Liu, Wenyu Guan, Sujian Li, Fei Cheng 0002, Daisuke Kawahara, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Dependency Enhanced Contextual Representations for Japanese Temporal Relation Classification
Chenjing Geng, Fei Cheng 0002, Masayuki Asahara, Lis Pereira, Ichiro Kobayashi 0001 |
PACLIC | 2 |
| 2021 | ALICE++: Adversarial Training for Robust and Effective Temporal Reasoning
Lis Pereira, Fei Cheng 0002, Masayuki Asahara, Ichiro Kobayashi 0001 |
PACLIC | 2 |
| 2020 | Towards a Versatile Medical-Annotation Guideline Feasible Without Heavy Medical Knowledge: Starting From Critical Lung DiseasesabstractApplying natural language processing (NLP) to medical and clinical texts can bring important social benefits by mining valuable information from unstructured text. A popular application for that purpose is named entity recognition (NER), but the annotation policies of existing clinical corpora have not been standardized across clinical texts of different types. This paper presents an annotation guideline aimed at covering medical documents of various types such as radiography interpretation reports and medical records. Furthermore, the annotation was designed to avoid burdensome requirements related to medical knowledge, thereby enabling corpus development without medical specialists. To achieve these design features, we specifically focus on critical lung diseases to stabilize linguistic patterns in corpora. After annotating around 1100 electronic medical records following the annotation scheme, we demonstrated its feasibility using an NER task. Results suggest that our guideline is applicable to large-scale clinical NLP projects. Shuntaro Yada, Ayami Joh, Ribeka Tanaka, Fei Cheng 0002, Eiji Aramaki, Sadao Kurohashi |
LREC | 4 |
| 2018 | Inducing Temporal Relations from Time Anchor AnnotationabstractRecognizing temporal relations among events and time expressions has been an essential but challenging task in natural language processing.Conventional annotation of judging temporal relations puts a heavy load on annotators.In reality, the existing annotated corpora include annotations on only "salient" event pairs, or on pairs in a fixed window of sentences.In this paper, we propose a new approach to obtain temporal relations from absolute time value (a.k.a.time anchors), which is suitable for texts containing rich temporal information such as news articles.We start from time anchors for events and time expressions, and temporal relation annotations are induced automatically by computing relative order of two time anchors.This proposal shows several advantages over the current methods for temporal relation annotation: it requires less annotation effort, can induce inter-sentence relations easily, and increases informativeness of temporal relations.We compare the empirical statistics and automatic recognition results with our data against a previous temporal relation corpus.We also reveal that our data contributes to a significant improvement of the downstream time anchor prediction task, demonstrating 14.1 point increase in overall accuracy. Fei Cheng 0002, Yusuke Miyao |
NAACL-HLT | 1 |
| 2018 | Automatic Error Correction on Japanese Functional Expressions Using Character-based Neural Machine Translation
Fei Cheng 0002, Yiran Wang 0006, Hiroyuki Shindo, Yuji Matsumoto 0001 |
PACLIC | 2 |
| 2015 | A Hybrid Ranking Approach to Chinese Spelling CheckabstractWe propose a novel framework for Chinese Spelling Check (CSC), which is an automatic algorithm to detect and correct Chinese spelling errors. Our framework contains two key components: candidate generation and candidate ranking . Our framework differs from previous research, such as Statistical Machine Translation (SMT) based model or Language Model (LM) based model, in that we use both SMT and LM models as components of our framework for generating the correction candidates, in order to obtain maximum recall; to improve the precision, we further employ a Support Vector Machines (SVM) classifier to rank the candidates generated by the SMT and the LM. Experiments show that our framework outperforms other systems, which adopted the same or similar resources as ours in the SIGHAN 7 shared task; even comparing with the state-of-the-art systems, which used more resources, such as a considerable large dictionary, an idiom dictionary and other semantic information, our framework still obtains competitive results. Furthermore, to address the resource scarceness problem for training the SMT model, we generate around 2 million artificial training sentences using the Chinese character confusion sets, which include a set of Chinese characters with similar shapes and similar pronunciations, provided by the SIGHAN 7 shared task. Xiaodong Liu 0003, Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2014 | Parsing Chinese Synthetic Words with a Character-based Dependency Model
Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001 |
LREC | 1 |