EDBT 2026 Demo / reviewers in the wild / expert
Eiji Aramaki
dblp:02/2649
· DBLP profile ↗
45ranked-venue papers
7as first author
15since 2021 · last 2027
0000-0003-0201-3609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | RED: Retrieval-enhanced knowledge distillation for smaller language modelsabstractWhile large language models (LLMs) excel in open-domain question answering (QA) through in-context learning (ICL), their performance in domain-specific QA, where specialized knowledge is essential, remains suboptimal. Despite the technical feasibility of fine-tuning LLMs, the prohibitive computational cost during training and inference limits their practicality. Therefore, a more practical approach under limited computational resources is to enhance smaller language models for domain-specific QA through knowledge distillation (KD), leveraging both external domain knowledge and the knowledge embedded in LLMs. KD aims to improve student model performance by guiding it to mimic the prediction process of teacher models. With the rise of LLMs, student models can further benefit from chain of thought (CoT) rationales generated by LLMs. However, while LLMs contain extensive general knowledge, they often lack the specialized expertise required for domain-specific tasks that demand deep domain knowledge. To address this, we propose a novel KD scenario: R etrieval- E nhanced knowledge D istillation (RED), which expands the knowledge sources in KD by incorporating external knowledge. While external knowledge can enhance the model’s performance, retrieved text does not always perfectly align with the information needed to answer a given question due to the limitations in retrieval techniques and the content of external sources. Therefore, the key challenge in the retrieval-enhanced KD scenario is to effectively extract useful knowledge from external text while minimizing the impact of irrelevant information. In the proposed RED framework, LLMs are required not only to generate rationale for reasoning but also to evaluate the utility of retrieved text in facilitating reasoning. Additionally, we introduce compartment, filter, and traction mechanisms to mitigate the challenges associated with external text insertion. Experimental results demonstrate that our approach achieves consistent improvements for smaller models across four datasets in the scientific and biomedical domains. Particularly, under the perturbation setting where the retrieved text and the input text are mismatched, our method outperforms the competitive baseline of LLM inference by 9.14%, demonstrating its robustness. Xinbai Li, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki |
Expert Syst. Appl. | 4 |
| 2026 | Estimating Shared Mental Models via Communication-Categorized Directed GraphsabstractCorporate organizations face increasingly complex tasks that demand effective team management. A key concept is the Shared Mental Model (SMM), which enables members to maintain performance despite limited communication. Traditional measurements rely on interviews or questionnaires, which are labor-intensive, context-specific, and unsuitable for continuous monitoring. Consequently, leaders lack practical tools to track shared cognition in real time. This paper’s empirical analysis shows that only specific categories (e.g., informative exchanges) correlate strongly with SMM, clarifying which forms of communication can influence shared cognition. This insight leads to our proposed approach, which estimates SMM from instant messaging systems like Slack. Our approach categorizes messages into communicative acts using large language models, constructs category-wise communication graphs, and applies a graph neural network for estimation. The model outperforms baselines, demonstrating the feasibility of continuous, scalable monitoring without intrusive surveys. While validated in corporate contexts, the approach extends to education, healthcare, and disaster response domains. Wataru Yamada, Keiichi Ochiai, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki |
CHI | 6 |
| 2026 | NAIST LIFE STORY: A Seven-Year Crowdsourced Dataset of Japanese Emotion-related EpisodesabstractExisting emotion datasets have supported a wide range of NLP tasks, but most are static resources that capture language use only at the time of their creation. As a result, they cannot represent how emotional meanings shift in response to cultural and social change. To address this limitation, we present NAIST LIFE STORY, a seven-year collection of Japanese emotion-related episodes that reflect contemporary topics across multiple years. Since 2017, 1,000 crowdsourced participants per quarter have written short texts describing personal experiences associated with seven emotions: anger, anxiety, disgust, trust, joy, sadness, and surprise. The dataset currently spans 28 periods and includes gender and age information for each participant. Analyses reveal systematic differences in text length and lexical diversity across emotions, as well as clear temporal trends linked to major events such as the COVID-19 pandemic. A preliminary experiment with a large language model shows that using this dataset as contextual evidence improves time-aware emotion inference, demonstrating its value for studying the evolving relationship between emotion and language. Kazuhiro Ito, Junko Hayashi, Hiroyuki Nagai, Shoko Wakamiya, Eiji Aramaki |
LREC | 5 |
| 2026 | JPPB: Automatic Construction of a Soft-Labeled Japanese Patient Phrase Bank for Symptom NormalizationabstractPatient-generated symptom expressions are linguistically diverse, often deviating from standardized medical terminology. This paper introduces the Japanese Patient Phrase Bank (JPPB), the first automatically constructed phrase-level normalization resource for Japanese patient language. JPPB introduces an embedding-based soft labeling framework that transforms traditional one-to-one dictionary mappings into graded and ambiguity-aware associations. This framework represents a shift from word-level to phrase-level normalization in Japanese. The resource covers 7,035 phrase–term pairs across 412 symptoms. Evaluation on the KEEPHA and MedNLP-SC datasets shows that soft labels consistently improve Top-1 accuracy and better approximate gold label distributions compared with hard labels. While LLM-based normalization achieved the highest scores, JPPB provides a lightweight and transparent alternative suitable for local deployment. This work demonstrates that large-scale, automatically generated phrase banks can achieve competitive performance relative to manually curated resources and serve as practical, scalable resources for medical natural language processing in Japanese. Tomohiro Nishiyama, Mana Kuramoto, Shoko Wakamiya, Eiji Aramaki |
LREC | 4 |
| 2026 | J-ClinicalBench: A Benchmark for Evaluating Large Language Models on Practical Clinical Tasks in JapaneseabstractRecent advances in large language models (LLMs) have accelerated the NLP applications in the medical and clinical domains. However, evaluations remain limited for non-English languages, such as Japanese, where clinical corpora are particularly scarce. To address this gap, we present J-ClinicalBench, a publicly available benchmark designed to reflect realistic Japanese clinical tasks. We first created 227 expert-authored clinical documents and newly constructed five datasets for core clinical tasks. Building on these datasets, J-ClinicalBench comprises nine clinical tasks spanning clinical language reasoning, generation, and understanding. We establish baseline performance on J-ClinicalBench by evaluating state-of-the-art proprietary and Japanese open-source LLMs, providing the first assessment of their utility in practical clinical scenarios. By releasing this benchmark, we aim to foster the development and evaluation of clinically applicable LLMs in Japanese healthcare, bridging the current gap between clinical NLP research and clinical practice. Seiji Shimizu, Tomohiro Nishiyama, H. I S. A D. A Shohei, Himi Yamato, Shoko Wakamiya, Yuki Yanagisawa, Masami Tsuchiya, Satoko Hori, Eiji Aramaki |
LREC | 9 |
| 2025 | Investigating Neurons and Heads in Transformer-based LLMs for Typographical ErrorsabstractThis paper investigates how LLMs encode inputs with typos. We hypothesize that specific neurons and attention heads recognize typos and fix them internally using local and global contexts. We introduce a method to identify typo neurons and typo heads that work actively when inputs contain typos. Our experimental results suggest the following: 1) LLMs can fix typos with local contexts when the typo neurons in either the early or late layers are activated, even if those in the other are not. 2) Typo neurons in the middle layers are the core of typo-fixing with global contexts. 3) Typo heads fix typos by widely considering the context not focusing on specific tokens. 4) Typo neurons and typo heads work not only for typo-fixing but also for understanding general contexts. Kohei Tsuji, Tatsuya Hiraoka, Yuchang Cheng, Eiji Aramaki, Tomoya Iwakura |
EMNLP | 4 |
| 2025 | GenKP: generative knowledge prompts for enhancing large language modelsabstractLarge language models (LLMs) have demonstrated extensive capabilities across various natural language processing (NLP) tasks. Knowledge graphs (KGs) harbor vast amounts of facts, furnishing external knowledge for language models. The structured knowledge extracted from KGs must undergo conversion into sentences to align with the input format required by LLMs. Previous research has commonly utilized methods such as triple conversion and template-based conversion. However, sentences converted using existing methods frequently encounter issues such as semantic incoherence, ambiguity, and unnaturalness, which distort the original intent, and deviate the sentences from the facts. Meanwhile, despite the improvement that knowledge-enhanced pre-training and prompt-tuning methods have achieved in small-scale models, they are difficult to implement for LLMs in the absence of computational resources. The advanced comprehension of LLMs facilitates in-context learning (ICL), thereby enhancing their performance without the need for additional training. In this paper, we propose a knowledge prompts generation method, GenKP, which injects knowledge into LLMs by ICL. Compared to inserting triple-conversion or templated-conversion knowledge without selection, GenKP entails generating knowledge samples using LLMs in conjunction with KGs and makes a trade-off of knowledge samples through weighted verification and BM25 ranking, reducing knowledge noise. Experimental results illustrate that incorporating knowledge prompts enhances the performance of LLMs. Furthermore, LLMs augmented with GenKP exhibit superior improvements compared to the methods utilizing triple and template-based knowledge injection. Xinbai Li, Shaowen Peng, Shuntaro Yada, Shoko Wakamiya, Eiji Aramaki |
Appl. Intell. | 5 |
| 2024 | A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across LanguagesabstractUser-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages. Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas 0001, Tomohiro Nishiyama, Sebastian Möller 0001, Eiji Aramaki, Yuji Matsumoto 0001, Roland Roller, Pierre Zweigenbaum |
LREC/COLING | 11 |
| 2024 | QA-based Event Start-Points Ordering for Clinical Temporal Relation AnnotationabstractTemporal relation annotation in the clinical domain is crucial yet challenging due to its workload and the medical expertise required. In this paper, we propose a novel annotation method that integrates event start-points ordering and question-answering (QA) as the annotation format. By focusing only on two points on a timeline, start-points ordering reduces ambiguity and simplifies the relation set to be considered during annotation. QA as annotation recasts temporal relation annotation into a reading comprehension task, allowing annotators to use natural language instead of the formalisms commonly adopted in temporal relation annotation. Based on our method, most of the relations in a document are inferable from a significantly smaller number of explicitly annotated relations, showing the efficiency of our proposed method. Using these inferred relations, we develop a temporal relation classification model that achieves a 0.72 F1 score. Also, by decomposing the annotation process into QA generation and QA validation, our method enables collaboration among medical experts and non-experts. We obtained high inter-annotator agreement (IAA) scores, which indicate the positive prospect of such collaboration in the annotation process. Our annotated corpus, annotation tool, and trained model are publicly available: https://github.com/seiji-shimizu/qa-start-ordering. Seiji Shimizu, Lis Pereira, Shuntaro Yada, Eiji Aramaki |
LREC/COLING | 4 |
| 2023 | Comparative evaluation of boundary-relaxed annotation for Entity Linking performanceabstractEntity Linking performance has a strong reliance on having a large quantity of high-quality annotated training data available.Yet, manual annotation of named entities, especially their boundaries, is ambiguous, error-prone, and raises many inconsistencies between annotators.While imprecise boundary annotation can degrade a model's performance, there are applications where accurate extraction of entities' surface form is not necessary.For those cases, a lenient annotation guideline could relieve the annotators' workload and speed up the process.This paper presents a case study designed to verify the feasibility of such annotation process and evaluate the impact of boundary-relaxed annotation in an Entity Linking pipeline.We first generate a set of noisy versions of the widely used AIDA CoNLL-YAGO dataset by expanding the boundaries subsets of annotated entity mentions and then train three Entity Linking models on this data and evaluate the relative impact of imprecise annotation on entity recognition and disambiguation performances.We demonstrate that the magnitude of effects caused by noise in the Named Entity Recognition phase is dependent on both model complexity and noise ratio, while Entity Disambiguation components are susceptible to entity boundary imprecision due to strong vocabulary dependency. Gabriel Herman Bernardim Andrade, Shuntaro Yada, Eiji Aramaki |
ACL (1) | 3 |
| 2023 | Visualyre: multimodal album art generation for independent musiciansabstractAbstract Album art often reflects the trends and themes of the songs in a given collection, and even the identities of the musicians who produced it. It therefore plays a central role in fomenting a potential listener’s first impression of the work. As such, musicians strive to find suitable images for this purpose, and those with limited financial resources or design skills may struggle to do so. Here, we report the development of Visualyre, a deep learning–based application that generates album art images from users’ song lyrics and audio files. This tool relies on generative adversarial network models to generate images from textual input (lyrics) and style transfer models to adjust the image according to the mood of the audio. We then report the results of a user study involving 35 amateur and independent musicians who tested the system. Results suggest that Visualyre was generally well received and largely effective in its intended purpose: providing musicians with a resource for generating their own album art. Gamar Azuaje, Kongmeng Liew, Elena V. Epure, Shuntaro Yada, Shoko Wakamiya, Eiji Aramaki |
Pers. Ubiquitous Comput. | 6 |
| 2022 | Graph Neural Network Tells Us Who is the Communication EnhancerabstractMany managers and human resource departments transfer a person in an attempt to increase the performance of an organization or team. When we define performance as team efficiency, the performance is influenced by the density of team communications. However, whether or not the candidate transferee will actually increase the density of team communications is an unknown. In this paper, we propose a new approach that estimates whether or not a person who joins a team will improve the density of team communications based on an instant messaging system (IMS). In the proposed approach, we embed people in feature space using graph neural networks. In this embedding process, we do not use the content of text communications but utilize only communication graph architectural information that expresses who is talking to whom and how often. Additionally, the proposed approach does not require a questionnaire to indicate the density of team communication as in some previous studies. In the proposed approach, we develop a machine learning model classifying whether or not a transferee will improve the density of team communications. The model classifies transferees at an accuracy of 0.57 and precision of 0.58. Since this model does not use text contents from the IMS, it is valuable in actual business situations. Yuuma Jitsunari, Wataru Yamada, Keiichi Ochiai, Shoko Wakamiya, Eiji Aramaki |
IEEE Big Data | 6 |
| 2022 | JaMIE: A Pipeline Japanese Medical Information Extraction System with Novel Relation AnnotationabstractIn the field of Japanese medical information extraction, few analyzing tools are available and relation extraction is still an under-explored topic. In this paper, we first propose a novel relation annotation schema for investigating the medical and temporal relations between medical entities in Japanese medical reports. We experiment with the practical annotation scenarios by separately annotating two different types of reports. We design a pipeline system with three components for recognizing medical entities, classifying entity modalities, and extracting relations. The empirical results show accurate analyzing performance and suggest the satisfactory annotation quality, the superiority of the latest contextual embedding models. and the feasible annotation strategy for high-accuracy demand. Fei Cheng 0002, Shuntaro Yada, Ribeka Tanaka, Eiji Aramaki, Sadao Kurohashi |
LREC | 4 |
| 2022 | Annotation-Scheme Reconstruction for "Fake News" and Japanese Fake News DatasetabstractFake news provokes many societal problems; therefore, there has been extensive research on fake news detection tasks to counter it. Many fake news datasets were constructed as resources to facilitate this task. Contemporary research focuses almost exclusively on the factuality aspect of the news. However, this aspect alone is insufficient to explain “fake news,” which is a complex phenomenon that involves a wide range of issues. To fully understand the nature of each instance of fake news, it is important to observe it from various perspectives, such as the intention of the false news disseminator, the harmfulness of the news to our society, and the target of the news. We propose a novel annotation scheme with fine-grained labeling based on detailed investigations of existing fake news datasets to capture these various aspects of fake news. Using the annotation scheme, we construct and publish the first Japanese fake news dataset. The annotation scheme is expected to provide an in-depth understanding of fake news. We plan to build datasets for both Japanese and other languages using our scheme. Our Japanese dataset is published at https://hkefka385.github.io/dataset/fakenews-japanese/. Taichi Murayama, Shohei Hisada, Makoto Uehara, Shoko Wakamiya, Eiji Aramaki |
LREC | 5 |
| 2021 | Single Model for Influenza Forecasting of Multiple Countries by Multi-task Learning
Taichi Murayama, Shoko Wakamiya, Eiji Aramaki |
ECML/PKDD (4) | 3 |
| 2020 | Offensive Language Detection on Video Live Streaming ChatabstractThis paper presents a prototype of a chat room that detects offensive expressions in a video live streaming chat in real time.Focusing on Twitch, one of the most popular live streaming platforms, we created a dataset for the task of detecting offensive expressions.We collected 2,000 chat posts across four popular game titles with genre diversity (e.g., competitive, violent, peaceful).To make use of the similarity in offensive expressions among different social media platforms, we adopted state-of-the-art models trained on offensive expressions from Twitter for our Twitch data (i.e., transfer learning).We investigated two similarity measurements to predict the transferability, textual similarity, and game-genre similarity.Our results show that the transfer of features from social media to live streaming is effective.However, the two measurements show less correlation in the transferability prediction. Shuntaro Yada, Shoko Wakamiya, Eiji Aramaki |
COLING | 4 |
| 2020 | Towards a Versatile Medical-Annotation Guideline Feasible Without Heavy Medical Knowledge: Starting From Critical Lung DiseasesabstractApplying natural language processing (NLP) to medical and clinical texts can bring important social benefits by mining valuable information from unstructured text. A popular application for that purpose is named entity recognition (NER), but the annotation policies of existing clinical corpora have not been standardized across clinical texts of different types. This paper presents an annotation guideline aimed at covering medical documents of various types such as radiography interpretation reports and medical records. Furthermore, the annotation was designed to avoid burdensome requirements related to medical knowledge, thereby enabling corpus development without medical specialists. To achieve these design features, we specifically focus on critical lung diseases to stabilize linguistic patterns in corpora. After annotating around 1100 electronic medical records following the annotation scheme, we demonstrated its feasibility using an NER task. Results suggest that our guideline is applicable to large-scale clinical NLP projects. Shuntaro Yada, Ayami Joh, Ribeka Tanaka, Fei Cheng 0002, Eiji Aramaki, Sadao Kurohashi |
LREC | 5 |
| 2020 | Chat-type Manzai application: mobile daily comedy presentations based on automatically generated manzai scenariosabstractWe have proposed and demonstrate Manzai robots that automatically generate Manzai scenarios. Manzai is a Japanese traditional comedy consisting of two comedians with funny dialogues like stand up comedy. Our Manzai robots are a huge system, but it is not easy for users to watch Manzai every day using our Manzai robots. As described herein, we propose a mobile style of presenting our automatically generated Manzai anytime and anywhere using a web application. We designate the application as a Chat-type Manzai Application. The Chat-type Manzai Application is intended to make people healthier by laughing as they relax and watch the Manzai. However, the Chat-type Manzai Application loses direction. Therefore, we propose a new dialogue component: the Name list component. We used experiments of three types to measure the benefits of our proposed application and component. Kazuki Haraguchi, Kazuki Yane, Akira Sato, Eiji Aramaki, Isao Miyashiro, Akiyo Nadamoto |
MoMM | 4 |
| 2019 | Learning to Select, Track, and Generate for Data-to-TextabstractHayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi 0001, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura |
ACL (1) | 5 |
| 2019 | City Link: Finding Similar Areas in Two Cities Using Twitter Data
Wannita Takerngsaksiri, Shoko Wakamiya, Eiji Aramaki |
W2GIS | 3 |
| 2019 | Pleasant Route Suggestion based on Color and Object RatesabstractFor a tourist who wishes to stroll in an unknown city, it is useful to have a recommendation of not just the shortest routes but also routes that are pleasant. This paper demonstrates a system that provides pleasant route recommendation. Currently, we focus on routes that have much green and bright views. The system measures pleasure scores by extracting colors or objects in Google Street View panorama images and re-ranks shortest paths in the order of the computed pleasure scores. The current prototype provides route recommendation for city areas in Tokyo, Kyoto and San Francisco. Shoko Wakamiya, Panote Siriaraya, Yihong Zhang 0001, Yukiko Kawai, Eiji Aramaki, Adam Jatowt |
WSDM | 5 |
| 2018 | Wisdom in Adversity: A Twitter Study of the Japanese TsunamiabstractSophisticated data science techniques have recently been applied to social networks data to study social phenomena and people. Recognizing that social psychology research has witnessed a renewed interest in the notion of wisdom, with an emphasis to its contextual dimensions, this study looks at the expression of wisdom in twitter messages. Specifically, it examines the relation between wisdom in adversity and cultural influences using Twitter data from the tragic Japanese tsunami of 2011. The study employs natural language processing and data science to detect the expression of wisdom. Two categories for wisdom in adversity are used: recognition of uncertainty and change, and cognitive empathy. Data processing is applied to 1,000 annotated tweets and extended to 43,436 tweets. The results show that it is viable to study wisdom in context using social networking sites data. This short paper discusses some of the findings. Paolo Casani, Hayate Iso, Shoko Wakamiya, Eiji Aramaki |
ASONAM | 4 |
| 2018 | How Information Sharing about Care Recipients by Family Caregivers Impacts Family CommunicationabstractPrevious research has shown that tracking technologies have the potential to help family caregivers optimize their coping strategies and improve their relationships with care recipients. In this paper, we explore how sharing the tracked data (i.e., caregiving journals and patient's conditions) with other family caregivers affects home care and family communication. Although previous works suggested that family caregivers may benefit from reading the records of others, sharing patients' private information might fuel negative feelings of surveillance and violation of trust for care recipients. To address this research question, we added a sharing feature to the previously developed tracking tool and deployed it for six weeks in the homes of 15 family caregivers who were caring for a depressed family member. Our findings show how the sharing feature attracted the attention of care recipients and helped the family caregivers discuss sensitive issues with care recipients. Naomi Yamashita, Hideaki Kuzuoka, Takashi Kudo, Keiji Hirata 0001, Eiji Aramaki, Kazuki Hattori |
CHI | 5 |
| 2018 | J-MeDic: A Japanese Disease Name Dictionary based on Real Clinical Usage
Kaoru Ito, Hiroyuki Nagai, Taro Okahisa, Shoko Wakamiya, Tomohide Iwao, Eiji Aramaki |
LREC | 6 |
| 2017 | Changing Moods: How Manual Tracking by Family Caregivers Improves Caring and Family CommunicationabstractPrevious research on healthcare technologies has shown how health tracking promotes desired behavior changes and effective health management. However, little is known about how the family caregivers' use of tracking technologies impacts the patient-caregiver relationship in the home. In this paper, we explore how health-tracking technologies could be designed to support family caregivers cope better with a depressed family member. Based on an interview study, we designed a simple tracking tool called Family Mood and Care Tracker (FMCT) and deployed it for six weeks in the homes of 14 family caregivers who were caring for a depressed family member. FMCT is a tracking tool designed specifically for family caregivers to record their caregiving activities and patient's conditions. Our findings demonstrate how caregivers used it to better understand the illness and cope with depressed family members. We also show how our tool improves family communication, despite the initial concerns about patient-caregiver conflicts. Naomi Yamashita, Hideaki Kuzuoka, Keiji Hirata 0001, Takashi Kudo, Eiji Aramaki, Kazuki Hattori |
CHI | 5 |
| 2016 | Forecasting Word Model: Twitter-based Influenza Surveillance and PredictionabstractBecause of the increasing popularity of social media, much information has been shared on the internet, enabling social media users to understand various real world events. Particularly, social media-based infectious disease surveillance has attracted increasing attention. In this work, we specifically examine influenza: a common topic of communication on social media. The fundamental theory of this work is that several words, such as symptom words (fever, headache, etc.), appear in advance of flu epidemic occurrence. Consequently, past word occurrence can contribute to estimation of the number of current patients. To employ such forecasting words, one can first estimate the optimal time lag for each word based on their cross correlation. Then one can build a linear model consisting of word frequencies at different time points for nowcasting and for forecasting influenza epidemics. Experimentally obtained results (using 7.7 million tweets of August 2012 – January 2016), the proposed model achieved the best nowcasting performance to date (correlation ratio 0.93) and practically sufficient forecasting performance (correlation ratio 0.91 in 1-week future prediction, and correlation ratio 0.77 in 3-weeks future prediction). This report is the first of the relevant literature to describe a model enabling prediction of future epidemics using Twitter. Hayate Iso, Shoko Wakamiya, Eiji Aramaki |
COLING | 3 |
| 2016 | Lets not stare at smartphones while walking: memorable route recommendation by detecting effective landmarksabstractNavigation in unfamiliar cities often requires frequent map checking, which is troublesome for wayfinders. We propose a novel approach for improving real-world navigation by generating short, memorable and intuitive routes. To do so we detect useful landmarks for effective route navigation. This is done by exploiting not only geographic data but also crowd footprints in Social Network Services (SNS) and Location Based Social Networks (LBSN). Specifically, we detect point, area, and line landmarks by using three indicators to measure landmark's utility: visit popularity, direct visibility, and indirect visibility. We then construct an effective route graph based on the extracted landmarks, which facilitates optimal path search. In the experiments, we show that landmark-based routes out-perform the ones created by baseline from the perspectives of the lap time and the number of references necessary to check self-positions for adjusting route directions. Shoko Wakamiya, Hiroshi Kawasaki, Yukiko Kawai, Adam Jatowt, Eiji Aramaki, Toyokazu Akiyama |
UbiComp | 5 |
| 2015 | Who caught a cold ? - Identifying the subject of a symptomabstractShin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa 0004 |
ACL (1) | 4 |
| 2014 | Comparative Analysis of Sizzle Words on the InternetabstractPeople post their impressions of foods on Twitter in real time after eating. On many user-generated recipe sites on the Internet, ordinary cooks such as homemakers can post their original recipes. They make full use of the recipe titles that incorporate tasty words such as authentic, homey, and spicy because they want to their recipe to become popular. Web sites of companies and restaurants use tasty words such as 'healthy' and 'old-fashioned' to sell their products. In this way, tasty words of many kinds have come to be used on the Internet. We consider that tasty words differ among those used on Internet media which are Twitter, user-generated recipe sites, and ordinary web sites. As described herein, we designate such tasty words as 'Sizzle Words' and then compare these three media using the Sizzle Words. Daisuke Kato, Mai Miyabe, Eiji Aramaki, Akiyo Nadamoto |
iiWAS | 3 |
| 2013 | Classification and characterization of clinical finding expressions in medical literatureabstractText mining of clinical findings has been employed to extract clinical information contained in “electronic medical records” without the need for labor intensive work by medical experts. However, the automated building of disease ontology necessitates knowledge acquisition of clinical findings documented in “medical literature” that requires an independent strategy. This study performs a preliminary analysis of clinical finding expressions in medical literature to enable the automated acquisition of disease knowledge. To this end, we selected descriptions of 20 diseases in a free-text format and annotated the texts to extract expressions of clinical findings. This resulted in 1368 expressions with varying lengths and syntactic features, and 161 annotator comments. The comments suggested that certain types of expressions, which were further classified into 10 categories. Also, in-depth analyses of their syntactic and semantic characteristics were performed, resulting in the following observations. First, expressions of clinical findings have certain patterns, syntactic and semantic, which can be exploited for appropriate knowledge acquisition. Second, clinical knowledge may guide the knowledge acquisition process in a top-down manner. Third, natural language processing of medical literature requires specific considerations compared with the processing of health records, namely, i) distinction of subjects, ii) handling of generalized knowledge, and iii) processing of expressions for examination results. This preliminary survey on the expressions in medical literature provides helpful insights for future corpus design. Takashi Okumura, Yuka Tateisi, Eiji Aramaki |
BIBM | 3 |
| 2013 | Analysis of Microblog Rumors and Correction Texts for Disaster SituationsabstractMicroblogging systems such as Twitter have become popular. They are especially useful and helpful for users in disaster situations. Microblogs have facilitated the spread of information of all kinds, even rumors. Rumors block adequate information sharing and cause severe problems. Several studies have analyzed rumors, but it remains unclear how rumors are spread on microblogging systems. As described in this paper, we present a case study of how rumors are spread on Twitter in a recent disaster situation, that of the Great East Japan earthquake in March 11 2011, based on comparison to a normal situation. We also specifically examine the correction of rumors because automatic extraction of rumors is difficult, but extracting rumor-correction is easier than extracting the rumors themselves. We (1) classify tweets in disaster situations, (2) analyze tweets in disaster situations based on user's impression, and (3) compare the spread of rumor tweets in a disaster situation to that in a normal situation. Akiyo Nadamoto, Mai Miyabe, Eiji Aramaki |
iiWAS | 3 |
| 2013 | Word in a Dictionary is used by Numerous Users
Eiji Aramaki, Sachiko Maskawa, Mai Miyabe, Mizuki Morita, Sachi Yasuda |
IJCNLP | 1 |
| 2012 | Which is Stronger? : Discriminative Learning of Sound Symbolism
Eiji Aramaki, Sachi Yasuda, Mai Miyabe, Satoshi Miura, Masaki Murata |
CogSci | 1 |
| 2012 | Ad hoc creature: Lost and added in translation from description to depiction
Sachi Yasuda, Masashi Okamoto, Eiji Aramaki |
CogSci | 3 |
| 2011 | Twitter Catches The Flu: Detecting Influenza Epidemics using Twitter
Eiji Aramaki, Sachiko Maskawa, Mizuki Morita |
EMNLP | 1 |
| 2010 | Outline of Community-Type Content Based on Wikipedia
Akiyo Nadamoto, Eiji Aramaki, Takeshi Abekawa, Yohei Murakami |
DASFAA (2) | 2 |
| 2010 | Extracting the gist of social network services using WikipediaabstractSocial Network Services(SNSs), which are maintained by a community of people, are among the popular Web 2.0 tools. Multiple users freely post their comments to an SNS thread. It is difficult to understand the gist of the comments because the dialog in an SNS thread is complicated. In this paper, we propose a system that presents the gist of information at a glance and basic information about an SNS thread by using Wikipedia. We focus on the table of contents (TOC) of the relevant articles on Wikipedia. Our system compares the comments in a thread with the information in the TOC and identifies contents that are similar. We consider the similar contents in the TOC as the gist of the thread and paragraphs in Wikipedia similar to the comments in the thread as comprising basic information about the thread. Thus, a user can obtain the gist of an SNS thread by viewing a table with similar contents. Akiyo Nadamoto, Eiji Aramaki, Takeshi Abekawa, Yohei Murakami |
iiWAS | 2 |
| 2010 | Using Various Features in Machine Learning to Obtain High Levels of Performance for Recognition of Japanese Notational Variants
Masahiro Kojima, Masaki Murata, Jun'ichi Kazama, Kow Kuroda, Atsushi Fujita, Eiji Aramaki, Masaaki Tsuchida, Yasuhiko Watanabe, Kentaro Torisawa |
PACLIC | 6 |
| 2009 | Content hole search in community-type content using WikipediaabstractSNSs and blogs, both of which are maintained by a community of people, have become popular in Web 2.0. We call these content as "Community-type content." This community is associated with the content, and those who use or contribute to community-type content are considered as members of the community. Occasionally, the members of a community do not understand the theme of the content from multiple viewpoints, hence, the amount of information is often insufficient. It is convenient to present the user missed information. In this way, when Web 2.0 became popular, the content on the Internet and type of users are changed. We believe that there is a need for next-generation search engines in Web 2.0. We require a search engine that can search for information users are unaware of; we call such information as "content holes." In this paper, we propose a method for searching content holes in community-type content. We attempt to extract and represent content holes from discussions on SNSs and blogs. Conventional Web search technique is generally based on similarities. On the other hand, our content-hole search is a different search. In this paper, we classify and represent a number of images for different searching methods; we define content holes and as the first step toward realizing our aim, we propose a content-hole search system using Wikipedia. Akiyo Nadamoto, Eiji Aramaki, Takeshi Abekawa, Yohei Murakami |
iiWAS | 2 |
| 2009 | Content hole search in community-type contentabstractIn community-type content such as blogs and SNSs, we call the user's unawareness of information as a "content hole" and the search for this information as a "content hole search." A content hole search differs from similarity searching and has a variety of types. In this paper, we propose different types of content holes and define each type. We also propose an analysis of dialogue related to community-type content and introduce content hole search by using Wikipedia as an example. Akiyo Nadamoto, Eiji Aramaki, Takeshi Abekawa, Yohei Murakami |
WWW | 2 |
| 2008 | Orthographic Disambiguation Incorporating Transliterated Probability
Eiji Aramaki, Takeshi Imai, Kengo Miyo, Kazuhiko Ohe |
IJCNLP | 1 |
| 2005 | Probabilistic Model for Example-based Machine TranslationabstractExample-based machine translation (EBMT) systems, so far, rely on heuristic measures in retrieving translation examples. Such a heuristic measure costs time to adjust, and might make its algorithm unclear. This paper presents a probabilistic model for EBMT. Under the proposed model, the system searches the translation example combination which has the highest probability. The proposed model clearly formalizes EBMT process. In addition, the model can naturally incorporate the context similarity of translation examples. The experimental results demonstrate that the proposed model has a slightly better translation quality than state-of-the-art EBMT systems. Eiji Aramaki, Sadao Kurohashi, Hideki Kashioka, Naoto Katoh |
MTSummit | 1 |
| 2004 | Example-Based Machine Translation Without Saying Inferable Predicate
Eiji Aramaki, Sadao Kurohashi, Hideki Kashioka, Hideki Tanaka |
IJCNLP | 1 |
| 2001 | Finding translation correspondences from parallel parsed corpus for example-based translationabstractThis paper describes a system for finding phrasal translation correspondences from parallel parsed corpus that are collections paired English and Japanese sentences. First, the system finds phrasal correspondences by Japanese-English translation dictionary consultation. Then, the system finds correspondences in remaining phrases by using sentences dependency structures and the balance of all correspondences. The method is based on an assumption that in parallel corpus most fragments in a source sentence have corresponding fragments in a target sentence. Eiji Aramaki, Sadao Kurohashi, Satoshi Sato, Hideo Watanabe |
MTSummit | 1 |
| 2000 | Finding Structural Correspondences from Bilingual Parsed Corpus for Corpus-based Translation
Hideo Watanabe, Sadao Kurohashi, Eiji Aramaki |
COLING | 3 |