Teruko Mitamura

dblp:90/785 · DBLP profile ↗
← Back
63ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Databases, data management, data science and information retrieval · 8Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
abstract
Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system development, we propose a new multimodal QA evaluation dataset on assembly activities. Our dataset, ProMQA-Assembly, consists of 646 QA pairs that require multimodal understanding of human activity videos and their instruction manuals in an online-style manner. For cost effectiveness in the data creation, we adopt a semi-automated QA annotation approach, where LLMs generate candidate QA pairs and humans verify them. We further improve QA generation by integrating fine-grained action labels to diversify question types. Additionally, we create 81 instruction task graphs for our target assembly tasks. These newly created task graphs are used in our benchmarking experiment, as well as in facilitating the human verification process. With our dataset, we benchmark models, including competitive proprietary multimodal models. We find that ProMQA-Assembly contains challenging multimodal questions, where reasoning models showcase promising results. We believe our new evaluation dataset contributes to the further development of procedural-activity assistants.
Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm, Xuanang Zhou, Ken Fukuda, Teruko Mitamura
LREC8
2026 VDAct 2.0: Scaling Video-Grounded Dialogue for Event-driven Activity Understanding with LLM-Assisted Filtering
Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Ken Fukuda, Teruko Mitamura
LREC5
2025 A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
abstract
This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and more complex video sequences that depict a variety of event-driven activities that require advanced contextual understanding for accurate response generation. The dataset comprises 3,000 dialogues with over 30,000 question-and-answer pairs, derived from 1,000 videos with diverse activity scenarios. VDAct displays a notably challenging characteristic due to its broad spectrum of activity scenarios and wide range of question types. Empirical studies on state-of-the-art vision foundation models highlight their limitations in addressing certain question types on our dataset. Furthermore, VDEval, which integrates dialogue session history and video content summaries extracted from our supplementary Knowledge Graphs to evaluate individual responses, demonstrates a significantly higher correlation with human assessments on the VDAct dataset than existing evaluation metrics that rely solely on the context of single dialogue turns.
Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Zhi-Qi Cheng, Ken Fukuda, Teruko Mitamura
AAAI6
2025 ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
abstract
Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Yuran Wang, Ken Fukuda, Teruko Mitamura. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Ken Fukuda, Teruko Mitamura
NAACL (Long Papers)8
2025 What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
abstract
Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LoGra that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LogIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LoGra achieves competitive accuracy against more expensive baselines while showing up to 6,500x improvement in throughput and 5x reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset.
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger B. Grosse, Eric P. Xing
NeurIPS9
2024 Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
abstract
Vision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-and-Language Navigation (HA-VLN), extending traditional VLN by incorporating dynamic human activities and relaxing key assumptions. We propose the Human-Aware 3D (HA3D) simulator, which combines dynamic human activities with the Matterport3D dataset, and the Human-Aware Room-to-Room (HA-R2R) dataset, extending R2R with human activity descriptions. To tackle HA-VLN challenges, we present the Expert-Supervised Cross-Modal (VLN-CM) and Non-Expert-Supervised Decision Transformer (VLN-DT) agents, utilizing cross-modal fusion and diverse training strategies for effective navigation in dynamic human environments. A comprehensive evaluation, including metrics considering human activities, and systematic analysis of HA-VLN's unique challenges, underscores the need for further research to enhance HA-VLN agents' real-world robustness and adaptability. Ultimately, this work provides benchmarks and insights for future research on embodied AI and Sim2Real transfer, paving the way for more realistic and applicable VLN systems in human-populated environments.
Zhi-Qi Cheng, Yifei Dong 0002, Yuxuan Zhou 0004, Jun-Yan He, Qi Dai 0001, Teruko Mitamura, Alex Hauptmann 0001
NeurIPS8
2023 Hierarchical Event Grounding
abstract
Event grounding aims at linking mention references in text corpora to events from a knowledge base (KB). Previous work on this task focused primarily on linking to a single KB event, thereby overlooking the hierarchical aspects of events. Events in documents are typically described at various levels of spatio-temporal granularity. These hierarchical relations are utilized in downstream tasks of narrative understanding and schema construction. In this work, we present an extension to the event grounding task that requires tackling hierarchical event structures from the KB. Our proposed task involves linking a mention reference to a set of event labels from a subevent hierarchy in the KB. We propose a retrieval methodology that leverages event hierarchy through an auxiliary hierarchical loss. On an automatically created multilingual dataset from Wikipedia and Wikidata, our experiments demonstrate the effectiveness of the hierarchical loss against retrieve and re-rank baselines. Furthermore, we demonstrate the systems' ability to aid hierarchical discovery among unseen events. Code is available at https://github.com/JefferyO/Hierarchical-Event-Grounding
Jiefu Ou, Adithya Pratapa, Rishubh Gupta, Teruko Mitamura
AAAI4
2022 Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models
abstract
We investigate the use of multimodal information contained in images as an effective method for enhancing the commonsense of Transformer models for text generation. We perform experiments using BART and T5 on concept-to-text generation, specifically the task of generative commonsense reasoning, or CommonGen. We call our approach VisCTG: Visually Grounded Concept-to-Text Generation. VisCTG involves captioning images representing appropriate everyday scenarios, and using these captions to enrich and steer the generation process. Comprehensive evaluation and analysis demonstrate that VisCTG noticeably improves model performance while successfully addressing several issues of the baseline generations, including poor commonsense, fluency, and specificity.
Steven Y. Feng, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy, Varun Gangal
AAAI5
2022 NAREOR: The Narrative Reordering Problem
abstract
Many implicit inferences exist in text depending on how it is structured that can critically impact the text's interpretation and meaning. One such structural aspect present in text with chronology is the order of its presentation. For narratives or stories, this is known as the narrative order. Reordering a narrative can impact the temporal, causal, event-based, and other inferences readers draw from it, which in turn can have strong effects both on its interpretation and interestingness. In this paper, we propose and investigate the task of Narrative Reordering (NAREOR) which involves rewriting a given story in a different narrative order while preserving its plot. We present a dataset, NAREORC, with human rewritings of stories within ROCStories in non-linear orders, and conduct a detailed analysis of it. Further, we propose novel task-specific training methods with suitable evaluation metrics. We perform experiments on NAREORC using state-of-the-art models such as BART and T5 and conduct extensive automatic and human evaluations. We demonstrate that although our models can perform decently, NAREOR is a challenging task with potential for further exploration. We also investigate two applications of NAREOR: generation of more interesting variations of stories and serving as adversarial sets for temporal/event-related tasks, besides discussing other prospective ones, such as for pedagogical setups related to language skills like essay writing and applications to medicine involving clinical narratives.
Varun Gangal, Steven Y. Feng, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy
AAAI4
2022 PRO-CS : An Instance-Based Prompt Composition Technique for Code-Switched Tasks
abstract
Code-switched (CS) data is ubiquitous in today's globalized world, but the dearth of annotated datasets in code-switching poses a significant challenge for learning diverse tasks across different language pairs.Parameter-efficient prompt-tuning approaches conditioned on frozen language models have shown promise for transfer learning in limited-resource setups.In this paper, we propose a novel instancebased prompt composition technique, PRO-CS, for CS tasks that combine language and task knowledge.We compare our approach with prompt-tuning and fine-tuning for codeswitched tasks on 10 datasets across 4 language pairs.Our model outperforms the prompttuning approach by significant margins across all datasets and outperforms or remains at par with fine-tuning by using just 0.18% of total parameters.We also achieve competitive results when compared with the fine-tuned model in the low-resource cross-lingual and crosstask setting, indicating the effectiveness of our approach to incorporate new code-switched tasks.
Srijan Bansal, Suraj Tripathi, Sumit Agarwal, Teruko Mitamura, Eric Nyberg
EMNLP4
2022 GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement
abstract
Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like'' event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but also predicts all corresponding semantic roles (e.g. agent and goods). Inspired by object detection and image captioning tasks, existing methods typically employ a two-stage framework: 1) detect the activity verb, and then 2) predict semantic roles based on the detected verb. Obviously, this illogical framework constitutes a huge obstacle to semantic understanding. First, pre-detecting verbs solely without semantic roles inevitably fails to distinguish many similar daily activities (e.g., offering and giving, buying and selling). Second, predicting semantic roles in a closed auto-regressive manner can hardly exploit the semantic relations among the verb and roles. To this end, in this paper we propose a novel two-stage framework that focuses on utilizing such bidirectional relations within verbs and roles. In the first stage, instead of pre-detecting the verb, we postpone the detection step and assume a pseudo label, where an intermediate representation for each corresponding semantic role is learned from images. In the second stage, we exploit transformer layers to unearth the potential semantic relations within both verbs and semantic roles. With the help of a set of support images, an alternate learning scheme is designed to simultaneously optimize the results: update the verb using nouns corresponding to the image, and update nouns using verbs from support images. Extensive experimental results on challenging SWiG benchmarks show that our renovated framework outperforms other state-of-the-art methods under various metrics.
Zhi-Qi Cheng, Qi Dai 0001, Siyao Li, Teruko Mitamura, Alex Hauptmann 0001
ACM Multimedia4
2021 Cross-document Event Identity via Dense Annotation
abstract
In this paper, we study the identity of textual events from different documents.While the complex nature of event identity is previously studied (Hovy et al., 2013), the case of events across documents is unclear.Prior work on cross-document event coreference has two main drawbacks.First, they restrict the annotations to a limited set of event types.Second, they insufficiently tackle the concept of event identity.Such annotation setup reduces the pool of event mentions and prevents one from considering the possibility of quasiidentity relations.We propose a dense annotation approach for cross-document event coreference, comprising a rich source of event mentions and a dense annotation effort between related document pairs.To this end, we design a new annotation workflow with careful quality control and an easy-to-use annotation interface.In addition to the links, we further collect overlapping event contexts, including time, location, and participants, to shed some light on the relation between identity decisions and context.We present an open-access dataset for cross-document event coreference, CDEC-WN, collected from English Wikinews and open-source our annotation toolkit to encourage further research on cross-document tasks. 1
Adithya Pratapa, Zhengzhong Liu 0001, Kimihiro Hasegawa, Yukari Yamakawa, Shikun Zhang, Teruko Mitamura
CoNLL7
2020 Extraction of the Argument Structure of Tokyo Metropolitan Assembly Minutes: Segmentation of Question-and-Answer Sets
abstract
In this study, we construct a corpus of Japanese local assembly minutes. All speeches in an assembly were transcribed into a local assembly minutes based on the local autonomy law. Therefore, the local assembly minutes form an extremely large amount of text data. Our ultimate objectives were to summarize and present the arguments in the assemblies, and to use the minutes as primary information for arguments in local politics. To achieve this, we structured all statements in assembly minutes. We focused on the structure of the discussion, i.e., the extraction of question and answer pairs. We organized the shared task “QA Lab-PoliInfo” in NTCIR 14. We conducted a “segmentation task” to identify the scope of one question and answer in the minutes as a sub task of the shared task. For the segmentation task, 24 runs from five teams were submitted. Based on the obtained results, the best recall was 1.000, best precision was 0.940, and best F-measure was 0.895.
Keiichi Takamaru, Yasutomo Kimura, Hideyuki Shibuki, Hokuto Ototake, Yuzu Uchida, Kotaro Sakamoto, Madoka Ishioroshi, Teruko Mitamura, Noriko Kando
LREC8
2018 Open-Domain Event Detection using Distant Supervision
abstract
This paper introduces open-domain event detection, a new event detection paradigm to address issues of prior work on restricted domains and event annotation. The goal is to detect all kinds of events regardless of domains. Given the absence of training data, we propose a distant supervision method that is able to generate high-quality training data. Using a manually annotated event corpus as gold standard, our experiments show that despite no direct supervision, the model outperforms supervised models. This result indicates that the distant supervision enables robust event detection in various domains, while obviating the need for human annotation of events.
Jun Araki, Teruko Mitamura
COLING2
2018 Graph Based Decoding for Event Sequencing and Coreference Resolution
abstract
Events in text documents are interrelated in complex ways. In this paper, we study two types of relation: Event Coreference and Event Sequencing. We show that the popular tree-like decoding structure for automated Event Coreference is not suitable for Event Sequencing. To this end, we propose a graph-based decoding algorithm that is applicable to both tasks. The new decoding algorithm supports flexible feature sets for both tasks. Empirically, our event coreference system has achieved state-of-the-art performance on the TAC-KBP 2015 event coreference task and our event sequencing system beats a strong temporal-based, oracle-informed baseline. We discuss the challenges of studying these event relations.
Zhengzhong Liu 0001, Teruko Mitamura, Eduard H. Hovy
COLING2
2018 Low-resource Cross-lingual Event Type Detection via Distant Supervision with Minimal Effort
abstract
The use of machine learning for NLP generally requires resources for training. Tasks performed in a low-resource language usually rely on labeled data in another, typically resource-rich, language. However, there might not be enough labeled data even in a resource-rich language such as English. In such cases, one approach is to use a hand-crafted approach that utilizes only a small bilingual dictionary with minimal manual verification to create distantly supervised data. Another is to explore typical machine learning techniques, for example adversarial training of bilingual word representations. We find that in event-type detection task—the task to classify [parts of] documents into a fixed set of labels—they give about the same performance. We explore ways in which the two methods can be complementary and also see how to best utilize a limited budget for manual annotation to maximize performance gain.
Aldrian Obaja Muis, Naoki Otani, Nidhi Vyas, Ruochen Xu, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy
COLING6
2018 Automatic Event Salience Identification
abstract
Identifying the salience (i.e.importance) of discourse units is an important task in language understanding.While events play important roles in text documents, little research exists on analyzing their saliency status.This paper empirically studies the Event Salience task and proposes two salience detection models based on content similarities and discourse relations.The first is a feature based salience model that incorporates similarities among discourse units.The second is a neural model that captures more complex relations between discourse units.Tested on our new largescale event salience corpus, both methods significantly outperform the strong frequency baseline, while our neural model further improves the feature based one by a large margin.Our analyses demonstrate that our neural model captures interesting connections between salience and discourse unit relations (e.g., scripts and frame structures).
Zhengzhong Liu 0001, Chenyan Xiong, Teruko Mitamura, Eduard H. Hovy
EMNLP3
2018 Parser combinators for Tigrinya and Oromo morphology
Patrick Littell, Tom McCoy 0001, Na-Rae Han, Shruti Rijhwani, Zaid Sheikh, David R. Mortensen, Teruko Mitamura, Lori S. Levin
LREC7
2018 The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approach
Patrick Littell, Ruochen Xu, Zaid Sheikh, David R. Mortensen, Lori S. Levin, Francis M. Tyers, Hiroaki Hayashi, Graham Horwood, Steve Sloto, Emily Tagtow, Alan W. Black, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy
Mach. Transl.14
2017 SIGIR 2017 Workshop on Open Knowledge Base and Question Answering (OKBQA2017)
abstract
Over the past years, several challenges and calls for research projects have pointed out the dire need for pushing natural language interfaces. In this context, the importance of Semantic Web data as a premier knowledge source is rapidly increasing. But we are still far from having accurate natural language interfaces that allow handling complex information needs in a user-centric and highly performant manner. The development of such interfaces requires collaboration of a range of different fields, including natural language processing, information extraction, knowledge base construction and population, reasoning, and question answering. With the goal to join forces in the collaborative development of natural language QA systems, the second OKBQA workshop is organized within the 40th SIGIR conference.
Key-Sun Choi, Teruko Mitamura, Piek Vossen, Jin-Dong Kim, Axel-Cyrille Ngonga Ngomo
SIGIR2
2016 Generating Questions and Multiple-Choice Answers using Semantic Analysis of Texts
abstract
We present a novel approach to automated question generation that improves upon prior work both from a technology perspective and from an assessment perspective. Our system is aimed at engaging language learners by generating multiple-choice questions which utilize specific inference steps over multiple sentences, namely coreference resolution and paraphrase detection. The system also generates correct answers and semantically-motivated phrase-level distractors as answer choices. Evaluation by human annotators indicates that our approach requires a larger number of inference steps, which necessitate deeper semantic understanding of texts than a traditional single-sentence approach.
Jun Araki, Dheeraj Rajagopal, Sreecharan Sankaranarayanan, Susan Holm, Yukari Yamakawa, Teruko Mitamura
COLING6
2015 Joint Event Trigger Identification and Event Coreference Resolution with Structured Perceptron
abstract
Events and their coreference offer useful semantic and discourse resources.We show that the semantic and discourse aspects of events interact with each other.However, traditional approaches addressed event extraction and event coreference resolution either separately or sequentially, which limits their interactions.This paper proposes a document-level structured learning model that simultaneously identifies event triggers and resolves event coreference.We demonstrate that the joint model outperforms a pipelined model by 6.9 BLANC F1 and 1.8 CoNLL F1 points in event coreference resolution using a corpus in the biology domain.
Jun Araki, Teruko Mitamura
EMNLP2
2015 Bridging the Ultimate Semantic Gap: A Semantic Search Engine for Internet Videos
abstract
Semantic search in video is a novel and challenging problem in information and multimedia retrieval. Existing solutions are mainly limited to text matching, in which the query words are matched against the textual metadata generated by users. This paper presents a state-of-the-art system for event search without any textual metadata or example videos. The system relies on substantial video content understanding and allows for semantic search over a large collection of videos. The novelty and practicality is demonstrated by the evaluation in NIST TRECVID 2014, where the proposed system achieves the best performance. We share our observations and lessons in building such a state-of-the-art system, which may be instrumental in guiding the design of the future system for semantic search in video.
Lu Jiang 0004, Shoou-I Yu, Deyu Meng, Teruko Mitamura, Alex Hauptmann 0001
ICMR4
2015 Fast and Accurate Content-based Semantic Search in 100M Internet Videos
abstract
Large-scale content-based semantic search in video is an interesting and fundamental problem in multimedia analysis and retrieval. Existing methods index a video by the raw concept detection score that is dense and inconsistent, and thus cannot scale to "big data" that are readily available on the Internet. This paper proposes a scalable solution. The key is a novel step called concept adjustment that represents a video by a few salient and consistent concepts that can be efficiently indexed by the modified inverted index. The proposed adjustment model relies on a concise optimization framework with interpretations. The proposed index leverages the text-based inverted index for video retrieval. Experimental results validate the efficacy and the efficiency of the proposed method. The results show that our method can scale up the semantic search while maintaining state-of-the-art search performance. Specifically, the proposed method (with reranking) achieves the best result on the challenging TRECVID Multimedia Event Detection (MED) zero-example task. It only takes 0.2 second on a single CPU core to search a collection of 100 million Internet videos.
Lu Jiang 0004, Shoou-I Yu, Deyu Meng, Yi Yang 0001, Teruko Mitamura, Alex Hauptmann 0001
ACM Multimedia5
2014 Detecting Subevent Structure for Event Coreference Resolution
Jun Araki, Zhengzhong Liu 0001, Eduard H. Hovy, Teruko Mitamura
LREC4
2014 Resources for the Detection of Conventionalized Metaphors in Four Languages
Lori S. Levin, Teruko Mitamura, Brian MacWhinney, Davida Fromm, Jaime G. Carbonell, Weston Feely, Robert E. Frederking, Anatole Gershman
LREC2
2014 Supervised Within-Document Event Coreference using Information Propagation
Zhengzhong Liu 0001, Jun Araki, Eduard H. Hovy, Teruko Mitamura
LREC4
2014 Zero-Example Event Search using MultiModal Pseudo Relevance Feedback
abstract
We propose a novel method MultiModal Pseudo Relevance Feedback (MMPRF) for event search in video, which requires no search examples from the user. Pseudo Relevance Feedback has shown great potential in retrieval tasks, but previous works are limited to unimodal tasks with only a single ranked list. To tackle the event search task which is inherently multimodal, our proposed MMPRF takes advantage of multiple modalities and multiple ranked lists to enhance event search performance in a principled way. The approach is unique in that it leverages not only semantic features, but also non-semantic low-level features for event search in the absence of training data. Evaluated on the TRECVID MEDTest dataset, the approach improves the baseline by up to 158% in terms of the mean average precision. It also significantly contributes to CMU Team's final submission in TRECVID-13 Multimedia Event Detection.
Lu Jiang 0004, Teruko Mitamura, Shoou-I Yu, Alex Hauptmann 0001
ICMR2
2014 Easy Samples First: Self-paced Reranking for Zero-Example Multimedia Search
abstract
Reranking has been a focal technique in multimedia retrieval due to its efficacy in improving initial retrieval results. Current reranking methods, however, mainly rely on the heuristic weighting. In this paper, we propose a novel reranking approach called Self-Paced Reranking (SPaR) for multimodal data. As its name suggests, SPaR utilizes samples from easy to more complex ones in a self-paced fashion. SPaR is special in that it has a concise mathematical objective to optimize and useful properties that can be theoretically verified. It on one hand offers a unified framework providing theoretical justifications for current reranking methods, and on the other hand generates a spectrum of new reranking schemes. This paper also advances the state-of-the-art self-paced learning research which potentially benefits applications in other fields. Experimental results validate the efficacy and the efficiency of the proposed method on both image and video search tasks. Notably, SPaR achieves by far the best result on the challenging TRECVID multimedia event search task.
Lu Jiang 0004, Deyu Meng, Teruko Mitamura, Alex Hauptmann 0001
ACM Multimedia3
2013 An English Reading Tool as a NLP Showcase
Mahmoud Azab, Ahmed Salama, Kemal Oflazer, Hideki Shima 0001, Jun Araki, Teruko Mitamura
IJCNLP6
2012 Diversifiable Bootstrapping for Acquiring High-Coverage Paraphrase Resource
Hideki Shima 0001, Teruko Mitamura
LREC2
2012 Multimodal knowledge-based analysis in multimedia event detection
abstract
Multimedia Event Detection (MED) is a multimedia retrieval task with the goal of finding videos of a particular event in a large-scale Internet video archive, given example videos and text descriptions. We focus on the multimodal knowledge-based analysis in MED where we utilize meaningful and semantic features such as Automatic Speech Recognition (ASR) transcripts, acoustic concept indexing (i.e. 42 acoustic concepts) and visual semantic indexing (i.e. 346 visual concepts) to characterize videos in archive. We study two scenarios where we either do or do not use the provided example videos. In the former, we propose a novel Adaptive Semantic Similarity (ASS) to measure textual similarity between ASR transcripts of videos. We also incorporate acoustic concept indexing and classification to retrieve test videos, specially with too few spoken words. In the latter 'ad-hoc' scenario where we do not have any example video, we use only the event kit description to retrieve test videos ASR transcripts and visual semantics. We also propose an event-specific fusion scheme to combine textual and visual retrieval outputs. Our results show the effectiveness of the proposed ASS and acoustic concept indexing methods and their complimentary role. We also conduct a set of experiments to assess the proposed framework for the 'ad-hoc' scenario.
Ehsan Younessian, Teruko Mitamura, Alex Hauptmann 0001
ICMR2
2012 Introduction to the Special Issue on RITE
abstract
No abstract available.
Teruko Mitamura, Noriko Kando, Koichi Takeda 0002
ACM Trans. Asian Lang. Inf. Process.1
2012 Evaluating Textual Entailment Recognition for University Entrance Examinations
abstract
The present article addresses an attempt to apply questions in university entrance examinations to the evaluation of textual entailment recognition. Questions in several fields, such as history and politics, primarily test the examinee’s knowledge in the form of choosing true statements from multiple choices. Answering such questions can be regarded as equivalent to finding evidential texts from a textbase such as textbooks and Wikipedia. Therefore, this task can be recast as recognizing textual entailment between a description in a textbase and a statement given in a question. We focused on the National Center Test for University Admission in Japan and converted questions into the evaluation data for textual entailment recognition by using Wikipedia as a textbase. Consequently, it is revealed that nearly half of the questions can be mapped into textual entailment recognition; 941 text pairs were created from 404 questions from six subjects. This data set is provided for a subtask of NTCIR RITE (Recognizing Inference in Text), and 16 systems from six teams used the data set for evaluation. The evaluation results revealed that the best system achieved a correct answer ratio of 56%, which is significantly better than a random choice baseline.
Yusuke Miyao, Hideki Shima 0001, Hiroshi Kanayama, Teruko Mitamura
ACM Trans. Asian Lang. Inf. Process.4
2011 Effects of Adaptive Prompted Self-explanation on Robust Learning of Second Language Grammar
Ruth Wylie, Melissa Sheng, Teruko Mitamura, Kenneth R. Koedinger
AIED3
2010 Analogies, Explanations, and Practice: Examining How Task Types Affect Second Language Grammar Learning
Ruth Wylie, Kenneth R. Koedinger, Teruko Mitamura
Intelligent Tutoring Systems (1)3
2010 Interlingual annotation of parallel text corpora: a new framework for annotation and evaluation
abstract
Abstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines.
Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan
Nat. Lang. Eng.10
2010 Probabilistic models for answer-ranking in multilingual question-answering
abstract
This article presents two probabilistic models for answering ranking in the multilingual question-answering (QA) task, which finds exact answers to a natural language question written in different languages. Although some probabilistic methods have been utilized in traditional monolingual answer-ranking, limited prior research has been conducted for answer-ranking in multilingual question-answering with formal methods. This article first describes a probabilistic model that predicts the probabilities of correctness for individual answers in an independent way. It then proposes a novel probabilistic method to jointly predict the correctness of answers by considering both the correctness of individual answers as well as their correlations. As far as we know, this is the first probabilistic framework that proposes to model the correctness and correlation of answer candidates in multilingual question-answering and provide a novel approach to design a flexible and extensible system architecture for answer selection in multilingual QA. An extensive set of experiments were conducted to show the effectiveness of the proposed probabilistic methods in English-to-Chinese and English-to-Japanese cross-lingual QA, as well as English, Chinese, and Japanese monolingual QA using TREC and NTCIR questions.
Jeongwoo Ko, Luo Si, Eric Nyberg, Teruko Mitamura
ACM Trans. Inf. Syst.4
2008 Introduction to the NTCIR-6 Special Issue
abstract
introduction Introduction to the NTCIR-6 Special Issue Share on Authors: Noriko Kando National Institute of Informatics National Institute of InformaticsView Profile , Teruko Mitamura Carnegie Mellon University Carnegie Mellon UniversityView Profile , Tetsuya Sakai News Watch, Co. News Watch, Co.View Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 7Issue 2June 2008 Article No.: 4pp 1–3https://doi.org/10.1145/1362782.1362783Online:01 April 2008Publication History 6citation234DownloadsMetricsTotal Citations6Total Downloads234Last 12 Months5Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Noriko Kando, Teruko Mitamura, Tetsuya Sakai
ACM Trans. Asian Lang. Inf. Process.2
2007 Language-independent Probabilistic Answer Ranking for Question Answering
Jeongwoo Ko, Teruko Mitamura, Eric Nyberg
ACL2
2007 What is the Jeopardy Model? A Quasi-Synchronous Grammar for QA
Mengqiu Wang, Noah A. Smith, Teruko Mitamura
EMNLP-CoNLL3
2006 A Fast, Accurate Deterministic Parser for Chinese
abstract
We present a novel classifier-based deterministic parser for Chinese constituency parsing. Our parser computes parse trees from bottom up in one pass, and uses classifiers to make shift-reduce decisions. Trained and evaluated on the standard training and test sets, our best model (using stacked classifiers) runs in linear time and has labeled precision and recall above 88% using gold-standard part-of-speech tags, surpassing the best published results. Our SVM parser is 2-13 times faster than state-of-the-art parsers, while producing more accurate results. Our Maxent and DTree parsers run at speeds 40-270 times faster than state-of-the-art parsers, but with 5-6% losses in accuracy.
Mengqiu Wang, Kenji Sagae, Teruko Mitamura
ACL3
2006 Analyzing the Effects of Spoken Dialog Systems on Driving Behavior
Jeongwoo Ko, Fumihiko Murase, Teruko Mitamura, Eric Nyberg, Masahiko Tateishi, Ichiro Akahori
LREC3
2006 Parallel Syntactic Annotation of Multiple Languages
Owen Rambow, Bonnie J. Dorr, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Flo Reeder, Advaith Siddharthan
LREC10
2006 Modular Approach to Error Analysis and Evaluation for Multilingual Question Answering
Hideki Shima 0001, Mengqiu Wang, Frank Lin, Teruko Mitamura
LREC4
2005 Capturing knowledge from domain text with controlled language
abstract
This paper describes a prototype system which captures semantic knowledge from domain text using controlled language. The KANTOO system is used to analyze input sentences from college-level science textbooks, producing sentence-level meaning representations (interlingua). The interlingua expressions are mapped into F-logic statements, which are be stored in a separate knowledge base to support reasoning in the domain.
Eric Nyberg, Teruko Mitamura, Justin Betteridge
K-CAP2
2004 Robust speech dialog interface for car telematics service
abstract
We describe new consumer services based on speech processing technologies to support a new digital/mobile era of ubiquitous communication. First, we propose, a compact and noise robust embedded speech recognition middleware implemented on microprocessors focused on sophisticated HMIs (human machine interfaces) for car information systems (i.e. car telematics). Second, we report on a novel and sophisticated dialog management/manager (DM) system, based on VoiceXML (voice extensible markup language), called CAMMIA (conversational agent for multimedia mobile information access). The proposed DM handles two important issues: an automatic generation scheme for lexicons and grammars, and an effective combination/merger between automatic speech recognition (ASR) and natural language processing (NLP). The new DM scheme has been evaluated for an application of the car telematics service task after integration with ASR and a VoiceXML interpreter (VXI).
Nobuo Hataoka, Yasunari Obuchi, Teruko Mitamura, Eric Nyberg
CCNC3
2004 A comparison of confirmation styles for error handling in a speech dialog system
abstract
Speech recognition errors are inevitable in a speech dialog system. It is important to provide a dialog flow that allows the user to correct system errors and quickly return to the original dialog. This paper describes explicit, final and implicit confirmation styles that are implemented in the CAMMIA speech dialog system, and compares them from the viewpoint of usability. Our results show that a final confirmation with fewer confirmation turns is preferred by the user when there is no error in the dialog. On the other hand, an explicit confirmation is preferred when an error occurs.
Hirohiko Sagawa, Teruko Mitamura, Eric Nyberg
INTERSPEECH2
2004 Pronominal Anaphora Resolution for Unrestricted Text
Anna Kupsc, Teruko Mitamura, Benjamin Van Durme, Eric Nyberg
LREC2
2004 An Information Repository Model for Advanced Question Answering Systems
Vasco Pedro, Jeongwoo Ko, Eric Nyberg, Teruko Mitamura
LREC4
2003 Source language diagnostics for MT
abstract
This paper presents a source language diagnostic system for controlled translation. Diagnostics were designed and implemented to address the most difficult rewrites for authors, based on an empirical analysis of log files containing over 180,000 sentences. The design and implementation of the diagnostic system are presented, along with experimental results from an empirical evaluation of the completed system. We found that the diagnostic system can correctly identify the problem in 90.2% of the cases. In addition, depending on the type of grammar problem, the diagnostic system may offer a rewritten sentence. We found that 89.4% of the rewritten sentences were correctly rewritten. The results suggest that these methods could be used as the basis for an automatic rewriting system in the future.
Teruko Mitamura, Kathrin Baker, David Svoboda, Eric Nyberg
MTSummit1
2002 Knowledge-based extraction of named entities
abstract
The usual approach to named-entity detection is to learn extraction rules that rely on linguistic, syntactic, or document format patterns that are consistent across a set of documents. However, when there is no consistency among documents, it may be more effective to learn document-specific extraction rules.This paper presents a knowledge-based approach to learning rules for named-entity extraction. Document-specific extraction rules are created using a generate-and-test paradigm and a database of known named-entities. Experimental results show that this approach is effective on Web documents that are difficult for the usual methods.
Jamie Callan, Teruko Mitamura
CIKM2
2001 Pronominal anaphora resolution in KANTOO English-to-Spanish machine translation system
abstract
We describe the automatic resolution of pronominal anaphora using KANT Controlled English (KCE) and the KANTOO English-to-Spanish MT system. Our algorithm is based on a robust, syntax-based approach that applies a set of restrictions and preferences to select the correct antecedent. We report a success rate of 89.6% on a training corpus with 289 anaphors, and 87.5% on held-out data containing 145 anaphors. Resolution of anaphors is important in translation, due to gender mismatches among languages; our approach translates anaphors to Spanish with 97.2% accuracy.
Teruko Mitamura, Eric Nyberg, Enrique Torrejón, David Svoboda, Kathrin Baker
MTSummit1
2001 Conference Information Management System: Towards a Personal Assistant System
Tsunenori Mine, Makoto Amamiya, Teruko Mitamura
Web Intelligence3
1999 Controlled language for multilingual machine translation
Teruko Mitamura
MTSummit1
1994 Coping With Ambiguity in a Large-Scale Machine Translation System
Kathrin Baker, Alexander Franz, Pamela W. Jordan, Teruko Mitamura, Eric Nyberg
COLING4
1994 Evaluation Metrics for Knowledge-Based Machine Translation
Eric Nyberg, Teruko Mitamura, Jaime G. Carbonell
COLING2
1994 Acquisition of large lexicons for practical knowledge-based MT
Deryle W. Lonsdale, Teruko Mitamura, Eric Nyberg
Mach. Transl.2
1992 Hierarchical Lexical Structure And Interpretive Mapping In Machine Translation
Teruko Mitamura, Eric Nyberg
COLING1
1992 The Kant System: Fast, Accurate, High-Quality Translation In Practical Domains
Eric Nyberg, Teruko Mitamura
COLING2
1989 A massively parallel model of speech-to-speech dialog translation: a step toward interpreting telephony
Hiroaki Kitano, Hideto Tomabechi, Teruko Mitamura, Hitoshi Iida
EUROSPEECH3
1989 Lexicons
Donna Gates, Dawn Haberlach, Todd Kaufmann, Marion Kee, Rita McCardell Doerr, Teruko Mitamura, Ira Monarch, Stephen Morrisson, Sergei Nirenburg, Eric Nyberg, Koichi Takeda 0002, Margalit Zabludowski
Mach. Transl.6
1989 Analysis and generation grammars
Donna Gates, Koichi Takeda 0002, Teruko Mitamura, Lori S. Levin, Marion Kee
Mach. Transl.3