EDBT 2026 Demo / reviewers in the wild / expert
Yi Luan
dblp:125/7491
· DBLP profile ↗
19ranked-venue papers
9as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Vision and language · 46% Question answering and dialogue systems · 20% Information extraction and text analysis · 14% | |
| Databases, data mining, and information retrieval
8 papers |
Information retrieval · 71% Knowledge graphs · 21% Query processing and optimization · 8% |
Topics — the 27 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language pretraining |
1.3 | 2 | 2023 | Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities · ICCV 2023 Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.2 | 2 | 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples · ICLR 2023 Large Dual Encoders Are Generalizable Retrievers · EMNLP 2022 |
Information retrieval
retrieval models |
0.8 | 2 | 2023 | Large Dual Encoders Are Generalizable Retrievers · EMNLP 2022 Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023 |
Information retrieval
image retrieval |
0.8 | 1 | 2024 | MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions · ICML 2024 |
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering |
0.7 | 1 | 2023 | Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
multimodal knowledge |
0.7 | 1 | 2023 | Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.7 | 1 | 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples · ICLR 2023 |
Computer vision › Vision and language
visual entity recognition |
0.7 | 1 | 2023 | Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities · ICCV 2023 |
Computer vision › Vision and language
visual question answering |
0.7 | 1 | 2023 | Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › robust question answering
ambiguous question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
factoid question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
long-form question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Information retrieval › retrieval models › neural retrieval › dense retrieval
bi-encoder retrieval |
0.6 | 1 | 2022 | Large Dual Encoders Are Generalizable Retrievers · EMNLP 2022 |
Information retrieval
evaluation |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Query processing and optimization
query rewriting |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Information retrieval › evaluation › task-based evaluation
question answering evaluation |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Information retrieval › machine learning for information retrieval
reinforcement learning for retrieval |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis › relation extraction › joint extraction
joint entity, relation and event extraction |
0.4 | 1 | 2019 | Entity, Relation, and Event Extraction with Contextualized Span Representations · EMNLP/IJCNLP (1) 2019 |
Knowledge graphs
domain-specific knowledge graph |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Knowledge graphs
knowledge graph construction |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Knowledge graphs
link prediction |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › document analysis › scholarly text analysis
scientific information extraction |
0.3 | 1 | 2018 | Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction · EMNLP 2018 |
Knowledge graphs › knowledge graph construction
scientific knowledge graph construction |
0.3 | 1 | 2018 | Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › sequence labeling
semi-supervised sequence labeling |
0.3 | 1 | 2017 | Scientific Information Extraction with Semi-supervised Neural Tagging · EMNLP 2017 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.1 | 1 | 2019 | Entity, Relation, and Event Extraction with Contextualized Span Representations · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2017 | Scientific Information Extraction with Semi-supervised Neural Tagging · EMNLP 2017 |
Methods — techniques the papers use, named apart from their topics
foundation model instruction synthesis · 1.5contrastive learning · 1.5prompt generation · 1.3fine-tuning · 1.3entity recognition · 1.3dense retrieval · 1.3summarization · 1.1automated metric · 1.1multimodal pretraining · 0.7autoregressive recognition · 0.7transfer learning · 0.6dual encoder · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsabstractImage retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures. Recent works leverage text instructions to allow users to more freely express their search intents. However, they primarily focus on image pairs that are visually similar and/or can be characterized by a small set of pre-defined relations. The core thesis of this paper is that text instructions can enable retrieving images with richer relations beyond visual similarity. To show this, we introduce MagicLens, a series of self-supervised image retrieval models that support open-ended instructions. MagicLens is built on a key novel insight: image pairs that naturally occur on the same web pages contain a wide range of implicit relations (e.g., inside view of), and we can bring those implicit relations explicit by synthesizing instructions via foundation models. Trained on 36.7M (query image, instruction, target image) triplets with rich semantic relations mined from the web, MagicLens achieves results comparable with or better than prior best on eight benchmarks of various image retrieval tasks, while maintaining high parameter efficiency with a significantly smaller model size. Additional human analyses on a 1.4M-image unseen corpus further demonstrate the diversity of search intents supported by MagicLens. Code and models are publicly available at the https://open-vision-language.github.io/MagicLens/. Kai Zhang 0033, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su 0001, Ming-Wei Chang |
ICML | 2 |
| 2023 | Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?abstractPre-trained vision and language models (Chen et al., 2023b,a;Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering.However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and informationseeking.In this study, we introduce INFOS-EEK 1 , a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics.Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but finetuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training.Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.* Work done when interned at Google 1 Our dataset is available at https:// open-vision-language.github.io/infoseek/.Dataset OK-VQA ViQuAE INFOSEEK PaLM (Q-only) 23.8 31.5 5.6 Current SotA 66.1 22.1 18.2 Require Knowledge † 29.2% 95.2% 95.6% † :% of questions that require knowledge to answer.PaLM (Q-only): a question-only baseline using PaLM. Yang Chen 0065, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, Ming-Wei Chang |
EMNLP | 3 |
| 2023 | Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesabstractLarge-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki‡by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities. Hexiang Hu, Yi Luan, Yang Chen 0065, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang |
ICCV | 2 |
| 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma 0004, Yi Luan, Jianmo Ni, Jing Lu 0014, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang |
ICLR | 4 |
| 2022 | Large Dual Encoders Are Generalizable RetrieversabstractJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, Yinfei Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jianmo Ni, Chen Qu 0001, Jing Lu 0014, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma 0004, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang |
EMNLP | 8 |
| 2022 | ASQA: Factoid Questions Meet Long-Form AnswersabstractAn abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA).This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations.The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality.In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA.Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation.Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity.In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question.We use this notion of correctness to define an automated metric of performance for ASQA.Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines. Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang |
EMNLP | 2 |
| 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningabstractCompared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision. Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar |
EMNLP | 2 |
| 2021 | Sparse, Dense, and Attentional Representations for Text RetrievalabstractAbstract Dual encoders perform retrieval by encoding documents and queries into dense low-dimensional vectors, scoring each document by its inner product with the query. We investigate the capacity of this architecture relative to sparse bag-of-words models and attentional neural networks. Using both theoretical and empirical analysis, we establish connections between the encoding dimension, the margin between gold and lower-ranked documents, and the document length, suggesting limitations in the capacity of fixed-length encodings to support precise retrieval of long documents. Building on these insights, we propose a simple neural model that combines the efficiency of dual encoders with some of the expressiveness of more costly attentional architectures, and explore sparse-dense hybrids to capitalize on the precision of sparse retrieval. These models outperform strong alternatives in large-scale retrieval. Yi Luan, Jacob Eisenstein, Kristina Toutanova, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Seismic Time-Frequency Analysis Based on Entropy-Optimized Paul Wavelet TransformabstractA new method of time-frequency analysis called the entropy-optimized Paul wavelet transform (WT) (EOPWT), is carried out with the use of Paul wavelet-based continuous WT (CWT) on seismic data with the help of the Rényi entropy measurement. Despite its complicated formulation and not being physically intuitive, we extend the parameter definition of Paul wavelet to the real-valued domain, where we demonstrate its superiority over commonly used wavelet functions because of its flexibility in controlling the time-frequency resolution. It can be even regarded as a generalized waveform or an alternative choice with respect to both Ricker and Morlet wavelets. Synthetic seismic trace analysis shows that EOPWT objectively produces a more compact representation of time-frequency localization than other conventional tools in capturing the nonlinear and nonstationary features of input signals. It also offers new possibilities in yielding more reliable seismic attributes in the field of reservoir characterization. We implement spectral decomposition using the proposed method upon a data set from a deeply buried carbonate gas reservoir in southwestern China. Combining the energy absorption analysis method, we obtain the attenuation gradient attribute section indicating gas accumulation. The field data example demonstrates the effectiveness of direct hydrocarbon detection using EOPWT. Yi Luan, Jixing Cheng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | PaperRobot: Incremental Draft Generation of Scientific IdeasabstractWe present a PaperRobot who performs as an automatic research assistant by (1) conducting deep understanding of a large collection of human-written papers in a target domain and constructing comprehensive background knowledge graphs (KGs); (2) creating new ideas by predicting links from the background KGs, by combining graph attention and contextual text attention; (3) incrementally writing some key elements of a new paper based on memory-attention networks: from the input title along with predicted related entities to generate a paper abstract, from the abstract to generate conclusion and future work, and finally from future work to generate a title for a follow-on paper.Turing Tests, where a biomedical domain expert is asked to compare a system output and a human-authored string, show PaperRobot generated abstracts, conclusion and future work sections, and new titles are chosen over human-written ones up to 30%, 24% and 12% of the time, respectively. 1 keeps almost the same across years.In 2012, US scientists estimated that they read, on average, only 264 papers per year (1 out of 5000 available papers), which is, statistically, not different from what they reported in an identical survey last conducted in 2005.PaperRobot automatically reads existing papers to build background knowledge graphs (KGs), in which nodes are entities/concepts and edges are the relations between these entities (Section 2.2). Qingyun Wang 0005, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji 0001, Mohit Bansal, Yi Luan |
ACL (1) | 7 |
| 2019 | Entity, Relation, and Event Extraction with Contextualized Span RepresentationsabstractDavid Wadden, Ulme Wennberg, Yi Luan, Hannaneh Hajishirzi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dave Wadden, Ulme Wennberg, Yi Luan, Hannaneh Hajishirzi |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph ConstructionabstractWe introduce a multi-task setup of identifying and classifying entities, relations, and coreference clusters in scientific articles.We create SCIERC, a dataset that includes annotations for all three tasks and develop a unified framework called Scientific Information Extractor (SCIIE) for with shared span representations.The multi-task setup reduces cascading errors between tasks and leverages cross-sentence relations through coreference links.Experiments show that our multi-task model outperforms previous models in scientific information extraction without using any domain-specific features.We further show that the framework supports construction of a scientific knowledge graph, which we use to analyze information in scientific literature. 1 Extracting nodes (entities) The SCIIE model extracts entities, their relations, and coreference Yi Luan, Luheng He, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 1 |
| 2017 | Scientific Information Extraction with Semi-supervised Neural TaggingabstractThis paper addresses the problem of extracting keyphrases from scientific articles and categorizing them as corresponding to a task, process, or material.We cast the problem as sequence tagging and introduce semi-supervised methods to a neural tagging model, which builds on recent advances in named entity recognition.Since annotated training data is scarce in this domain, we introduce a graph-based semi-supervised algorithm together with a data selection scheme to leverage unannotated articles.Both inductive and transductive semi-supervised learning strategies outperform state-of-the-art information extraction performance on the 2017 SemEval Task 10 ScienceIE task. Yi Luan, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 1 |
| 2017 | Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation ModelsabstractBuilding a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training neural conversation models that leverages both conversation data across speakers and other types of data pertaining to the speaker and speaker roles to be modeled. Experiments show that our approach leads to significant improvements over baseline model quality, generating responses that capture more precisely speakers’ traits and speaking styles. The model offers the benefits of being algorithmically simple and easy to implement, and not relying on large quantities of data representing specific individual speakers. Yi Luan, Chris Brockett, William B. Dolan, Jianfeng Gao 0001, Michel Galley |
IJCNLP(1) | 1 |
| 2015 | Efficient learning for spoken language understanding tasks with word embedding based pre-training
Yi Luan, Shinji Watanabe 0001, Bret Harsham |
INTERSPEECH | 1 |
| 2014 | Semi-supervised noise dictionary adaptation for exemplar-based noise robust speech recognitionabstractThe exemplar-based approaches, which model signals as a sparse linear combination of exemplars of signals, are proved to have state-of-the-art performance in noise robust ASR, especially on low SNRs. However, since both the speech exemplars and noise exemplars are built from training data and are fixed throughout the process of enhancing speech features, the conventional approach is especially weak for unknown types of noise. Therefore, in this paper, we propose a semi-supervised approach which automatically adapt noise exemplars to the target noise, while keeping the speech exemplars fixed. Continuous digits recognition experiments show that this approach is much more robust for unknown noise. The recognition errors are reduced by 36.2%. Yi Luan, Daisuke Saito, Yosuke Kashiwagi, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 1 |
| 2014 | Relating automatic vowel space estimates to talker intelligibilityabstractDifferences in pronunciation have been shown to underlie significant talker-dependent intelligibility differences. There are several dimensions of variability that are correlated with talker intelligibility including pitch range, vowel-space expansion, and rhythmic patterns. Prior work has shown that some of the better predictors of individual intelligibility are based on the talker’s F1 by F2 vowel space, but findings are based on handcorrected measurements on carefully balanced sets of vowels, making large scale analysis impractical. This paper proposes a novel method for automatic estimation of a talker’s vowel space using sparse expanded vowel space representations, including an approximate convex hull sampling, which are projected to a low dimensional space for intelligibility scoring. Both supervised and unsupervised mappings are used to generate an intelligibility score. Automatic intelligibility rankings are assessed in terms of correlation with an intelligibility score based on human transcription accuracy. We find that including a larger sample of vowels (beyond point vowels) leads to improved performance, obtaining correlations of roughly 0.6 for this feature alone, which is a strong result given that there are other factors that may also contribute to a talker’s intelligibility in addition to a talker’s vowel space area. Yi Luan, Richard A. Wright, Mari Ostendorf, Gina-Anne Levow |
INTERSPEECH | 1 |
| 2014 | Recognition of stance strength and polarity in spontaneous speechabstractFrom activities as simple as scheduling a meeting to those as complex as balancing a national budget, people take stances in negotiations and decision making. While the related areas of subjectivity and sentiment analysis have received significant attention, work has focused almost exclusively on text, and much stance-taking activity is carried out verbally. This paper investigates automatic recognition of stance-taking in spontaneous speech. It first presents a new annotated corpus of spontaneous, conversational speech designed to elicit high densities of stance-taking at different strengths. Speaker spurts are annotated both for strength of stance-taking behavior and polarity of stance. Based on this annotated corpus, we develop classifiers for automatic recognition of stance-taking behavior in speech. We employ a range of lexical, speaking style, and prosodic features in a boosting framework. The classifiers achieve strong accuracies on both binary detection of stance and four-way recognition of stance strength, well above most common class assignment. Finally, we classify the polarity of stance-taking spurts, obtaining accuracies around 80%. The best classifiers rely primarily on word unigram features, with speaking style and prosodic features yielding lower accuracies but still well above common class assignment. Gina-Anne Levow, Valerie Freeman, Alena Hrynkevich, Mari Ostendorf, Richard A. Wright, Julian Chan, Yi Luan, Trang Tran 0001 |
SLT | 7 |
| 2012 | Performance improvement of automatic pronunciation assessment in a noisy classroomabstractIn recent years Computer-Assisted Language Learning (CALL) systems have been widely used in foreign language education. Some systems use automatic speech recognition (ASR) technologies to detect pronunciation errors and estimate the proficiency level of individual students. When speech recording is done in a CALL classroom, however, utterances of a student are always recorded with those of the others in the same class. The latter utterances are just background noise, and the performance of automatic pronunciation assessment is degraded especially when a student is surrounded with very active students. To solve this problem, we apply a noise reduction technique, Stereo-based Piecewise Linear Compensation for Environments (SPLICE), and the compensated feature sequences are input to a Goodness Of Pronunciation (GOP) assessment system. Results show that SPLICE-based noise reduction works very well as a means to improve the assessment performance in a noisy classroom. Yi Luan, Masayuki Suzuki, Yutaka Yamauchi, Nobuaki Minematsu, Shuhei Kato, Keikichi Hirose |
SLT | 1 |