Yi-Cheng Wang

dblp:222/8809 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-6819-0400ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents.Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domainspecific content such as full-length books.To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning.However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision.On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts.To address this issue, we introduce a multimodal knowledge graphbased RAG that enables cross-modal reasoning for better content understanding.Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process.Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing approaches on both textual and multimodal benchmarks.
Chi-Hsiang Hsiao, Yi-Cheng Wang, Tzung-Sheng Lin, Yi-Ren Yeh, Chu-Song Chen
ACL (1)2
2025 PRIME: Novel Prompting Strategies for Effective Biasing Word Recognition in Contextualized ASR
abstract
Accurately recognizing domain-specific words remains an arduous challenge facing current automatic speech recognition (ASR) systems. While there are a number of prior arts managing to incorporate domain context through crossattention mechanisms, these methods often add additional components that increase model complexity and focus solely on wordlevel cues, overlooking broader domain-level topic information. To address these limitations, we put forward PRIME, a simple yet effective prompt-tuning method that enriches ASR with domain context using LLM-generated topic descriptions. To mitigate the limited context window of the ASR decoder, we introduce a biasing word retriever that selects the most relevant domain-specific words to construct informative prompts. Notably, PRIME requires no architectural modifications and offers a lightweight, scalable solution for contextualized ASR. A series of experiments counducted on the AISHELL and SlideSpeech benchmark datasets show that PRIME considerably promotes biasing word recognition, outperforming some strong baselines.
Yu-Chun Liu, Li-Ting Pai, Yi-Cheng Wang, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Juan-Wei Xu, Berlin Chen
ASRU3
2025 ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
abstract
Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, wherein the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while factoring in the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our regularizer. A series of experiments conducted on the speechocean762 benchmark dataset suggests the feasibility and effectiveness of our approach in relation to several competitive baselines.
Bi-Cheng Yan, Yi-Cheng Wang, Jiun-Ting Li, Meng-Shin Lin, Hsin-Wei Wang, Wei-Cheng Chao, Berlin Chen
ICASSP2
2024 An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training Strategies
abstract
Bi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Bi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin-Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen
ACL (1)3
2024 DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition
abstract
End-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently been proposed, which normally build on pho-netic-level edit distance algorithms and have shown impressive NEC performance. However, as the named entity (NE) list grows, the problems of phonetic confusion in the NE list are exacerbated; for example, homophone ambiguities increase substantially. In view of this, we proposed a novel Description Augmented Named entity CorrEctoR (dubbed DANCER), which leverages entity descriptions to provide additional information to facilitate mitigation of phonetic con-fusion for NEC on ASR transcription. To this end, an efficient entity description augmented masked language model (EDA-MLM) comprised of a dense retrieval model is introduced, enabling MLM to adapt swiftly to domain-specific entities for the NEC task. A series of experiments conducted on the AISHELL-1 and Homophone datasets confirm the effectiveness of our modeling approach. DANCER outperforms a strong baseline, the phonetic edit-distance-based NEC model (PED-NEC), by a character error rate (CER) reduction of about 7% relatively on AISHELL-1 for named entities. More notably, when tested on Homophone that contain named entities of high phonetic confusion, DANCER offers a more pronounced CER reduction of 46% relatively over PED-NEC for named entities. The code is available at https://github.com/Amiannn/Dancer.
Yi-Cheng Wang, Hsin-Wei Wang, Bi-Cheng Yan, Chi-Han Lin, Berlin Chen
LREC/COLING1
2024 An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
abstract
With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the code-switching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the code-switching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pre-trained dual-encoders, meanwhile having access only to the code-switching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders’ output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture.
Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Chi-Han Lin, Berlin Chen
ICASSP3
2024 Automated Speaking Assessment of Conversation Tests with Novel Graph-Based Modeling on Spoken Response Coherence
abstract
Automated speaking assessment in conversation tests (ASAC) aims to evaluate the overall speaking proficiency of an L2 (second-language) speaker in a setting where an interlocutor interacts with one or more candidates. Although prior ASAC approaches have shown promising performance on their respective datasets, there is still a dearth of research specifically focused on incorporating the coherence of the logical flow within a conversation into the grading model. To address this critical challenge, we propose a hierarchical graph model that aptly incorporates both broad inter-response interactions (e.g., discourse relations) and nuanced semantic information (e.g., semantic words and speaker intents), which is subsequently fused with contextual information for the final prediction. Extensive experimental results on the NICT-JLE benchmark dataset suggest that our proposed modeling approach can yield considerable improvements in prediction accuracy with respect to various assessment metrics, as compared to some strong baselines. This also sheds light on the importance of investigating coherence-related facets of spoken responses in ASAC.
Jiun-Ting Li, Bi-Cheng Yan, Tien-Hong Lo, Yi-Cheng Wang, Yung-Chang Hsu, Berlin Chen
SLT4
2024 An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
abstract
End-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to thetoken due to higher-frequency words not being present in the context list. Second, the long-tailed distribution within the context list itself still causes the model to perform poorly on low-frequency context words. In light of this, we explore in-depth the impact of altering the context list to have words with different frequency distributions on model performance, and meanwhile extend CA with a simple yet effective context-balanced learning objective1. A series of experiments conducted on the AISHELL-1 benchmark dataset suggests that using all vocabulary words from the training corpus as the context list and pairing them with our balanced objective yields the best performance, demonstrating a significant reduction in character error rate (CER) by up to 1.21% and a more pronounced 9.44% reduction in the error rate of zero-shot words.1The code is available at: https://github.com/Amiannn/espnet/tree/context_balanced_adapter
Yi-Cheng Wang, Li-Ting Pai, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Berlin Chen
SLT1
2024 Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
abstract
Code-switching—where multilingual speakers alternately switch between languages during conversations—still poses significant challenges to end-to-end (E2E) automatic speech recognition (ASR) systems due to phenomena of both acoustic and semantic confusion. This issue arises because ASR systems struggle to handle the rapid alternation of languages effectively, which often leads to significant performance degradation. Our main contributions are at least threefold: First, we incorporate language identification (LID) information into several intermediate layers of the encoder, aiming to enrich output embeddings with more detailed language information. Secondly, through the novel application of language boundary alignment loss, the subsequent ASR modules are enabled to more effectively utilize the knowledge of internal language posteriors. Third, we explore the feasibility of using language posteriors to facilitate deep interaction between shared encoder and language-specific encoders. Through comprehensive experiments on the SEAME corpus, we have verified that our proposed method outperforms the prior-art method, disentangle based mixture-of-experts (D-MoE), further enhancing the acuity of the encoder to languages.
Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen
SLT3
2024 Directional Land Surface Emissivity Retrieval From Combined MERSI/FY-3D/E and MODIS Data
abstract
This paper addresses the directional land surface emissivity (LSE) retrieval from the data acquired by the MEdium Resolution Spectral Imager (MERSI) on Fengyun 3D and 3E (FY-3D/E) satellites and the Moderate-resolution Imaging Spectroradiometer (MODIS) on Terra and Aqua satellites. First, a method to retrieve directional LSEs from multi-satellite data is developed based on the radiative transfer model. Then, the directional LSEs are retrieved from the combined MERSI/FY-3D/E, MODIS/Aqua and MODIS/Terra data in January 1~16 and July 1~16 of 2022 over a study area with longitude from 100°E to 130°E and latitude from 20°N to 50°N. Finally, the retrieved LSEs are, respectively, cross-validated with the MODIS/Terra land surface temperature (LST) and emissivity 8-day level 3 global 0.05° V61 (MOD11C2) product and the MODIS/Terra LST/3-band emissivity 8-day level 3 global 0.05° V61 (MOD21C2) product over the entire study area, and validated against the in-situ data at three true desert and semi-arid sites. The results show that the multi-satellite data provide more information of view angles and solar angles, which makes the determination of bi-directional reflectance distribution function model and the LSE retrieval more robust. The LSEs retrieved in this work have strong dependence on time and land cover types. Over the vegetated areas, the LSEs retrieved in this work basically agree with the MOD11C2 and MOD21C2 products, while over the true desert and semi-arid areas, the LSEs in the MOD11C2 and MOD21C2 products are obviously overestimated, especially the MOD11C2 product, but the LSEs in this work are consistent with the in-situ data. In general, the method developed in this work is valid and the retrieved LSEs are accurate.
Yi-Cheng Wang, Geng-Ming Jiang
IEEE Trans. Geosci. Remote. Sens.1
2023 Preserving Phonemic Distinctions For Ordinal Regression: A Novel Loss Function For Automatic Pronunciation Assessment
abstract
Automatic pronunciation assessment (APA) manages to quantify the pronunciation proficiency of a second language (L2) learner in a language. Prevailing approaches to APA normally leverage neural models trained with a regression loss function, such as the mean-squared error (MSE) loss, for proficiency level prediction. Despite most regression models can effectively capture the ordinality of proficiency levels in the feature space, they are confronted with a primary obstacle that different phoneme categories with the same proficiency level are inevitably forced to be close to each other, retaining less phoneme-discriminative information. On account of this, we devise a phonemic contrast ordinal (PCO) loss for training regression-based APA models, which aims to preserve better phonemic distinctions between phoneme categories meanwhile considering ordinal relationships of the regression target output. Specifically, we introduce a phoneme-distinct regularizer into the MSE loss, which encourages feature representations of different phoneme categories to be far apart while simultaneously pulling closer the representations belonging to the same phoneme category by means of weighted distances. An extensive set of experiments carried out on the speechocean 762 benchmark dataset demonstrate the feasibility and effectiveness of our model in relation to some existing state-of-the-art models.
Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Jiun-Ting Li, Chi-Han Lin, Berlin Chen
ASRU3
2023 Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and Diagnosis
abstract
Mispronunciation detection and diagnosis (MDD) manages to pinpoint phone-level erroneous pronunciation segmentations and provide instant and informative diagnostic feedback to L2 (second-language) learners. Among the various modeling paradigms for MDD, dictation-based neural methods have recently become a de facto standard, which identifies pronunciation errors and returns diagnostic feedback at the same time by aligning the recognized phone sequence uttered by an L2 learner to the corresponding canonical phone sequence of a given text prompt. Despite their decent efficacy, dictation-based methods have at least two downsides. First, the dictation process and alignment process are made independent of each other, often resulting in a poor diagnostic feedback. Second, prior knowledge about the articulation traits of the canonical phones in the text prompt is not fully utilized in MDD. On account of this, we propose a novel end-to-end MDD method that can streamline the dictation process and the alignment process in a non-autoregressive manner. In addition, knowledge about phone-level articulation traits are extracted with a graph convolutional network (GCN) to obtain more discriminative phonetic embeddings so as to promote the MDD performance. An extensive set of experiments conducted on the L2-ARCTIC benchmark dataset suggest the feasibility and effectiveness of our approach in relation to competitive baselines.
Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen
ICASSP3
2023 Directional Land Surface Emissivity Retrieval from Combined FY-3d/E MERSI-2 and Modis Infrared Data
abstract
In this paper, a method to retrieve directional land surface emissivity (LSE) from combined infrared data acquired by the advanced MEdium Resolution Spectral Imager (MERSI-2) on Fengyun 3D and 3E (FY-3D/E) and the Moderate-resolution Imaging Spectroradiometer (MODIS) on Terra and Aqua is developed. First, the observation differences due to different spectral response functions are removed by radiative transfer modeling, and FY-3D/E MERSI-2 observations and Aqua MODIS observations are transferred to effective MODIS measurements. Then, according to radiative transfer equation, the radiances at top of atmosphere are atmospherically corrected to ground level, in which the atmospheric parameters are calculated using the MODerate spectral resolution atmospheric TRANsmittance and radiance code (MODTRAN) and the fifth generation of European Centre for Medium-Range Weather Forecast (ECMWF) atmospheric reanalysis (ERA5) data. Next, based on the concept of temperature independent spectral indices (TISI), the bi-directional reflectances in the middle infrared (MIR) channel are derived. After that, according to the RossThick-LiSparse-R model and Kirchhoff’s law, the LSE in the MIR channel is retrieved. Finally, the LSE in the thermal infrared channel is deduced from the LSE in the MIR channel using the TISI concept again.
Yi-Cheng Wang, Geng-Ming Jiang
IGARSS1
2019 Serpentine: A Self-Powered Reversibly Deformable Cord Sensor for Human Input
abstract
We introduce Serpentine, a self-powered sensor that is a reversibly deformable cord capable of sensing a variety of human input. The material properties and structural design of Serpentine allow it to be flexible, twistable, stretchable and squeezable, enabling a broad variety of expressive input modalities. The sensor operates using the principle of Triboelectric Nanogenerators (TENG), which allows it to sense mechanical deformation without an external power source. The affordances of the cord include six interactions---Pluck, Twirl, Stretch, Pinch, Wiggle and Twist. Serpentine demonstrates the ability to simultaneously recognize these inputs through a single physical interface. A 12-participant user study illustrates 95.7% accuracy for a user-dependent recognition model using a realtime system and 92.17% for user-independent offline detection. We conclude by demonstrating how Serpentine can be employed in everyday ubiquitous computing applications.
Fereshteh Shahmiri, Chaoyu Chen, Anandghan Waghmare, Dingtian Zhang, Shivan Mittal, Steven L. Zhang, Yi-Cheng Wang, Zhong Lin Wang, Thad Starner, Gregory D. Abowd
CHI7