EDBT 2026 Demo / reviewers in the wild / expert
Liang-Chih Yu
dblp:68/5894
· DBLP profile ↗
52ranked-venue papers
21as first author
19since 2021 · last 2026
0000-0003-1443-4347ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 15 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningabstractLarge reasoning models (LRMs) improve performance at test time by thinking longer, but this often leads to overthinking and high computational cost. To address this, recent reinforcement learning (RL) methods adopt outcome-level rewards, such as rule- or prompt-based signals, that favor shorter correct reasoning paths but often overlook reasoning quality. While such rewards neglect intermediate reasoning, dense supervision from process reward models (PRMs) has proven more effective in promoting coherent and high-quality reasoning. However, static PRM supervision introduces two challenges: reward hacking, since fixed rewards poorly capture global reasoning objectives, and the high training cost of obtaining dense reward labels at scale. To overcome these issues, we propose Step Group Relative Policy Optimization (Step-GRPO), a GRPO-based method that integrates step-level PRM signals into sparse trajectory-level feedback, avoiding costly step-level supervision while improving reasoning quality beyond accuracy. In addition, Step-GRPO employs a step-attention mechanism that captures inter-step dependencies and emphasizes critical reasoning steps, effectively mitigating reward hacking. We apply Step-GRPO to train large language models and observe consistent gains in reasoning quality, accuracy, and shorter reasoning traces across multiple math benchmarks, outperforming reinforcement learning baselines at substantially lower cost. Notably, the proposed model achieves 36.7 percent accuracy on AIME 2024 with 11,000 training samples and a training cost of 38 US dollars, surpassing baselines that require over 1,000 US dollars and more than 40,000 samples, demonstrating strong cost-effectiveness and scalability. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
AAAI | 3 |
| 2026 | DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment AnalysisabstractLung-Hao Lee, Liang-Chih Yu, Natalia V Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng, Jin Wang, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lung-Hao Lee, Liang-Chih Yu, Natalia V. Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng 0001, Jin Wang 0008, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad |
ACL (1) | 2 |
| 2026 | Self-verified user simulator via code-based interpretation in task-oriented dialogues
Xiang Luo 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian PerspectiveabstractCurrent neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods affect pre-trained language models (PLMs) remains unclear. This study revisits the assumption that non-IID information enhances PLMs to achieve performance improvements from a Bayesian perspective, which unearths and integrates non-IID and IID features. Furthermore, we proposed a multi-attribute multi-grained framework for PLM adaptations (M2A), which combines multi-attribute and multi-grained views to mitigate uncertainty in a lightweight manner. We evaluate M2A through prevalent text-understanding datasets and demonstrate its superior performance, mainly when data are implicitly non-IID, and PLMs scale larger. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 3 |
| 2025 | Topology-of-Question-Decomposition: Enhancing Large Language Models with Information Retrieval for Knowledge-Intensive TasksabstractLarge language models (LLMs) are increasingly deployed for general problem-solving across various domains yet remain constrained to chaining immediate reasoning steps and depending solely on parametric knowledge. Integrating an information retrieval system directly into the reasoning process of LLMs can improve answer accuracy but might disrupt the natural reasoning sequence. Consequently, LLMs may underperform in complex, knowledge-intensive tasks requiring multiple reasoning steps, extensive real-world knowledge, or critical initial decisions. To overcome these challenges, we introduce a novel framework, Topology-of-Question-Decomposition (ToQD), which activates retrieval only when necessary. Globally, ToQD guides LLMs in constructing a topology graph from the input question, each node representing a sub-question. Locally, ToQD employs self-verify inference to determine whether a sub-question should retrieve relevant documents, necessitate further decomposition, or directly provide an answer. Experiments demonstrate that ToQD achieves superior performance and robustness in complex, knowledge-intensive tasks, significantly enhancing system response efficiency. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 3 |
| 2024 | Personalized LoRA for Human-Centered Text UnderstandingabstractEffectively and efficiently adapting a pre-trained language model (PLM) for human-centered text understanding (HCTU) is challenging since user tokens are million-level in most personalized applications and do not have concrete explicit semantics. A standard and parameter-efficient approach (e.g., LoRA) necessitates memorizing numerous suits of adapters for each user. In this work, we introduce a personalized LoRA (PLoRA) with a plug-and-play (PnP) framework for the HCTU task. PLoRA is effective, parameter-efficient, and dynamically deploying in PLMs. Moreover, a personalized dropout and a mutual information maximizing strategies are adopted and hence the proposed PLoRA can be well adapted to few/zero-shot learning scenarios for the cold-start issue. Experiments conducted on four benchmark datasets show that the proposed method outperforms existing methods in full/few/zero-shot learning scenarios for the HCTU task, even though it has fewer trainable parameters. For reproducibility, the code for this paper is available at: https://github.com/yoyo-yun/PLoRA. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 3 |
| 2024 | SoftMCL: Soft Momentum Contrastive Learning for Fine-grained Sentiment-aware Pre-trainingabstractThe pre-training for language models captures general language understanding but fails to distinguish the affective impact of a particular context to a specific word. Recent works have sought to introduce contrastive learning (CL) for sentiment-aware pre-training in acquiring affective information. Nevertheless, these methods present two significant limitations. First, the compatibility of the GPU memory often limits the number of negative samples, hindering the opportunities to learn good representations. In addition, using only a few sentiment polarities as hard labels, e.g., positive, neutral, and negative, to supervise CL will force all representations to converge to a few points, leading to the issue of latent space collapse. This study proposes a soft momentum contrastive learning (SoftMCL) for fine-grained sentiment-aware pre-training. Instead of hard labels, we introduce valence ratings as soft-label supervision for CL to fine-grained measure the sentiment similarities between samples. The proposed SoftMCL conducts CL on both the word- and sentence-level to enhance the model’s ability to learn affective information. A momentum queue was introduced to expand the contrastive samples, allowing storing and involving more negatives to overcome the limitations of hardware platforms. Extensive experiments were conducted on four different sentiment-related tasks, which demonstrates the effectiveness of the proposed SoftMCL method. The code and data of the proposed SoftMCL is available at: https://www.github.com/wangjin0818/SoftMCL/. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
LREC/COLING | 2 |
| 2024 | Improving Personalized Sentiment Representation with Knowledge-enhanced and Parameter-efficient Layer NormalizationabstractExisting studies on personalized sentiment classification consider a document review as an overall text unit and incorporate backgrounds (i.e., user and product information) to learn sentiment representation. However, it is difficult when these methods meet the current pretrained language models (PLMs) owing to quadratic costs that increase with text length and heterogeneous mixes of randomly initialized background information and textual information initialized from well-pretrained checkpoints during information incorporation. To address these problems, we propose a knowledge-enhanced and parameter-efficient layer normalization (E2LN) for efficient and effective review modeling via leveraging LN in transformer structures. Initially, a knowledge base is introduced that stores well-pretrained checkpoints, structured text information, and background information. Based on such a knowledge base, the ability of LN can be magnified as being a crucial component of transformer structure and then improve the performance of PLMs in downstream tasks. Moreover, the proposed E2LN can make PLMs capable of modeling long document reviews and incorporating background information with parameter-efficient fine-tuning and knowledge injecting. Extensive experimental results were obtained for three document-level sentiment classification benchmark datasets. By comparing the results, the effectiveness and efficiency of the proposed model was demonstrated. Code and Data are released at https://github.com/yoyo-yun/E2LN. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
LREC/COLING | 3 |
| 2024 | Encoding Syntactic Information into Transformers for Aspect-Based Sentiment Triplet ExtractionabstractAspect-based sentiment triplet extraction(ASTE) aims to extract triplets consisting of aspect terms and their associated opinion terms and sentiment polarities from sentences, a relatively new and challenging subtask of aspect-based sentiment analysis (ABSA). Previous studies have used either pipeline models or unified tagging schema models. These models ignore the syntactic relationships between the aspect and its corresponding opinion words, which leads them to mistakenly focus on syntactically unrelated words. One feasible option is to use a graph convolution network (GCN) to exploit syntactic information by propagating the representation from the opinion words to the aspect. However, such a method considers all syntactic dependencies to be of the same type and thus may still incorrectly associate unrelated words to the target aspect through the iterations of graph convolutional propagation. Herein, a syntax-aware transformer (SA-Transformer) is proposed to extend the GCN strategy by fully exploiting the dependency types of edges to block inappropriate propagation. The proposed approach can obtain different representations and weights even for edges with the same dependency type according to their adjacent dependency type of edges. Instead of using a GCN layer, we used anL-layer SA transformer to encode syntactic information in the word-pair representation to improve performance. Experimental results on four benchmark datasets show that the proposed model outperforms various previous models for ASTE. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Multimodality Self-distillation for Fast Inference of Vision and Language Pretrained ModelsabstractThe computational cost of the vision and language pretrained models (VL-PTMs) limits their deployment in resource-constrained devices that require low latency. One existing solution is to apply the early exiting (EE) strategy to accelerate the inference. This technique can force model prediction using only a few former transformer layers. However, these former layers behave differently with the final classifier, inevitably resulting in performance decline. To counter such limitation, self-distillation has been commonly introduced to enhance the representation abilities of the EE classifiers. This results in a semantic gap since EE classifiers are directly trained to mimic the outputs of the final classifier without access to the modality-specific behaviors. This study proposes a multimodality self-distillation method for the fast inference of VL-PTMs. To fill the semantic gap between modalities, we split the multimodalities into separate modalities and added them as extra inputs to encourage the effective distillation of each modality. Furthermore, the mean squared error (MSE) is introduced to minimize the distance of feature maps and further enhance the representation ability of the EE classifiers. Experiments show that the proposed method outperforms the previous EE strategies with the same inference time, and performs competitively even if the model exited very early. Jun Kong 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2023 | Learning to Memorize Entailment and Discourse Relations for Persona-Consistent DialoguesabstractMaintaining engagement and consistency is particularly important in dialogue systems. Existing works have improved the performance of dialogue systems by intentionally learning interlocutor personas with sophisticated network structures. One issue with this approach is that it requires more personal corpora with annotations. Additionally, these models typically perform the next utterance prediction to generate a response but neglect the discourse coherence in the entire conversation. To address these issues, this study proposes a method of learning to memorize entailment and discourse relations for persona-consistent dialogue tasks. Entailment text pairs in natural language inference dataset were applied to learn latent entailment relations as external memories by premise-to-hypothesis generation task. Furthermore, an internal memory with a similar architecture was applied to the discourse information in the dialogue. Placing orthogonality restrictions on these two memory spaces ensures that the latent entailment relations remain dialogue-independent. Both memories collaborate to obtain entailment and discourse representation for the generation, allowing a deeper understanding of both consistency and coherence. Experiments on two large public datasets, PersonaChat and DSTC7-AVSD, demonstrated the effectiveness of the proposed method. Both automatic and human evaluations indicate that the proposed model outperforms several strong baselines in terms of both persona consistency and response coherence. Our source code is availabled at https://github.com/Chenrj233/LMEDR. Ruijun Chen 0001, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
AAAI | 3 |
| 2023 | Decoupled variational autoencoder with interactive attention for affective text generation
Ruijun Chen 0001, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Accelerating Inference for Pretrained Language Models by Unified Multi-Perspective Early ExitingabstractConditional computation algorithms, such as the early exiting (EE) algorithm, can be applied to accelerate the inference of pretrained language models (PLMs) while maintaining competitive performance on resource-constrained devices. However, this approach is only applied to the vertical architecture to decide which layers should be used for inference. Conversely, the operation of the horizontal perspective is ignored, and the determination of which tokens in each layer should participate in the computation fails, leading to a high redundancy for adaptive inference. To address this limitation, a unified horizontal and vertical multi-perspective early exiting (MPEE) framework is proposed in this study to accelerate the inference of transformer-based models. Specifically, the vertical architecture uses recycling EE classifier memory and weighted self-distillation to enhance the performance of the EE classifiers. Then, the horizontal perspective uses recycling class attention memory to emphasize the informative tokens. Conversely, the tokens with less information are truncated by weighted fusion and isolated from the following computation. Based on this, both horizontal and vertical EE are unified to obtain a better tradeoff between performance and efficiency. Extensive experimental results show that MPEE can achieve higher acceleration inference with competent performance than existing competitive methods. Jun Kong 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 3 |
| 2022 | Knowledge Distillation with Reptile Meta-Learning for Pretrained Language Model CompressionabstractThe billions, and sometimes even trillions, of parameters involved in pre-trained language models significantly hamper their deployment in resource-constrained devices and real-time applications. Knowledge distillation (KD) can transfer knowledge from the original model (i.e., teacher) into a compact model (i.e., student) to achieve model compression. However, previous KD methods have usually frozen the teacher and applied its immutable output feature maps as soft labels to guide the student’s training. Moreover, the goal of the teacher is to achieve the best performance on downstream tasks rather than knowledge transfer. Such a fixed architecture may limit the teacher’s teaching and student’s learning abilities. Herein, a knowledge distillation method with reptile meta-learning is proposed to facilitate the transfer of knowledge from the teacher to the student. The teacher can continuously meta-learn the student’s learning objective to adjust its parameters for maximizing the student’s performance throughout the distillation process. In this way, the teacher learns to teach, produces more suitable soft labels, and transfers more appropriate knowledge to the student, resulting in improved performance. Unlike previous KD using meta-learning, the proposed method only needs to calculate the first-order derivatives to update the teacher, leading to lower computational cost but better convergence. Extensive experiments on the GLUE benchmark show the competitive performance achieved by the proposed method. For reproducibility, the code for this paper is available at: https://github.com/maxinge8698/ReptileDistil. Xinge Ma, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 3 |
| 2022 | Hierarchical template transformer for fine-grained sentiment controllable generation
Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Inf. Process. Manag. | 3 |
| 2022 | Contextual sentiment embeddings via bi-directional GRU language modelabstractCompared with conventional word embeddings, sentiment embeddings can distinguish words with similar contexts but opposite sentiment. They can be used to incorporate sentiment information from labeled corpora or lexicons by either end-to-end training or sentiment refinement. However, these methods present two major limitations. First, traditional approaches provide a fixed representation to each word but ignore the alternation of word meaning in different contexts. As a result, the polarity of a certain emotional word may vary with context, but will be assigned with a same representation. Another problem is the handling of out-of-vocabulary (OOV) or informal-writing sentiment words that would be assigned generic vectors (e.g., ). In addition, if affective words are not included in affective corpora or lexicons, they would be treated as neutral. Using such low-quality embeddings for building a neural model will reduce performance. This study proposes a training model of contextual sentiment embeddings. A stacked two-layer GRU model was used as the language model, simultaneously trained to incorporate semantic and sentiment information from labeled corpora and lexicons. To deal with OOV or informal-writing sentiment words, the WordPiece tokenizer was used to divide the text into subwords. The resulting model can be transferred to downstream applications by either feature extractor or fine-tuning. The results show that the proposed model can handle unseen or informal writing sentiment words and thus outperforms previously proposed methods. Jin Wang 0008, You Zhang 0002, Liang-Chih Yu, Xuejie Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Explainable detection of adverse drug reaction with imbalanced data distributionabstractAnalysis of health-related texts can be used to detect adverse drug reactions (ADR). The greatest challenge for ADR detection lies in imbalanced data distributions where words related to ADR symptoms are often minority classes. As a result, trained models tend to converge to a point that strongly biases towards the majority class and then ignores the minority class. Since the most used cross-entropy criteria is an approximation to accuracy, the model focuses more readily on the majority class to achieve high accuracy. To address this issue, existing methods apply either oversampling or down-sampling strategies to balance the data distribution and exploit the most difficult samples of the minority class. However, increasing or reducing the number of individual tokens alone in sequence labeling tasks will result in the loss of the syntactic relations of the sentence. This paper proposes a weighted variant of conditional random field (CRF) for data-imbalanced sequence labeling tasks. Such a weighting strategy can alleviate data distribution imbalances between majority and minority classes. Instead of using softmax in the output layer, the CRF can capture the relationship of labels between tokens. The locally interpretable model-agnostic explanations (LIME) algorithm was applied to investigate performance differences between models with and without the weighted loss function. Experimental results on two different ADR tasks show that the proposed model outperforms previously proposed sequence labeling methods. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
PLoS Comput. Biol. | 2 |
| 2022 | Chinese EmoBank: Building Valence-Arousal Resources for Dimensional Sentiment AnalysisabstractAn increasing amount of research has recently focused on dimensional sentiment analysis that represents affective states as continuous numerical values on multiple dimensions, such as valence-arousal (VA) space. Compared to the categorical approach that represents affective states as distinct classes (e.g., positive and negative), the dimensional approach can provide more fine-grained (real-valued) sentiment analysis. However, dimensional sentiment resources with valence-arousal ratings are very rare, especially for the Chinese language. Therefore, this study aims to: (1) Build a Chinese valence-arousal resource called Chinese EmoBank, the first Chinese dimensional sentiment resource featuring various levels of text granularity including 5,512 single words, 2,998 multi-word phrases, 2,582 single sentences, and 2,969 multi-sentence texts. The valence-arousal ratings are annotated by crowdsourcing based on the Self-Assessment Manikin (SAM) rating scale. A corpus cleanup procedure is then performed to improve annotation quality by removing outlier ratings and improper texts. (2) Evaluate the proposed resource using different categories of classifiers such as lexicon-based, regression-based, and neural-network-based methods, and comparing their performance to a similar evaluation of an English dimensional sentiment resource. Lung-Hao Lee, Jian-Hong Li, Liang-Chih Yu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | A multi-dimensional relation model for dimensional sentiment analysisabstractDimensional sentiment analysis has received considerable attention because it can represent affective states as continuous numerical values on multiple dimensions such as valence (positive–negative) and arousal (excited–calm). Compared to the categorical approach, which represents affective states as several discrete classes (e.g., positive and negative), the dimensional approach can provide more fine-grained (real-valued) sentiment analysis. Traditional approaches to predicting dimensional sentiment scores typically treat each dimension independently without consideration of relations between dimensions. In fact, different dimensions may correlate with each other. For example, expressions with a higher valence score usually have a higher arousal score, And higher irony expressions usually have a lower valence score. Such relations between dimensions are useful for dimension score prediction. To this end, this study proposes a multi-dimensional relation model to incorporate relations between dimensions into deep neural networks for dimension score prediction. The proposed method has two modes: internal and external. The internal mode incorporates the relations between dimensions into sentence representations before prediction, whereas the external mode builds a linear regression model that can capture the relations between dimensions to refine the predicted scores after prediction. To evaluate the proposed method, we created a Chinese three-dimensional corpus with valence-arousal-irony (VAI) ratings. Experiments using various neural network architectures demonstrate that the proposed multi-dimensional relation model outperformed those that treat each dimension independently. In addition, the internal mode outperformed the external mode, and a combination of the two modes achieved the best performance. Housheng Xie, Wei Lin 0023, Shuying Lin, Jin Wang 0008, Liang-Chih Yu |
Inf. Sci. | 5 |
| 2020 | Pipelined Neural Networks for Phrase-Level Sentiment Intensity PredictionabstractLinguistic modifiers such as negators (e.g., not), intensifiers (e.g., very) and modals (e.g., would) are commonly used in expressing opinions. These modifiers play an important role in recognizing the sentiment intensity of multi-word phrases because they may lead to an intensity shift and polarity reversal for the words they modify. Appropriately modeling the effect of such modifiers on the intensity shift can greatly improve the performance of phrase-level sentiment intensity prediction. To this end, this paper proposes two neural network (NN) models organized in a pipelined fashion to determine 1) the intensity of individual words and 2) the shift weights of modifiers representing the degrees of intensity change for the words they modify. The intensity of a phrase can then be determined by combining the intensity of the constituent word and the shift weight of the modifier within the phrase. When measuring the word intensity, the first NN model introduces a hidden layer as a filter to select appropriate similar seed words in the prediction process. Automatic word intensity prediction can address the unknown intensities of words not covered in sentiment lexicons. In learning the modifier weights, the second NN model considers both the weights of individual modifiers and groups of modifiers to capture various intensity shift effects caused by them. Experiments on a SemEval-2016 dataset showed that the proposed method yielded better prediction performance for both single words and multi-word phrases. Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Tree-Structured Regional CNN-LSTM Model for Dimensional Sentiment AnalysisabstractDimensional sentiment analysis aims to recognize continuous numerical values in multiple dimensions such as the valence-arousal (VA) space. Compared to the categorical approach that focuses on sentiment classification such as binary classification (i.e., positive and negative), the dimensional approach can provide a more fine-grained sentiment analysis. This article proposes a tree-structured regional CNN-LSTM model consisting of two parts: regional CNN and LSTM to predict the VA ratings of texts. Unlike a conventional CNN which considers a whole text as input, the proposed regional CNN uses a part of the text as a region, dividing an input text into several regions such that the useful affective information in each region can be extracted and weighted according to their contribution to the VA prediction. Such regional information is sequentially integrated across regions using LSTM for VA prediction. By combining the regional CNN and LSTM, both local (regional) information within sentences and long-distance dependencies across sentences can be considered in the prediction process. To further improve performance, a region division strategy is proposed to discover task-relevant phrases and clauses to incorporate structured information into VA prediction. Experimental results on different corpora show that the proposed method outperforms lexicon-, regression-, conventional NN and other structured NN methods proposed in previous studies. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Investigating Dynamic Routing in Tree-Structured LSTM for Sentiment AnalysisabstractJin Wang, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Refining Word Embeddings Using Intensity Scores for Sentiment AnalysisabstractWord embeddings that provide continuous low-dimensional vector representations of words have been extensively used for various natural language processing tasks. However, existing context-based word embeddings such as Word2vec and GloVe typically fail to capture sufficient sentiment information, which may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance. To tackle this problem, recent studies have suggested learning sentiment embeddings to incorporate the sentiment polarity (positive and negative) information from labeled corpora. This study adopts another strategy to learn sentiment embeddings. Instead of creating a new word embedding from labeled corpora, we propose a word vector refinement model to refine existing pretrained word vectors using real-valued sentiment intensity scores provided by sentiment lexicons. The idea of the refinement model is to improve each word vector such that it can be closer in the lexicon to both semantically and sentimentally similar words (i.e., those with similar intensity scores) and further away from sentimentally dissimilar words (i.e., those with dissimilar intensity scores). An obvious advantage of the proposed method is that it can be applied to any pretrained word embeddings. In addition, the intensity scores can provide more fine-grained (real-valued) sentiment information than binary polarity labels to guide the refinement process. Experimental results show that the proposed refinement model can improve both conventional word embeddings and previously proposed sentiment embeddings for binary, ternary, and fine-grained sentiment classification on the SemEval and Stanford Sentiment Treebank datasets. Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Refining Word Embeddings for Sentiment AnalysisabstractWord embeddings that can capture semantic and syntactic information from contexts have been extensively used for various natural language processing tasks.However, existing methods for learning contextbased word embeddings typically fail to capture sufficient sentiment information.This may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance.Therefore, this study proposes a word vector refinement model that can be applied to any pre-trained word vectors (e.g., Word2vec and GloVe).The refinement model is based on adjusting the vector representations of words such that they can be closer to both semantically and sentimentally similar words and further away from sentimentally dissimilar words.Experimental results show that the proposed method can improve conventional word embeddings and outperform previously proposed sentiment embeddings for both binary and fine-grained classification on Stanford Sentiment Treebank (SST). Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
EMNLP | 1 |
| 2017 | Chinese Grammatical Error Detection Using a CNN-LSTM Model
Lung-Hao Lee, Bo-Lin Lin, Liang-Chih Yu, Yuen-Hsien Tseng |
ICCE | 3 |
| 2016 | Building Chinese Affective Resources in Valence-Arousal Dimensions
Liang-Chih Yu, Lung-Hao Lee, Jin Wang 0008, Yunchao He, K. Robert Lai, Xuejie Zhang 0002 |
HLT-NAACL | 1 |
| 2016 | Locally weighted linear regression for cross-lingual valence-arousal prediction of affective words
Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
Neurocomputing | 2 |
| 2016 | Near-synonym substitution using a discriminative vector space model
Liang-Chih Yu, Lung-Hao Lee, Jui-Feng Yeh, Hsiu-Min Shih, Yu-Ling Lai |
Knowl. Based Syst. | 1 |
| 2016 | Community-Based Weighted Graph Model for Valence-Arousal Prediction of Affective WordsabstractCompared to the categorical approach that represents affective states as several discrete classes (e.g., positive and negative), the dimensional approach represents affective states as continuous numerical values in multiple dimensions, such as the valence-arousal (VA) space, thus allowing for more fine-grained sentiment analysis. In building dimensional sentiment applications, affective lexicons with VA ratings are useful resources but are still very rare. Several semi-supervised methods such as the kernel method, linear regression, and the pagerank algorithm have been investigated to automatically determine the VA ratings of affective words from a set of semantically similar seed words. These methods suffer from two major limitations. First, they apply an equal weight to all seeds similar to an unseen word in predicting its VA ratings. Second, even similar seeds may have quite different ratings (or an inverse polarity) of valence/arousal to the unseen word, thus reducing prediction performance. To overcome these limitations, this study proposes a community-based weighted graph model that can select seeds which are both similar to and have similar ratings (or the same polarity) with each unseen word to form a community (subgraph) so that its VA ratings can be estimated from such high-quality seeds using a weighted propagation scheme. That is, seeds more similar to unseen words contribute more to the estimation process. Experimental results show that the proposed method yields better prediction performance for both English and Chinese datasets. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A locally weighted method to improve linear regression for lexical-based valence-arousal predictionabstractText-based sentiment analysis is a growing research field in affective computing, driven by both commercial applications and academic interest. Continuous dimensional representations, such as valence-arousal (VA) space, can represent the affective state more precisely than discrete effective representations. In building dimensional sentiment applications, affective lexicons with valence-arousal ratings are useful resources but are still very rare. Therefore, recent studies have investigated the automatic development of VA lexicons using linear regression techniques. One of the major limitations of linear regression is the under-fitting problem which can cause a poor fit between the algorithm and the training data. To tackle this problem, this study proposes the use of a locally weighted linear regression (LWLR) model to predict the valence-arousal ratings of affective words. The locally weighted method performs a regression around the point of interest using only training data that are "local" to that point, and thus can reduce the impact of noise from unrelated training data. Experimental results show that the proposed method achieved better performance for VA word prediction. Jin Wang 0008, K. Robert Lai, Liang-Chih Yu, Xuejie Zhang 0002 |
ACII | 3 |
| 2015 | Who Will Pass? Analyzing Learner Behaviors in MOOCs
Shu-Fen Tseng, Yen-Wei Tsao, Liang-Chih Yu, Chien-Lung Chan, K. Robert Lai |
ICCE | 3 |
| 2014 | Identifying Emotion Labels from Psychiatric Social Texts Using Independent Component Analysis
Liang-Chih Yu, Chun-Yuan Ho |
COLING | 1 |
| 2014 | A Tagging Editor for Learner Corpora Annotation and Error AnalysisabstractIn this paper, we describe the development of the tagging editor for learner corpora annotation and computer-aided error analysis. We collect essays written by learners of Chinese as a foreign language for grammatical error annotation and correction. Our tagging editor is effective and enables the annotated corpus to be used in a shared task in ICCE 2014. Lung-Hao Lee, Kuei-Ching Lee, Li-Ping Chang, Yuen-Hsien Tseng, Liang-Chih Yu, Hsin-Hsi Chen |
ICCE | 5 |
| 2014 | Overview of Grammatical Error Diagnosis for Learning Chinese as a Foreign Language
Liang-Chih Yu, Lung-Hao Lee, Li-Ping Chang |
ICCE | 1 |
| 2013 | Independent component analysis for near-synonym choice
Liang-Chih Yu, Wei-Nan Chien |
Decis. Support Syst. | 1 |
| 2013 | Using a contextual entropy model to expand emotion words and their intensity for the sentiment classification of stock market news
Liang-Chih Yu, Jheng-Long Wu, Pei-Chann Chang, Hsuan-Shou Chu |
Knowl. Based Syst. | 1 |
| 2012 | Developing a Context-Supported Chinese Grammar Learning System in Mobile Environments
Yuen-Hsien Tseng, Chi-Hao Huang, Liang-Chih Yu, Yu-Ju Lan |
ICCE | 3 |
| 2012 | Combining Language and Speech Features to Predict Students' Emotions in E-Learning EnvironmentsabstractEmotions play an important role in e-learning environments. Text and speech have been recognized as convenient and natural means for expressing emotions, and are increasingly used in human-computer interaction interfaces for e-learning applications, indicating that language and speech could potentially be used to predict learner emotions. In this study, we investigate the use of speech and language features for automatic emotion recognition. A corpus of emotion-laden sentences was collected from student-teacher dialogs in the context of mathematics instruction. The corpus was then annotated to analyze emotion types as they occurred in e-learning applications. The speech and language features were then used to build several classifiers for emotion recognition. Experiments show that the two features combined yielded better results than either feature alone. In addition, among speech features, energy and formant are found to best contribute to successful classification. Liang-Chih Yu, Shou-Fang Liang, Wei-Hua Lin |
ICCE | 1 |
| 2011 | Analysis of Students' Emotion from a Text Corpus
Liang-Chih Yu, Shou-Fang Liang, Wei-Hua Lin, K. Robert Lai, Baw-Jhiune Liu |
ICCE | 1 |
| 2011 | A Baseline System for Chinese Near-Synonym Choice
Liang-Chih Yu, Wei-Nan Chien, Shih-Ting Chen |
IJCNLP | 1 |
| 2011 | Mining association language patterns using a distributional semantic model for negative life event classification
Liang-Chih Yu, Chien-Lung Chan, Chao-Cheng Lin, I-Chun Lin |
J. Biomed. Informatics | 1 |
| 2010 | Discriminative Training for Near-Synonym Substitution
Liang-Chih Yu, Hsiu-Min Shih, Yu-Ling Lai, Jui-Feng Yeh, Chung-Hsien Wu 0001 |
COLING | 1 |
| 2010 | Annotation and verification of sense pools in OntoNotes
Liang-Chih Yu, Chung-Hsien Wu 0001, Ru-Yng Chang, Chao-Hong Liu, Eduard H. Hovy |
Inf. Process. Manag. | 1 |
| 2010 | Sentence Correction Incorporating Relative Position and Parse Template Language ModelsabstractSentence correction has been an important emerging issue in computer-assisted language learning. However, existing techniques based on grammar rules or statistical machine translation are still not robust enough to tackle the common errors in sentences produced by second language learners. In this paper, a relative position language model and a parse template language model are proposed to complement traditional language modeling techniques in addressing this problem. A corpus of erroneous English-Chinese language transfer sentences along with their corrected counterparts is created and manually judged by human annotators. Experimental results show that compared to a state-of-the-art phrase-based statistical machine translation system, the error correction performance of the proposed approach achieves a significant improvement using human evaluation. Chung-Hsien Wu 0001, Chao-Hong Liu, Matthew Harris, Liang-Chih Yu |
IEEE Trans. Speech Audio Process. | 4 |
| 2009 | Psychiatric document retrieval using a discourse-aware model
Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
Artif. Intell. | 1 |
| 2008 | OntoNotes: Corpus Cleanup of Mistaken Agreement Using Word Sense Disambiguation
Liang-Chih Yu, Chung-Hsien Wu 0001, Eduard H. Hovy |
COLING | 1 |
| 2008 | HAL-Based Evolutionary Inference for Pattern Induction From Psychiatry Web ResourcesabstractNegative and stressful life events play a significant role in triggering depressive episodes. Psychiatric services that can identify such events efficiently are vital for mental health care and prevention. Meaningful patterns, e.g.,, must be extracted from psychiatric texts before these services can be provided. This study presents an evolutionary text-mining framework capable of inducing variable-length patterns from unannotated psychiatry Web resources. The proposed framework can be divided into two parts: 1) a cognitive motivated model such as hyperspace analog to language (HAL) and 2) an evolutionary inference algorithm (EIA). The HAL model constructs a high-dimensional context space to represent words as well as combinations of words. Based on the HAL model, the EIA bootstraps with a small set of seed patterns, and then iteratively induces additional relevant patterns. To avoid moving in the wrong direction, the EIA further incorporates relevance feedback to guide the induction process. Experimental results indicate that combining the HAL model and relevance feedback enables the EIA to not only induce patterns from the unannotated Web corpora, but also achieve useful results in a reasonable amount of time. The proposed framework thus significantly reduces reliance on annotated corpora. Liang-Chih Yu, Chung-Hsien Wu 0001, Jui-Feng Yeh, Fong-Lin Jang |
IEEE Trans. Evol. Comput. | 1 |
| 2008 | Extended probabilistic HAL with close temporal association for psychiatric query document retrievalabstractPsychiatric query document retrieval can assist individuals to locate query documents relevant to their depression-related problems efficiently and effectively. By referring to relevant documents, individuals can understand how to alleviate their depression-related symptoms according to recommendations from health professionals. This work presents an extended probabilistic Hyperspace Analog to Language ( epHAL ) model to achieve this aim. The epHAL incorporates the close temporal associations between words in query documents to represent word cooccurrence relationships in a high-dimensional context space. The information flow mechanism further combines the query words in the epHAL space to infer related words for effective information retrieval. The language model perplexity is considered as the criterion for model optimization. Finally, the epHAL is adopted for psychiatric query document retrieval, and indicates its superiority in information retrieval over traditional approaches. Jui-Feng Yeh, Chung-Hsien Wu 0001, Liang-Chih Yu, Yu-Sheng Lai |
ACM Trans. Inf. Syst. | 3 |
| 2007 | Topic Analysis for Psychiatric Document Retrieval
Liang-Chih Yu, Chung-Hsien Wu 0001, Chin-Yew Lin, Eduard H. Hovy, Chia-Ling Lin |
ACL | 1 |
| 2007 | Psychiatric Consultation Record Retrieval Using Scenario-Based Representation and Multilevel Mixture ModelabstractPsychiatric consultation record retrieval attempts to help people to efficiently and effectively locate the consultation records relevant to their depressive problems. Consultation records can also make people aware that they are not alone, because many individuals have suffered from the same or similar problems. Additionally, people can understand how to alleviate their depressive symptoms according to recommendations from health professionals. To achieve this goal, this paper proposes the use of a scenario-based representation, i.e., a symptom-based structural representation, to capture the depressive symptoms and their semantic relations, such as cause-effect and temporal relations, for understanding users' queries clearly. The symptoms and relations are identified from semantic mining and analysis of consultation records. The multilevel mixture model is adopted to estimate the relevance of queries and consultation records based on the structural information. Experimental results show that the proposed approach achieves higher precision than does a term-based flat representation. An experiment is also conducted to examine the effect of error propagation resulting from incorrect identification of symptoms and relations. Experimental results demonstrate that combining different approaches can improve the retrieval robustness. Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | HAL-Based Cascaded Model for Variable-Length Semantic Pattern Induction from Psychiatry Web Resources
Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
ACL | 1 |
| 2004 | Automated Alignment and Extraction of Bilingual Domain Ontology for Cross-Language Domain-Specific Applications
Jui-Feng Yeh, Chung-Hsien Wu 0001, Ming-Jun Chen, Liang-Chih Yu |
COLING | 4 |