EDBT 2026 Demo / reviewers in the wild / expert
Chu-Ren Huang
dblp:49/1619
· DBLP profile ↗
159ranked-venue papers
13as first author
39since 2021 · last 2026
0000-0002-8526-5520ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 154 · 13 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesabstractDo large language models (LLMs) truly acquire embodied cognition and cultural conventions from text?We introduce demonstratives, fundamental spatial expressions like "this/that" in English and "这/那" in Chinese, as a novel probe for grounded knowledge.Using 6,400 responses from 320 native speakers, we establish a human baseline: English speakers reliably distinguish proximal-distal referents but struggle with perspective-taking, while Chinese speakers switch perspectives fluently but tolerate distal ambiguity.In contrast, five state-ofthe-art LLMs fail to inherently understand the proximal-distal contrast and show no cultural differences, defaulting to English-centric reasoning.Our study contributes (i) a new task, based on demonstratives, as a new lens for evaluating embodied cognition and cultural conventions; (ii) empirical evidence of cross-cultural asymmetries in human interpretation; (iii) a new perspective on the egocentric-sociocentric debate, showing both orientations coexist but vary across languages; and (iv) a call to address individual variation in future model design.Model Gpt5. Emmanuele Chersoni, Chu-Ren Huang |
ACL (1) | 3 |
| 2026 | The Sensorimotor Norms for the Chinese Classifiers
Yimei Shao, Yu-Yin Hsu, Chu-Ren Huang |
LREC | 3 |
| 2026 | This One or That One? A Study on Accessibility via Demonstratives with Multimodal Large Language Models
Emmanuele Chersoni, Chu-Ren Huang |
LREC | 3 |
| 2026 | Sentimental image generation with image quality assessment
Xiaoyi Bao, Jinghang Gu, Chu-Ren Huang |
Pattern Recognit. | 4 |
| 2026 | Exploring Context-Free Opinion Grammar for Aspect-Based Sentiment AnalysisabstractUtilizing pre-trained generative models for sentiment element extraction has recently significantly enhanced aspect-based sentiment analysis benchmarks. Nonetheless, these models have two significant drawbacks: 1) high-computational cost in both the inference time and hardware requirement. 2) Lack of explicit modeling as they model the connections between sentiment elements with fragile natural or notational language target sequence. To overcome these challenges, we present a novel opinion tree parsing model designed to swiftly parse sentiment elements from an opinion tree. This approach not only accelerates the process but also explicitly unveils a more comprehensive and fully articulated aspect-level sentiment structure. Our method begins by introducing a pioneering context-free opinion grammar to standardize the opinion tree structure. Subsequently, we leverage a neural chart-based opinion tree parser to thoroughly explore the interconnections among sentiment elements and parse them into a structured opinion tree. Extensive experiments underscore the effectiveness of our proposed model and the capability of the opinion tree parser, particularly when coupled with the introduced context-free opinion grammar. Crucially, the results confirm the superior speed of our model compared to the SOTA baselines. Xiaoyi Bao, Jinghang Gu, Chu-Ren Huang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Revisiting Classical Chinese Event Extraction with Ancient Literature InformationabstractThe research on classical Chinese event extraction trends to directly graft the complex modeling from English or modern Chinese works, neglecting the utilization of the unique characteristic of this language. We argue that, compared with grafting the sophisticated methods from other languages, focusing on classical Chinese’s inimitable source of Ancient Literature could provide us with extra and comprehensive semantics in event extraction. Motivated by this, we propose a Literary Vision-Language Model (VLM) for classical Chinese event extraction, integrating with literature annotations, historical background and character glyph to capture the inner- and outer-context information from the sequence. Extensive experiments build a new state-of-the-art performance in the GuwenEE, CHED datasets, which underscores the effectiveness of our proposed VLM, and more importantly, these unique features can be obtained precisely at nearly zero cost. Xiaoyi Bao, Jinghang Gu, Chu-Ren Huang |
ACL (1) | 4 |
| 2025 | CalligraphicOCR for Chinese Calligraphy RecognitionabstractWith thousand years of history, calligraphy serve as one of the representative symbols of Chinese culture. Increasing works try to digitize calligraphy by recognizing the context of calligraphy for better preservation and propagation. However, previous works stick to isolated single character recognition, not only requires unpractical manual splitting into characters, but also abandon the enriched context information that could be supplementary. To this end, we construct the pioneering end-to-end calligraphy recognition benchmark dataset, this dataset is challenging due to both the visual variations such as different writing styles and the textual understanding such as the domain shift in semantics. We further propose CalligraphicOCR (COCR) equipped with calligraphic image augmentation and action-based corrector targeted at the challenging root of this setting. Experiments demonstrate the advantage of our proposed model over cutting-edge baselines, underscoring the necessity of introducing this new setting, thereby facilitating a solid precondition for protecting and propagating the already scarce resources. Xiaoyi Bao, Jinghang Gu, Chu-Ren Huang |
EMNLP | 4 |
| 2025 | The Evolutionary Mechanisms of Transitivization in Mandarin VO Compounds: A Corpus-Driven Study of Competing Alternations
Menghan Jiang, Chu-Ren Huang |
PACLIC | 2 |
| 2025 | Sensory and Affective Dimensions in Mandarin Monosyllabic Adjectives
Yimei Shao, Yu-Yin Hsu, Chu-Ren Huang |
PACLIC | 3 |
| 2025 | Computational Linguistic Approach to Empathy and its Language Communication Pattern
Mingyu Wan, Chu-Ren Huang |
PACLIC | 3 |
| 2024 | From Text to Historical Ecological Knowledge: The Construction and Application of the Shan Jing Knowledge BaseabstractTraditional Ecological Knowledge (TEK) has been recognized as a shared cultural heritage and a crucial instrument to tackle today’s environmental challenges. In this paper, we deal with historical ecological knowledge, a special type of TEK that is based on ancient language texts. In particular, we aim to build a language resource based on Shanhai Jing (The Classic of Mountains and Seas). Written 2000 years ago, Shanhai Jing is a record of flora and fauna in ancient China, anchored by mountains (shan) and seas (hai). This study focuses on the entities in the Shan Jing part and builds a knowledge base for them. We adopt a pattern-driven and bottom-up strategy to accommodate two features of the source: highly stylized narrative and juxtaposition of knowledge from multiple domains. The PRF values of both entity and relationship extraction are above 96%. Quality assurance measures like entity disambiguation and resolution were done by domain experts. Neo4j graph database is used to visualize the result. We think the knowledge base, containing 1432 systematically classified entities and 3294 relationships, can provide the foundation for the construction of a historical ecological knowledge base of China. Additionally, the ruled-based text-matching method can be helpful in ancient language processing. Chu-Ren Huang, Xin-Lan Jiang |
LREC/COLING | 2 |
| 2024 | Be Helpful but Don't Talk too Much - Enhancing Helpfulness in Conversations through Relevance in Multi-Turn Emotional SupportabstractFor a conversation to help and support, speakers should maintain an "effect-effort" trade-off.As outlined in the gist of "Cognitive Relevance Principle", helpful speakers should optimize the "cognitive relevance" through maximizing the "cognitive effects" and minimizing the "processing effort" imposed on listeners.Although preference learning methods provide a boon for studies concerning "effect-optimization", none have delved into "effort-optimization" which is pivotal to the acquisition of "optimal relevance" for emotional support conversation agents.To address this gap, we integrate the "Cognitive Relevance Principle" into emotional support agents in the environment of multi-turn conversation.The results demonstrate a significant and robust improvement against the baseline systems with respect to response quality, human-likedness, and supportiveness.This study offers compelling evidence for the effectiveness of the "Relevance Principle" in generating human-like, helpful, and harmless emotional support conversations. Yu-Yin Hsu, Chu-Ren Huang |
EMNLP | 4 |
| 2024 | Comparing Gender Bias in Lexical Semantics and World Knowledge: Deep-learning Models Pre-trained on Historical Corpus
Yingqiu Ge, Jinghang Gu, Chu-Ren Huang, Lifu Li |
PACLIC | 3 |
| 2024 | A Comparable Corpus-Driven Study on Dative Variation in Mandarin Chinese and the Pedagogical Implications
Menghan Jiang, Chu-Ren Huang |
PACLIC | 2 |
| 2024 | The Evolving Use of WAR Metaphors in Businesswomen-focused Media Discourse
Yanlin Li 0013, Jing Chen 0048, Kathleen Ahrens, Chu-Ren Huang |
PACLIC | 4 |
| 2024 | Analyzing the Gendered Power Dynamics in Addressing Practices: A Corpus-based Approach
Chu-Ren Huang |
PACLIC | 2 |
| 2024 | Word Boundary Decision: An Efficient Approach for Low-Resource Word Segmentation
Chu-Ren Huang |
PACLIC | 2 |
| 2024 | Comparing Professional and Common Literary Critics Using Multi-Dimensional Analysis
Chu-Ren Huang |
PACLIC | 2 |
| 2024 | Supervised Cross-Momentum Contrast: Aligning representations with prototypical examples to enhance financial sentiment analysisabstractFinancial sentiment analysis plays a pivotal role in understanding market dynamics and investor sentiment. In this paper, we propose the Supervised Cross-Momentum Contrast (SuCroMoCo) framework, a novel approach for financial sentiment analysis. SuCroMoCo leverages supervised contrastive learning and cross-momentum contrast to align financial text representations with prototypical representations based on sentiment categories. This alignment greatly improves classification performance, addressing the limitations of pre-trained language models (PLMs) in fully grasping the intricate nature of financial text. Through extensive experiments, we demonstrate that SuCroMoCo outperforms existing PLMs-based approaches and Large Language Models (LLMs) on diverse benchmark datasets. The code and datasets can be found in https://github.com/PengBO-O/SuCroMoCo. Emmanuele Chersoni, Yu-Yin Hsu, Le Qiu, Chu-Ren Huang |
Knowl. Based Syst. | 5 |
| 2024 | Perceptional and actional enrichment for metaphor detection with sensorimotor normsabstractAbstract Understanding the nature of meaning and its extensions (with metaphor as one typical kind) has been one core issue in figurative language study since Aristotle’s time. This research takes a computational cognitive perspective to model metaphor based on the assumption that meaning is perceptual, embodied, and encyclopedic. We model word meaning representation for metaphor detection with embodiment information obtained from behavioral experiments. Our work is the first attempt to incorporate sensorimotor knowledge into neural networks for metaphor detection, and demonstrates superiority, consistency, and interpretability compared to peer systems based on two general datasets. In addition, with cross-sectional analysis of different feature schemas, our results suggest that metaphor, as a device of cognitive conceptualization, can be ‘learned’ from the perceptual and actional information independent of several more explicit levels of linguistic representation. The access to such knowledge allows us to probe further into word meaning mapping tendencies relevant to our conceptualization and reaction to the physical world. Mingyu Wan, Qi Su 0001, Kathleen Ahrens, Chu-Ren Huang |
Nat. Lang. Eng. | 4 |
| 2024 | Fake News, Real Emotions: Emotion Analysis of COVID-19 Infodemic in WeiboabstractThe proliferation of COVID-19 fake news on social media poses a severe threat to the health information ecosystem. We show that affective computing can make significant contributions to combat this infodemic. Given that fake news is often presented with emotional appeals, we propose a new perspective on the role of emotion in the attitudes, perceptions, and behaviors of the dissemination of information. We study emotions in conjunction with fake news, and explore different aspects of their interaction. To process both emotion and ‘falsehood’ based on the same set of data, we auto-tag emotions on existing COVID-19 fake news datasets following an established emotion taxonomy. More specifically, based on the distribution of seven basic emotions (e.g.Happiness, Like, Fear, Sadness, Surprise, Disgust, Anger), we find across domains and styles that COVID-19 fake news is dominated by emotions ofFear(e.g., of coronavirus), andDisgust(e.g., of social conflicts). In addition, the framing of fake news in terms of gain-versus-loss reveals a close correlation between emotions, perceptions, and collective human reactions. Our analysis confirms the significant role of emotionFearin the spreading of the fake news, especially when contextualized in the loss frame. Our study points to a future direction of incorporating emotion footprints in models of automatic fake news detection, and establishes an affective computing approach to information quality in general and fake news detection in particular. Mingyu Wan, Sophia Yat Mei Lee, Chu-Ren Huang |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Existence Justifies Reason: A Data Analysis on Chinese Classifiers Based on Eye Tracking and Transformers
Emmanuele Chersoni, Chu-Ren Huang |
PACLIC | 3 |
| 2023 | Tracing Social Change through Metaphor: A Diachronic Corpus-Assisted Analysis
Huiheng Zeng, Kathleen Ahrens, Chu-Ren Huang |
PACLIC | 3 |
| 2022 | From Frying to Speculating: Google Ngram evidence to the meaning development of '?' in Mandarin Chinese
Jing Chen 0048, Chu-Ren Huang |
PACLIC | 2 |
| 2022 | Cross-strait Variations on Two Near-synonymous Loanwords xie2shang1 and tan2pan4: A Corpus-based Comparative Study
Yueyue Huang, Chu-Ren Huang |
PACLIC | 2 |
| 2022 | Gain-framed Buying or Loss-framed Selling? The Analysis of Near Synonyms in Mandarin in Prospect Theory
Chu-Ren Huang |
PACLIC | 2 |
| 2022 | Multi-probe attention neural network for COVID-19 semantic indexingabstractBACKGROUND: The COVID-19 pandemic has increasingly accelerated the publication pace of scientific literature. How to efficiently curate and index this large amount of biomedical literature under the current crisis is of great importance. Previous literature indexing is mainly performed by human experts using Medical Subject Headings (MeSH), which is labor-intensive and time-consuming. Therefore, to alleviate the expensive time consumption and monetary cost, there is an urgent need for automatic semantic indexing technologies for the emerging COVID-19 domain. RESULTS: In this research, to investigate the semantic indexing problem for COVID-19, we first construct the new COVID-19 Semantic Indexing dataset, which consists of more than 80 thousand biomedical articles. We then propose a novel semantic indexing framework based on the multi-probe attention neural network (MPANN) to address the COVID-19 semantic indexing problem. Specifically, we employ a k-nearest neighbour based MeSH masking approach to generate candidate topic terms for each input article. We encode and feed the selected candidate terms as well as other contextual information as probes into the downstream attention-based neural network. Each semantic probe carries specific aspects of biomedical knowledge and provides informatively discriminative features for the input article. After extracting the semantic features at both term-level and document-level through the attention-based neural network, MPANN adopts a linear multi-view classifier to conduct the final topic prediction for COVID-19 semantic indexing. CONCLUSION: The experimental results suggest that MPANN promises to represent the semantic features of biomedical texts and is effective in predicting semantic topics for COVID-19 related biomedical articles. Jinghang Gu, Rong Xiang, Jing Li 0049, Wenjie Li 0002, Longhua Qian, Guodong Zhou 0001, Chu-Ren Huang |
BMC Bioinform. | 8 |
| 2021 | Aspect or Manner? A Study of Reduplicated Adverbials in Mandarin Chinese
Siaw-Fong Chung, Chu-Ren Huang |
PACLIC | 2 |
| 2021 | Language change in Chinese political discourse based on the relationship between sentence and clause
Renkui Hou, Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 2 |
| 2021 | Spatial-temporal attributes in verbal semantics: A corpus-based lexical semantic study of discriminating Mandarin near synonyms of "tui1" and "la1"
Qiangmei Liang, Chu-Ren Huang |
PACLIC | 2 |
| 2021 | Animosity and suffering: Metaphors of BITTERNESS in English and Chinese
Gábor Parti, Andreas Liesenfeld, Chu-Ren Huang |
PACLIC | 3 |
| 2021 | A Corpus-based Lexical Semantic Study of Mandarin Verbs of "Tui" and "La"
Chu-Ren Huang |
PACLIC | 2 |
| 2021 | From Near-synonyms to Divergent Viewpoint Foci: A Corpus-based MARVS Driven Account of Two Verbs of Attention
Chu-Ren Huang |
PACLIC | 2 |
| 2021 | Automatic Analysis of Linguistic Features in Journal Articles of Different Academic Impacts with Feature Engineering Techniques
Ruiying Yang, Siyu Lei, Chu-Ren Huang |
PACLIC | 3 |
| 2021 | Scikit-talk: A toolkit for processing real-world conversational speech dataabstractWe present Scikit-talk, an open-source toolkit for processing collections of real-world conversational speech in Python.First of its kind, the toolkit equips those interested in studying or modeling conversations with an easyto-use interface to build and explore large collections of transcriptions and annotations of talk-in-interaction.Designed for applications in speech processing and Conversational AI, Scikit-talk provides tools to custombuild datasets for tasks such as intent prototyping, dialog flow testing, and conversation design.Its preprocessor module comes with several pre-built interfaces for common transcription formats, which aim to make working across multiple data sources more accessible.The explorer module provides a collection of tools to explore and analyse this data type via string matching and unsupervised machine learning techniques.Scikit-talk serves as a platform to collect and connect different transcription formats and representations of talk, enabling the user to quickly build multilingual datasets of varying detail and granularity.Thus, the toolkit aims to make working with authentic conversational speech data in Python more accessible and to provide the user with comprehensive options to work with representations of talk in appropriate detail for any downstream task.For the latest updates and information on currently supported languages and language resources, please refer to Andreas Liesenfeld, Gábor Parti, Chu-Ren Huang |
SIGDIAL | 3 |
| 2021 | Decoding Word Embeddings with Brain-Based Semantic FeaturesabstractWord embeddings are vectorial semantic representations built with either counting or predicting techniques aimed at capturing shades of meaning from word co-occurrences. Since their introduction, these representations have been criticized for lacking interpretable dimensions. This property of word embeddings limits our understanding of the semantic features they actually encode. Moreover, it contributes to the “black box” nature of the tasks in which they are used, since the reasons for word embedding performance often remain opaque to humans. In this contribution, we explore the semantic properties encoded in word embeddings by mapping them onto interpretable vectors, consisting of explicit and neurobiologically motivated semantic features (Binder et al. 2016). Our exploration takes into account different types of embeddings, including factorized count vectors and predict models (Skip-Gram, GloVe, etc.), as well as the most recent contextualized representations (i.e., ELMo and BERT). In our analysis, we first evaluate the quality of the mapping in a retrieval task, then we shed light on the semantic features that are better encoded in each embedding type. A large number of probing tasks is finally set to assess how the original and the mapped embeddings perform in discriminating semantic categories. For each probing task, we identify the most relevant semantic features and we show that there is a correlation between the embedding performance and how they encode those features. This study sets itself as a step forward in understanding which aspects of meaning are captured by vector spaces, by proposing a new and simple method to carve human-interpretable semantic representations from distributional vectors. Emmanuele Chersoni, Enrico Santus, Chu-Ren Huang, Alessandro Lenci |
Comput. Linguistics | 3 |
| 2021 | Lexical data augmentation for sentiment analysisabstractAbstract Machine learning methods, especially deep learning models, have achieved impressive performance in various natural language processing tasks including sentiment analysis. However, deep learning models are more demanding for training data. Data augmentation techniques are widely used to generate new instances based on modifications to existing data or relying on external knowledge bases to address annotated data scarcity, which hinders the full potential of machine learning techniques. This paper presents our work using part‐of‐speech (POS) focused lexical substitution for data augmentation (PLSDA) to enhance the performance of machine learning algorithms in sentiment analysis. We exploit POS information to identify words to be replaced and investigate different augmentation strategies to find semantically related substitutions when generating new instances. The choice of POS tags as well as a variety of strategies such as semantic‐based substitution methods and sampling methods are discussed in detail. Performance evaluation focuses on the comparison between PLSDA and two previous lexical substitution‐based data augmentation methods, one of which is thesaurus‐based, and the other is lexicon manipulation based. Our approach is tested on five English sentiment analysis benchmarks: SST‐2, MR, IMDB, Twitter, and AirRecord. Hyperparameters such as the candidate similarity threshold and number of newly generated instances are optimized. Results show that six classifiers (SVM, LSTM, BiLSTM‐AT, bidirectional encoder representations from transformers [BERT], XLNet, and RoBERTa) trained with PLSDA achieve accuracy improvement of more than 0.6% comparing to two previous lexical substitution methods averaged on five benchmarks. Introducing POS constraint and well‐designed augmentation strategies can improve the reliability of lexical data augmentation methods. Consequently, PLSDA significantly improves the performance of sentiment analysis algorithms. Rong Xiang, Emmanuele Chersoni, Qin Lu 0001, Chu-Ren Huang, Wenjie Li 0002 |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2021 | Affective awareness in neural sentiment analysis
Rong Xiang, Jing Li 0049, Mingyu Wan, Jinghang Gu, Qin Lu 0001, Wenjie Li 0002, Chu-Ren Huang |
Knowl. Based Syst. | 7 |
| 2021 | Improving Attention Model Based on Cognition Grounded Data for Sentiment AnalysisabstractAttention models are proposed in sentiment analysis and other classification tasks because some words are more important than others to train the attention models. However, most existing methods either use local context based information, affective lexicons, or user preference information. In this work, we propose a novel attention model trained by cognition grounded eye-tracking data. First,a reading prediction model is built using eye-tracking data as dependent data and other features in the context as independent data. The predicted reading time is then used to build a cognition grounded attention layer for neural sentiment analysis. Our model can capture attentions in context both in terms of words at sentence level as well as sentences at document level. Other attention mechanisms can also be incorporated together to capture other aspects of attentions, such as local attention, and affective lexicons. Results of our work include two parts. The first part compares our proposed cognition ground attention model with other state-of-the-art sentiment analysis models. The second part compares our model with an attention model based on other lexicon based sentiment resources. Evaluations show that sentiment analysis using cognition grounded attention model outperforms the state-of-the-art sentiment analysis methods significantly. Comparisons to affective lexicons also indicate that using cognition grounded eye-tracking data has advantages over other sentiment resources by considering both word information and context information. This work brings insight to how cognition grounded data can be integrated into natural language processing (NLP) tasks. Rong Xiang, Qin Lu 0001, Chu-Ren Huang, Minglei Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | Are Word Embeddings Really a Bad Fit for the Estimation of Thematic Fit?abstractWhile neural embeddings represent a popular choice for word representation in a wide variety of NLP tasks, their usage for thematic fit modeling has been limited, as they have been reported to lag behind syntax-based count models. In this paper, we propose a complete evaluation of count models and word embeddings on thematic fit estimation, by taking into account a larger number of parameters and verb roles and introducing also dependency-based embeddings in the comparison. Our results show a complex scenario, where a determinant factor for the performance seems to be the availability to the model of reliable syntactic information for building the distributional representations of the roles. Emmanuele Chersoni, Ludovica Pannitto, Enrico Santus, Alessandro Lenci, Chu-Ren Huang |
LREC | 5 |
| 2020 | Ciron: a New Benchmark Dataset for Chinese Irony DetectionabstractAutomatic Chinese irony detection is a challenging task, and it has a strong impact on linguistic research. However, Chinese irony detection often lacks labeled benchmark datasets. In this paper, we introduce Ciron, the first Chinese benchmark dataset available for irony detection for machine learning models. Ciron includes more than 8.7K posts, collected from Weibo, a micro blogging platform. Most importantly, Ciron is collected with no pre-conditions to ensure a much wider coverage. Evaluation on seven different machine learning classifiers proves the usefulness of Ciron as an important resource for Chinese irony detection. Rong Xiang, Emmanuele Chersoni, Qin Lu 0001, Chu-Ren Huang |
LREC | 7 |
| 2020 | Affection Driven Neural Networks for Sentiment AnalysisabstractDeep neural network models have played a critical role in sentiment analysis with promising results in the recent decade. One of the essential challenges, however, is how external sentiment knowledge can be effectively utilized. In this work, we propose a novel affection-driven approach to incorporating affective knowledge into neural network models. The affective knowledge is obtained in the form of a lexicon under the Affect Control Theory (ACT), which is represented by vectors of three-dimensional attributes in Evaluation, Potency, and Activity (EPA). The EPA vectors are mapped to an affective influence value and then integrated into Long Short-term Memory (LSTM) models to highlight affective terms. Experimental results show a consistent improvement of our approach over conventional LSTM models by 1.0% to 1.5% in accuracy on three large benchmark datasets. Evaluations across a variety of algorithms have also proven the effectiveness of leveraging affective terms for deep model enhancement. Rong Xiang, Mingyu Wan, Jinghang Gu, Qin Lu 0001, Chu-Ren Huang |
LREC | 6 |
| 2020 | Sketching the English Translations of Kumāraj\=\iva's The Diamond Sutra: A Comparison of Individual Translators and Translation Teams
Xi Chen 0111, Vincent Xian Wang, Chu-Ren Huang |
PACLIC | 3 |
| 2020 | Language change in Report on the Work of the Government by Premiers of the People's Republic of China
Renkui Hou, Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 2 |
| 2020 | Marking Trustworthiness with Near Synonyms: A Corpus-based Study of "Renwei" and "Yiwei" in Chinese
Chu-Ren Huang |
PACLIC | 2 |
| 2020 | Predicting gender and age categories in English conversations using lexical, non-lexical, and turn-taking features
Andreas Liesenfeld, Gábor Parti, Yu-Yin Hsu, Chu-Ren Huang |
PACLIC | 4 |
| 2020 | Abstract Meaning Representation for MWE: A study of the mapping of aspectuality based on Mandarin light verb jiayi
Nianwen Xue, Chu-Ren Huang |
PACLIC | 3 |
| 2020 | Sensorimotor Enhanced Neural Network for Metaphor Detection
Mingyu Wan, Baixi Xing, Qi Su 0001, Pengyuan Liu 0001, Chu-Ren Huang |
PACLIC | 5 |
| 2020 | A Parallel Corpus-driven Approach to Bilingual Oenology Term Banks: How Culture Differences Influence Wine Tasting Terms
Vincent Xian Wang, Xi Chen 0111, Songnan Quan, Chu-Ren Huang |
PACLIC | 4 |
| 2020 | Corpus-based Comparison of Verbs of Separation "Qie" and "Ge"
Nga-In Wu, Chu-Ren Huang, Lap-Kei Lee |
PACLIC | 2 |
| 2020 | Dual memory network model for sentiment analysis of review text
Jiaxing Shen, Mingyu Derek Ma, Rong Xiang, Qin Lu 0001, Elvira Perez, Chu-Ren Huang |
Knowl. Based Syst. | 7 |
| 2020 | Robust stylometric analysis and author attribution based on tones and rimesabstractAbstract In this article, we propose an innovative and robust approach to stylometric analysis without annotation and leveraging lexical and sub-lexical information. In particular, we propose to leverage the phonological information of tones and rimes in Mandarin Chinese automatically extracted from unannotated texts. The texts from different authors were represented by tones, tone motifs, and word length motifs as well as rimes and rime motifs. Support vector machines and random forests were used to establish the text classification model for authorship attribution. From the results of the experiments, we conclude that the combination of bigrams of rimes, word-final rimes, and segment-final rimes can discriminate the texts from different authors effectively when using random forests to establish the classification model. This robust approach can in principle be applied to other languages with established phonological inventory of onset and rimes. Renkui Hou, Chu-Ren Huang |
Nat. Lang. Eng. | 2 |
| 2020 | Classification of regional and genre varieties of Chinese: A correspondence analysis approach based on comparable balanced corporaabstractAbstract This paper proposes a robust text classification and correspondence analysis approach to identification of similar languages. In particular, we propose to use the readily available information of clauses and word length distribution to model similar languages. The modeling and classification are based on the hypothesis that languages are self-adaptive complex systems and hence can be classified by dynamic features describing the system, especially in terms of distributional relations of constituents of a system. For similar languages whose grammatical differences are often subtle, classification based on dynamic system features should be more effective. To test this hypothesis, we considered both regional and genre varieties of Mandarin Chinese for classification. The data are extracted from two comparable balanced corpora to minimize possible confounding factors. The two corpora are the Sinica Corpus from Taiwan and the Lancaster Corpus of Mandarin Chinese from Mainland China, and the two genres are reportage and review. Our text classification and correspondence analysis results show that the linguistically felicitous two-level constituency model combining power functions between word and clauses effectively classifies the two varieties of Chinese for both genres. In addition, we found that genres do have compounding effect on classification of regional varieties. In particular, reportage in two varieties is more likely to be classified than review, corroborating the complex system view of language variations. That is, language variations and changes typically do not take place evenly across the board for the complete language system. This further enhances our hypothesis that dynamic complex system features, such as the power functions captured by the Menzerath–Altmann law, provide effective models in classifications of similar languages. Renkui Hou, Chu-Ren Huang |
Nat. Lang. Eng. | 2 |
| 2019 | Neighborhood in Decay: Working Memory Modulates Effect of Phonological Similarity on Lexical Access
Karl David Neergaard, James Britton, Chu-Ren Huang |
CogSci | 3 |
| 2019 | A structured distributional model of sentence meaning and processingabstractAbstract Most compositional distributional semantic models represent sentence meaning with a single vector. In this paper, we propose a structured distributional model (SDM) that combines word embeddings with formal semantics and is based on the assumption that sentences represent events and situations. The semantic representation of a sentence is a formal structure derived from discourse representation theory and containing distributional vectors. This structure is dynamically and incrementally built by integrating knowledge about events and their typical participants, as they are activated by lexical items. Event knowledge is modelled as a graph extracted from parsed corpora and encoding roles and relationships between participants that are represented as distributional vectors. SDM is grounded on extensive psycholinguistic research showing that generalized knowledge about events stored in semantic memory plays a key role in sentence comprehension.We evaluate SDMon two recently introduced compositionality data sets, and our results show that combining a simple compositionalmodel with event knowledge constantly improves performances, even with dif ferent types of word embeddings. Emmanuele Chersoni, Enrico Santus, Ludovica Pannitto, Alessandro Lenci, Philippe Blache, Chu-Ren Huang |
Nat. Lang. Eng. | 6 |
| 2018 | Annotating Chinese Light Verb Constructions according to PARSEME guidelines
Menghan Jiang, Natalia Klyueva, Hongzhi Xu, Chu-Ren Huang |
LREC | 4 |
| 2018 | Facilitating and Blocking Conditions of Haplology: A comparative study of Hong Kong Cantonese and Taiwan Mandarin
Sam Yin Wong, I-Hsuan Chen, Chu-Ren Huang |
PACLIC | 3 |
| 2018 | Semantic Transparency of Radicals in Chinese Characters: An Ontological Perspective
Yike Yang, Chu-Ren Huang, Sicong Dong |
PACLIC | 2 |
| 2017 | Leveraging Eventive Information for Better Metaphor Detection and Classificationabstract202101 bcrc I-Hsuan Chen, Qin Lu 0001, Chu-Ren Huang |
CoNLL | 4 |
| 2017 | A Cognition Based Attention Model for Sentiment AnalysisabstractAttention models are proposed in sentiment analysis because some words are more important than others.However, most existing methods either use local context based text information or user preference information.In this work, we propose a novel attention model trained by cognition grounded eye-tracking data.A reading prediction model is first built using eye-tracking data as dependent data and other features in the context as independent data.The predicted reading time is then used to build a cognition based attention (CBA) layer for neural sentiment analysis.As a comprehensive model, We can capture attentions of words in sentences as well as sentences in documents.Different attention mechanisms can also be incorporated to capture other aspects of attentions.Evaluations show the CBA based method outperforms the state-of-the-art local context based attention methods significantly.This brings insight to how cognition grounded data can be brought into NLP tasks. Qin Lu 0001, Rong Xiang, Minglei Li 0001, Chu-Ren Huang |
EMNLP | 5 |
| 2017 | A Preliminary Phonetic Investigation of Alphabetic Words in Mandarin Chineseabstract202303 bcww Hongchao Liu, Chu-Ren Huang |
INTERSPEECH | 4 |
| 2017 | Stylometric Studies based on Tone and Word Length Motifs
Renkui Hou, Chu-Ren Huang |
PACLIC | 2 |
| 2017 | Lexicalization, Separation and transitivity: A comparative study of Mandarin VO compound Variations
Menghan Jiang, Chu-Ren Huang |
PACLIC | 2 |
| 2017 | Multi-dimensional Meanings of Subjective Adverbs - Case Study of Mandarin Chinese Adverb Pianpian
Chu-Ren Huang |
PACLIC | 3 |
| 2016 | Unsupervised Measure of Word Similarity: How to Outperform Co-Occurrence and Vector Cosine in VSMsabstractIn this paper, we claim that vector cosine – which is generally considered among the most efficient unsupervised measures for identifying word similarity in Vector Space Models – can be outperformed by an unsupervised measure that calculates the extent of the intersection among the most mutually dependent contexts of the target words. To prove it, we describe and evaluate APSyn, a variant of the Average Precision that, without any optimization, outperforms the vector cosine and the co-occurrence on the standard ESL test set, with an improvement ranging between +9.00% and +17.98%, depending on the number of chosen top contexts. Enrico Santus, Alessandro Lenci, Tin-Shing Chiu, Qin Lu 0001, Chu-Ren Huang |
AAAI | 5 |
| 2016 | ROOT13: Spotting Hypernyms, Co-Hyponyms and RandomsabstractIn this paper, we describe ROOT13, a supervised system for the classification of hypernyms, co-hyponyms and random words. The system relies on a Random Forest algorithm and 13 unsupervised corpus-based features. We evaluate it with a 10-fold cross validation on 9,600 pairs, equally distributed among the three classes and involving several Parts-Of-Speech (i.e. adjectives, nouns and verbs). When all the classes are present, ROOT13 achieves an F1 score of 88.3%, against a baseline of 57.6% (vector cosine). When the classification is binary, ROOT13 achieves the following results: hypernyms-co-hyponyms (93.4% vs. 60.2%), hypernyms-random (92.3% vs. 65.5%) and co-hyponyms-random (97.3% vs. 81.5%). Our results are competitive with state-of-the-art models. Enrico Santus, Alessandro Lenci, Tin-Shing Chiu, Qin Lu 0001, Chu-Ren Huang |
AAAI | 5 |
| 2016 | Domain-specific user preference prediction based on multiple user activitiesabstractInferring latent user preferences using both structured and unstructured data is an important social computing task. In this paper, we propose a user preference representation based on user activities embedded in unstructured data to better encode the homophily theory. The representation of an individual user is learned using a embedding based method to integrate latent user preferences in social media. The method has the ability to integrate a variety of user activities based cues from user comments, user social network (i.e; follower/followee connections) and user interested topics which are indicated by the topics a user has participated in. Experiments are conducted to evaluate the prediction of each user's favorite team as a part of user preferences in a dataset collected from the Hu-pu basketball discussion forum.1Results clearly indicate that our proposed user representation outperforms other user representation baselines. Integrating user social network and user interested topics with user comments can improve the overall performance of user preference prediction. Qin Lu 0001, Minglei Li 0001, Chu-Ren Huang |
IEEE BigData | 5 |
| 2016 | Representing Verbs with Rich Contexts: an Evaluation on Verb SimilarityabstractSeveral studies on sentence processing suggest that the mental lexicon keeps track of the mutual expectations between words.Current DSMs, however, represent context words as separate features, thereby loosing important information for word expectations, such as word interrelations.In this paper, we present a DSM that addresses this issue by defining verb contexts as joint syntactic dependencies.We test our representation in a verb similarity task on two datasets, showing that joint contexts achieve performances comparable to single dependencies or even better.Moreover, they are able to overcome the data sparsity problem of joint feature spaces, in spite of the limited size of our training corpus. Emmanuele Chersoni, Enrico Santus, Alessandro Lenci, Philippe Blache, Chu-Ren Huang |
EMNLP | 5 |
| 2016 | A lexicon of perception for the identification of synaesthetic metaphors in corpora
Francesca Strik Lievers, Chu-Ren Huang |
LREC | 2 |
| 2016 | EVALution-MAN: A Chinese Dataset for the Training and Evaluation of DSMs
Hongchao Liu, Karl David Neergaard, Enrico Santus, Chu-Ren Huang |
LREC | 4 |
| 2016 | Database of Mandarin Neighborhood Statistics
Karl David Neergaard, Hongzhi Xu, Chu-Ren Huang |
LREC | 3 |
| 2016 | Nine Features in a Random Forest to Learn Taxonomical Semantic Relations
Enrico Santus, Alessandro Lenci, Tin-Shing Chiu, Qin Lu 0001, Chu-Ren Huang |
LREC | 5 |
| 2016 | What a Nerd! Beating Students and Vector Cosine in the ESL and TOEFL Datasets
Enrico Santus, Alessandro Lenci, Tin-Shing Chiu, Qin Lu 0001, Chu-Ren Huang |
LREC | 5 |
| 2016 | The use of body part terms in Taiwan and China: Analyzing 血 xue 'blood' and 骨 gu 'bone' in Chinese Gigaword v. 2.0
Ren-Feng Duann, Chu-Ren Huang |
PACLIC | 2 |
| 2016 | Endurant vs Perdurant: Ontological Motivation for Language Variations
Chu-Ren Huang |
PACLIC | 1 |
| 2016 | Transitivity in Light Verb Variations in Mandarin Chinese - A Comparable Corpus-based Statistical Approach
Menghan Jiang, Dingxu Shi, Chu-Ren Huang |
PACLIC | 3 |
| 2016 | Testing APSyn against Vector Cosine on Similarity Estimation
Enrico Santus, Emmanuele Chersoni, Alessandro Lenci, Chu-Ren Huang, Philippe Blache |
PACLIC | 4 |
| 2016 | The Synaesthetic and Metaphorical Uses of 味 wei 'taste' in Chinese Buddhist Suttas
Jiajuan Xiong, Chu-Ren Huang |
PACLIC | 2 |
| 2015 | The Invertible Construction in Chinese
Yan Cong, Chu-Ren Huang, Lian-Hee Wee |
PACLIC | 2 |
| 2015 | When Embodiment Meets Generative Lexicon: The Human Body Part Metaphors in Sinica Corpus
Ren-Feng Duann, Chu-Ren Huang |
PACLIC | 2 |
| 2015 | Graph Theoretic Features of the Adult Mental lexicon Predict Language Production in Mandarin: Clustering Coefficient
Karl David Neergaard, Chu-Ren Huang |
PACLIC | 2 |
| 2015 | Sentiment Analyzer with Rich Features for Ironic and Sarcastic Tweets
Piyoros Tungthamthiti, Enrico Santus, Hongzhi Xu, Chu-Ren Huang, Kiyoaki Shirai |
PACLIC | 4 |
| 2015 | Mechanical Turk-based Experiment vs Laboratory-based Experiment: A Case Study on the Comparison of Semantic Transparency Rating Data
Shichang Wang, Chu-Ren Huang, Angel Chan |
PACLIC | 2 |
| 2015 | De-verbalization and Nominal Categories in Mandarin Chinese: A corpus-driven study in both Mainland Mandarin and Taiwan Mandarin
Jiajuan Xiong, Chu-Ren Huang |
PACLIC | 2 |
| 2015 | Auditory Synaesthesia and Near Synonyms: A Corpus-Based Analysis of sheng1 and yin1 in Mandarin Chinese
Chu-Ren Huang, Hongzhi Xu |
PACLIC | 2 |
| 2014 | Annotating Events in an Emotion Corpus
Sophia Yat Mei Lee, Shoushan Li, Chu-Ren Huang |
LREC | 3 |
| 2014 | Taking Antonymy Mask off in Vector Space
Enrico Santus, Qin Lu 0001, Alessandro Lenci, Chu-Ren Huang |
PACLIC | 4 |
| 2014 | On the Argument Structures of the Transitive Verb 'annoy; be annoyed; bother to do': A Study Based on Two Comparable Corpora
Jiajuan Xiong, Chu-Ren Huang |
PACLIC | 2 |
| 2014 | An Analysis of Radicals-based Features in Subjectivity Classification on Simplified Chinese Sentences
Chu-Ren Huang |
PACLIC | 2 |
| 2013 | Joint learning on sentiment and emotion classificationabstractSentiment and emotion classification have been popularly but separately studied in natural language processing. In this paper, we address joint learning on sentiment and emotion classification where both the labeled data for sentiment and emotion classification are available. The objective of this joint-learning is to benefit the two tasks from each other for improving their performances. Specifically, an extra data set that is annotated with both sentiment and emotion labels are employed to estimate the transformation probability between the two kinds of labels. Furthermore, the transformation probability is leveraged to transfer the classification labels to benefit the two tasks from each other. Empirical studies demonstrate the effectiveness of our approach for the novel joint learning task. Shoushan Li, Sophia Yat Mei Lee, Guodong Zhou 0001, Chu-Ren Huang |
CIKM | 5 |
| 2013 | A Rule System for Chinese Time Entity Recognition by Comprehensive Linguistic Study
Hongzhi Xu, Chu-Ren Huang |
IJCNLP | 2 |
| 2013 | Semi-supervised Text Categorization by Considering Sufficiency and Diversity
Shoushan Li, Sophia Yat Mei Lee, Chu-Ren Huang |
NLPCC | 4 |
| 2013 | Active Learning for Cross-Lingual Sentiment Classification
Shoushan Li, Chu-Ren Huang |
NLPCC | 4 |
| 2013 | Detecting Emotion Causes with a Linguistic Rule-Based ApproachabstractMost theories of emotion treat recognition of a triggering cause event as an integral part of emotion processing. This paper proposes emotion cause detection as a new research area in emotion processing. As a first step toward fully automatic inference of emotion‐cause correlation, we propose a text‐driven, rule‐based approach to emotion cause detection in Chinese. First, we constructed a Chinese emotion cause annotated corpus based on our proposed annotation scheme. Next, we analyzed the corpus data, which yielded the identification of seven groups of linguistic cues and two sets of generalized linguistic rules for the detection of emotion causes. We then developed a rule‐based system for emotion cause detection based on the linguistic rules. In addition, we proposed an evaluation scheme with two phases for performance assessment. The results of our experiments show that our system achieved a promising performance for cause occurrence detection, as well as for cause event detection. The current study should lay the groundwork for future research on the inferences of implicit information and the discovery of new information based on cause‐event relation. Sophia Yat Mei Lee, Ying Chen 0012, Chu-Ren Huang, Shoushan Li |
Comput. Intell. | 3 |
| 2012 | A Grammar-informed Corpus-based Sentence Database for Linguistic and Computational Studies
Hongzhi Xu, Helen Kai-Yun Chen, Chu-Ren Huang, Qin Lu 0001, Dingxu Shi, Tin-Shing Chiu |
LREC | 3 |
| 2012 | The Headedness of Mandarin Chinese Serial Verb Constructions: A Corpus-Based Study
Jingxia Lin, Chu-Ren Huang, Huarui Zhang, Hongzhi Xu |
PACLIC | 2 |
| 2012 | Type Construction of Event Nouns in Mandarin Chinese
Shan Wang 0002, Chu-Ren Huang |
PACLIC | 2 |
| 2012 | Compositionality of NN Compounds: A Case Study on [N1+Artifactual-Type Event Nouns]
Shan Wang 0002, Chu-Ren Huang, Hongzhi Xu |
PACLIC | 2 |
| 2012 | A robust web personal name information extraction system
Ying Chen 0012, Sophia Yat Mei Lee, Chu-Ren Huang |
Expert Syst. Appl. | 3 |
| 2011 | The Co-occurrence of Two Delimiters: An Investigation of Mandarin Chinese Resultatives
Jingxia Lin, Chu-Ren Huang |
PACLIC | 2 |
| 2011 | Compound Event Nouns of the 'Modifier-head' Type in Mandarin Chinese
Shan Wang 0002, Chu-Ren Huang |
PACLIC | 2 |
| 2011 | Multi-Domain Sentiment Classification with Classifier Combination
Shoushan Li, Chu-Ren Huang, Chengqing Zong |
J. Comput. Sci. Technol. | 2 |
| 2010 | Employing Personal/Impersonal Views in Supervised and Semi-Supervised Sentiment Classification
Shoushan Li, Chu-Ren Huang, Guodong Zhou 0001, Sophia Yat Mei Lee |
ACL | 2 |
| 2010 | Emotion Cause Detection with Linguistic Constructions
Ying Chen 0012, Sophia Yat Mei Lee, Shoushan Li, Chu-Ren Huang |
COLING | 4 |
| 2010 | Sentiment Classification and Polarity Shifting
Shoushan Li, Sophia Yat Mei Lee, Ying Chen 0012, Chu-Ren Huang, Guodong Zhou 0001 |
COLING | 4 |
| 2010 | Emotion Cause Events: Corpus Construction and Analysis
Sophia Yat Mei Lee, Ying Chen 0012, Shoushan Li, Chu-Ren Huang |
LREC | 4 |
| 2010 | Automatic Acquisition of Chinese Novel Noun Compounds
Chu-Ren Huang, Shiwen Yu |
LREC | 2 |
| 2010 | Using Corpus-based Linguistic Approaches in Sense Prediction Study
Jia-Fei Hong, Sue-jin Ker, Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 3 |
| 2010 | Cross-sortal Predication and Polysemy
Petr Simon, Chu-Ren Huang |
PACLIC | 2 |
| 2010 | Incorporate Credibility into Context for the Best Social Media Answers
Qi Su 0001, Helen Kai-Yun Chen, Chu-Ren Huang |
PACLIC | 3 |
| 2010 | Adjectival Modification to Nouns in Mandarin Chinese: Case Studies on "cháng+noun" and "adjective+tú shu gu n"
Shan Wang 0002, Chu-Ren Huang |
PACLIC | 2 |
| 2010 | Compositional Operations of Mandarin Chinese Perception Verb "kàn": A Generative Lexicon Approach
Shan Wang 0002, Chu-Ren Huang |
PACLIC | 2 |
| 2009 | A Framework of Feature Selection Methods for Text Categorization
Shoushan Li, Chengqing Zong, Chu-Ren Huang |
ACL/IJCNLP | 4 |
| 2009 | An Integrated Approach to Heterogeneous Data for Information Extraction
Ying Chen 0012, Sophia Yat Mei Lee, Chu-Ren Huang |
PACLIC | 3 |
| 2009 | Are Emotions Enumerable or Decomposable? And its Implications for Emotion Processing
Ying Chen 0012, Sophia Yat Mei Lee, Chu-Ren Huang |
PACLIC | 3 |
| 2009 | Bridging the Gap between Graph Modeling and Developmental Psycholinguistics: An Experiment on Measuring Lexical Proximity in Chinese Semantic Space
Shu-Kai Hsieh, Chun-Han Chang, Ivy Kuo, Hintat Cheung, Chu-Ren Huang, Bruno Gaume |
PACLIC | 5 |
| 2009 | Cause Event Representations for Happiness and Surprise
Sophia Yat Mei Lee, Ying Chen 0012, Chu-Ren Huang |
PACLIC | 3 |
| 2009 | Chinese WordNet Domains: Bootstrapping Chinese WordNet with Semantic Domain Labels
Lung-Hao Lee, Yu-Ting Yu, Chu-Ren Huang |
PACLIC | 3 |
| 2009 | Sentiment Classification Considering Negation and Contrast Transition
Shoushan Li, Chu-Ren Huang |
PACLIC | 2 |
| 2009 | Word Boundary Decision with CRF for Chinese Word Segmentation
Shoushan Li, Chu-Ren Huang |
PACLIC | 2 |
| 2008 | Constructing Taxonomy of Numerative Classifiers for Asian Languages
Kiyoaki Shirai, Takenobu Tokunaga, Chu-Ren Huang, Shu-Kai Hsieh, Tzu-Yi Kuo, Virach Sornlertlamvanich, Thatsanee Charoenporn |
IJCNLP | 3 |
| 2008 | The Extended Architecture of Hantology for Japan Kanji
Ya-Min Chou, Chu-Ren Huang, Jia-Fei Hong |
LREC | 2 |
| 2008 | Extracting Concrete Senses of Lexicon through Measurement of Conceptual Similarity in Ontologies
Siaw-Fong Chung, Laurent Prévot 0001, Kathleen Ahrens, Shu-Kai Hsieh, Chu-Ren Huang |
LREC | 6 |
| 2008 | Quality Assurance of Automatic Annotation of Very Large Corpora: a Study based on heterogeneous Tagging System
Chu-Ren Huang, Lung-Hao Lee, Jia-Fei Hong, Weiguang Qu, Shiwen Yu |
LREC | 1 |
| 2008 | Adapting International Standard for Asian Language Technologies
Takenobu Tokunaga, Dain Kaplan, Chu-Ren Huang, Shu-Kai Hsieh, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Kiyoaki Shirai, Virach Sornlertlamvanich, Thatsanee Charoenporn, Yingju Xia |
LREC | 3 |
| 2008 | KYOTO: a System for Mining, Structuring and Distributing Knowledge across Languages and Cultures
Piek Vossen, Eneko Agirre, Nicoletta Calzolari, Christiane Fellbaum, Shu-Kai Hsieh, Chu-Ren Huang, Hitoshi Isahara, Kyoko Kanzaki, Andrea Marchetti, Monica Monachini, Federico Neri, Remo Raffaelli, German Rigau, Maurizio Tesconi, Joop VanGent |
LREC | 6 |
| 2008 | Contrastive Approach towards Text Source Classification based on Top-Bag-of-Word Similarity
Chu-Ren Huang, Lung-Hao Lee |
PACLIC | 1 |
| 2008 | An Ontology of Chinese Radicals: Concept Derivation and Knowledge Representation based on the Semantic Symbols of the Four Hoofed-Mammals
Chu-Ren Huang, Ya-Jun Yang, Sheng-Yi Chen |
PACLIC | 1 |
| 2007 | Automatic Discovery of Named Entity Variants: Grammar-driven Approaches to Non-Alphabetical Transliterations
Chu-Ren Huang, Petr Simon, Shu-Kai Hsieh |
ACL | 1 |
| 2007 | Rethinking Chinese Word Segmentation: Tokenization, Character Classification, or Wordbreak Identification
Chu-Ren Huang, Petr Simon, Shu-Kai Hsieh, Laurent Prévot 0001 |
ACL | 1 |
| 2007 | Computing Thresholds of Linguistic Saliency
Siaw-Fong Chung, Kathleen Ahrens, Chung-Ping Cheng, Chu-Ren Huang, Petr Simon |
PACLIC | 4 |
| 2007 | The Polysemy of Da3: An ontology-based lexical semantic study
Jia-Fei Hong, Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 2 |
| 2006 | When Conset Meets Synset: A Preliminary Survey of an Ontological Lexical Resource Based on Chinese Characters
Shu-Kai Hsieh, Chu-Ren Huang |
ACL | 2 |
| 2006 | Infrastructure for Standardization of Asian Language Resources
Takenobu Tokunaga, Virach Sornlertlamvanich, Thatsanee Charoenporn, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Chu-Ren Huang, Yingju Xia, Hao Yu 0005, Laurent Prévot 0001, Kiyoaki Shirai |
ACL | 7 |
| 2006 | GuangQunFangPu: e-Humanities Combining Textual and Botanic InformationabstractIn this paper, we propose a lexicon-driven and ontologymerging methodology of constructing diachronic domain knowledge via Sinica BOW, a bilingual ontological lexical resource based on WordNet and SUMO ontology. The main domain knowledge that we model our specialized ontology on is GuangQunFangPu, a Chinese classic literature of botany. Our studies yields promising result, we believe that the proposed research scenario will boost the on-going development of e-Humanities. Shu-Kai Hsieh, Shu-Ming Chang, Chun-Han Chang, Yi-Shuan Zhou, Chu-Ren Huang, Feng-Ju Lo, Ru-Yng Chang |
e-Science | 5 |
| 2006 | Hantology-A Linguistic Resource for Chinese Language Processing and Studying
Ya-Min Chou, Chu-Ren Huang |
LREC | 2 |
| 2006 | Uniform and Effective Tagging of a Heterogeneous Giga-word Corpus
Wei-Yun Ma, Chu-Ren Huang |
LREC | 2 |
| 2006 | Using Chinese Gigaword Corpus and Chinese Word Sketch in linguistic Research
Jia-Fei Hong, Chu-Ren Huang |
PACLIC | 2 |
| 2006 | Knowledge-Rich Approach to Automatic Grammatical Information Acquisition: Enriching Chinese Sketch Engine with a Lexical Grammar
Chu-Ren Huang, Wei-Yun Ma, Yi-Ching Wu, Chih-Ming Chiu |
PACLIC | 1 |
| 2006 | Using the Swadesh list for creating a simple common taxonomy
Laurent Prévot 0001, Chu-Ren Huang, I-Li Su |
PACLIC | 2 |
| 2005 | In and Out: Senses and Meaning Extension of Mandarin Spatial Terms nei and wai
Yi-Ching Wu, Cui-Xia Weng, Chu-Ren Huang |
PACLIC | 3 |
| 2004 | Sinica BOW (Bilingual Ontological Wordnet): Integration of Bilingual WordNet and SUMO
Chu-Ren Huang, Ru-Yng Chang, Hshiang-Pin Lee |
LREC | 1 |
| 2004 | Distributional Consistency: As a General Method for Defining a Core Lexicon
Huarui Zhang, Chu-Ren Huang, Shiwen Yu |
LREC | 2 |
| 2004 | Ontology-based Prediction of Compound Relations : A Study Based on SUMO
Jia-Fei Hong, Xiang-Bing Li, Chu-Ren Huang |
PACLIC | 3 |
| 2004 | Text-based Construction and Comparison of Domain Ontology : A Study Based on Classical Poetry
Chu-Ren Huang |
PACLIC | 1 |
| 2003 | The Semantics of Onomatopoeic Speech Act Verbs
I-Ni Tsai, Chu-Ren Huang |
PACLIC | 2 |
| 2003 | The Semantics of Shapes : A Study based on Mandarin Quan1zi5(圈子)
Cui-Xia Weng, Chu-Ren Huang |
PACLIC | 2 |
| 2002 | The Structure of Polysemy : A Study of Multi-sense Words Based on WordNet
Jen-Yi Lin, Chang-Hua Yang, Shu-Chuan Tseng, Chu-Ren Huang |
PACLIC | 4 |
| 2001 | A Comparative Study of English and Chinese Synonym Pairs : An Approach based on The Module-Attribute Representation of Verbal Semantics
Kathleen Ahrens, Chu-Ren Huang |
PACLIC | 2 |
| 2000 | The Module-Attribute Representation of Verbal Semantics
Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 1 |
| 1999 | Alternation Across Semantic Fields : A Study of Mandarin Verbs of Emotion
Li-Li Chang, Keh-Jiann Chen, Chu-Ren Huang |
PACLIC | 3 |
| 1999 | Lexical Information and Beyond : Constructional Inferences in Semantic Representation
Mei-Chun Liu, Chu-Ren Huang, Ching-Yi Lee |
PACLIC | 2 |
| 1996 | Segmentation Standard for Chinese Natural Language Processing
Chu-Ren Huang, Keh-Jiann Chen, Li-Li Chang |
COLING | 1 |
| 1996 | Classifiers and Semantic Type Coercion : Motivating a New Classification of Classifiers
Kathleen Ahrens, Chu-Ren Huang |
PACLIC | 2 |
| 1996 | SINICA CORPUS : Design Methodology for Balanced Corpora
Keh-Jiann Chen, Chu-Ren Huang, Li-Ping Chang, Hui-Li Hsu |
PACLIC | 2 |
| 1995 | Construction as a Theoretical Entity : An Argument Based on Mandarin Existential Sentences
Chao-ran Chen, Chu-Ren Huang, Kathleen Ahrens |
PACLIC | 2 |
| 1994 | Character-based Collocation for Mandarin Chinese
Chu-Ren Huang, Keh-Jiann Chen, Yun-yan Yang |
COLING | 1 |
| 1992 | A Chinese Corpus for Linguistic Research
Chu-Ren Huang, Keh-Jiann Chen |
COLING | 1 |
| 1990 | Information-based Case Grammar
Keh-Jiann Chen, Chu-Ren Huang |
COLING | 2 |