VLDB 2026 Research / reviewers in the wild / expert
Shu-Kai Hsieh
dblp:05/864
· DBLP profile ↗
34ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-9674-1249ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Structure Matters: Cross-Lingual Hyperbolic Embeddings for Chinese and English Wordnets
Mao-Chang Ku, Da-Chen Lian, Pin-Er Chen, Po-Ya Angela Wang, Wei-Ling Chen, Shu-Kai Hsieh |
LREC | 6 |
| 2025 | Zero-Shot Evaluation of Conversational Language Competence in Data-Efficient LLMs Across English, Mandarin, and FrenchabstractLarge Language Models (LLMs) have achieved oustanding performance across various natural language processing tasks, including those from Discourse and Dialogue traditions. However, these achievements are typically obtained thanks to pretraining on huge datasets. In contrast, humans learn to speak and communicate through dialogue and spontaneous speech with only a fraction of the language exposure. This disparity has spurred interest in evaluating whether smaller, more carefully selected and curated pretraining datasets can support robust performance on specific tasks. Drawing inspiration from the BabyLM initiative, we construct small (10M-token) pretraining datasets from different sources, including conversational transcripts and Wikipedia-style text. To assess the impact of these datasets, we develop evaluation benchmarks focusing on discourse and interactional markers, extracted from high-quality spoken corpora in English, French, and Mandarin. Employing a zero-shot classification framework inspired by the BLiMP benchmark, we design tasks wherein the model must determine, between a genuine utterance extracted from a corpus and its minimally altered counterpart, which one is the authentic instance. Our findings reveal that the nature of pretraining data significantly influences model performance on discourse-related tasks. Models pretrained on conversational data exhibit a clear advantage in handling discourse and interactional markers compared to those trained on written or encyclopedic text. Furthermore, the models, trained on small amount spontaneous speech transcripts, perform comparably to standard LLMs. Sheng-Fu Wang, Ri-Sheng Huang, Shu-Kai Hsieh, Laurent Prévot 0001 |
SIGDIAL | 3 |
| 2023 | Lexical Retrieval Hypothesis in Multimodal Context
Po-Ya Angela Wang, Pin-Er Chen, Hsin-Yu Chou, Yu-Hsiang Tseng, Shu-Kai Hsieh |
LDK | 5 |
| 2023 | Exploring Affordance and Situated Meaning in Image Captions: A Multimodal Analysis
Pin-Er Chen, Po-Ya Angela Wang, Hsin-Yu Chou, Yu-Hsiang Tseng, Shu-Kai Hsieh |
PACLIC | 5 |
| 2023 | Vec2Gloss: definition modeling leveraging contextualized vectors with Wordnet gloss
Yu-Hsiang Tseng, Mao-Chang Ku, Wei-Ling Chen, Yu-Lin Chang, Shu-Kai Hsieh |
PACLIC | 5 |
| 2022 | Character Jacobian: Modeling Chinese Character Meanings with Deep Learning ModelabstractCompounding, a prevalent word-formation process, presents an interesting challenge for computational models. Indeed, the relations between compounds and their constituents are often complicated. It is particularly so in Chinese morphology, where each character is almost simultaneously bound and free when treated as a morpheme. To model such word-formation process, we propose the Notch (NOnlinear Transformation of CHaracter embeddings) model and the character Jacobians. The Notch model first learns the non-linear relations between the constituents and words, and the character Jacobians further describes the character’s role in each word. In a series of experiments, we show that the Notch model predicts the embeddings of the real words from their constituents but helps account for the behavioral data of the pseudowords. Moreover, we also demonstrated that character Jacobians reflect the characters’ meanings. Taken together, the Notch model and character Jacobians may provide a new perspective on studying the word-formation process and morphology with modern deep learning. Yu-Hsiang Tseng, Shu-Kai Hsieh |
COLING | 2 |
| 2022 | CxLM: A Construction and Context-aware Language ModelabstractConstructions are direct form-meaning pairs with possible schematic slots. These slots are simultaneously constrained by the embedded construction itself and the sentential context. We propose that the constraint could be described by a conditional probability distribution. However, as this conditional probability is inevitably complex, we utilize language models to capture this distribution. Therefore, we build CxLM, a deep learning-based masked language model explicitly tuned to constructions’ schematic slots. We first compile a construction dataset consisting of over ten thousand constructions in Taiwan Mandarin. Next, an experiment is conducted on the dataset to examine to what extent a pretrained masked language model is aware of the constructions. We then fine-tune the model specifically to perform a cloze task on the opening slots. We find that the fine-tuned model predicts masked slots more accurately than baselines and generates both structurally and semantically plausible word samples. Finally, we release CxLM and its dataset as publicly available resources and hope to serve as new quantitative tools in studying construction grammar. Yu-Hsiang Tseng, Cing-Fang Shih, Pin-Er Chen, Hsin-Yu Chou, Mao-Chang Ku, Shu-Kai Hsieh |
LREC | 6 |
| 2021 | Examine persuasion strategies in Chinese on social media
Yu-Yun Chang, Po-Ya Angela Wang, Han-Tang Hung, Ka-Sîng Khóo, Shu-Kai Hsieh |
PACLIC | 5 |
| 2021 | Exploring sentiment constructions: connecting deep learning models with linguistic construction
Shu-Kai Hsieh, Yu-Hsiang Tseng |
PACLIC | 1 |
| 2020 | Computational Modeling of Affixoid Behavior in Chinese MorphologyabstractThe morphological status of affixes in Chinese has long been a matter of debate.How one might apply the conventional criteria of free/bound and content/function features to distinguish word-forming affixes from bound roots in Chinese is still far from clear.Issues involving polysemy and diachronic dynamics further blur the boundaries.In this paper, we propose three quantitative features in a computational model of affixoid behavior in Mandarin Chinese.The results show that, except for in a very few cases, there are no clear criteria that can be used to identify an affix's status in an isolating language like Chinese.A diachronic check using contextualized embeddings with the WordNet Sense Inventory also demonstrates the possible role of the polysemy of lexical roots across diachronic settings. Yu-Hsiang Tseng, Shu-Kai Hsieh, Pei-Yi Chen, Sara Court |
COLING | 2 |
| 2020 | Do You Believe It Happened? Assessing Chinese Readers' Veridicality JudgmentsabstractThis work collects and studies Chinese readers’ veridicality judgments to news events (whether an event is viewed as happening or not). For instance, in “The FBI alleged in court documents that Zazi had admitted having a handwritten recipe for explosives on his computer”, do people believe that Zazi had a handwritten recipe for explosives? The goal is to observe the pragmatic behaviors of linguistic features under context which affects readers in making veridicality judgments. Exploring from the datasets, it is found that features such as event-selecting predicates (ESP), modality markers, adverbs, temporal information, and statistics have an impact on readers’ veridicality judgments. We further investigated that modality markers with high certainty do not necessarily trigger readers to have high confidence in believing an event happened. Additionally, the source of information introduced by an ESP presents low effects to veridicality judgments, even when an event is attributed to an authority (e.g. “The FBI”). A corpus annotated with Chinese readers’ veridicality judgments is released as the Chinese PragBank for further analysis. Yu-Yun Chang, Shu-Kai Hsieh |
LREC | 2 |
| 2020 | From Sense to Action: A Word-Action Disambiguation Task in NLP
Shu-Kai Hsieh, Yu-Hsiang Tseng, Chiung-Yu Chiang, Richard Lian, Yong-fu Liao, Mao-Chang Ku, Ching-Fang Shih |
PACLIC | 1 |
| 2020 | Exploring Discourse of Same-sex Marriage in Taiwan: A Case Study of Near-Synonym of HOMOSEXUAL in Opposing Stances
Han-Tang Hung, Shu-Kai Hsieh |
PACLIC | 2 |
| 2019 | The Secret to Popular Chinese Web Novels: A Corpus-Driven StudyabstractWhat is the secret to writing popular novels? The issue is an intriguing one among researchers from various fields. The goal of this study is to identify the linguistic features of several popular web novels as well as how the textual features found within and the overall tone interact with the genre and themes of each novel. Apart from writing style, non-textual information may also reveal details behind the success of web novels. Since web fiction has become a major industry with top writers making millions of dollars and their stories adapted into published books, determining essential elements of "publishable" novels is of importance. The present study further examines how non-textual information, namely, the number of hits, shares, favorites, and comments, may contribute to several features of the most popular published and unpublished web novels. Findings reveal that keywords, function words, and lexical diversity of a novel are highly related to its genres and writing style while dialogue proportion shows the narration voice of the story. In addition, relatively shorter sentences are found in these novels. The data also reveal that the number of favorites and comments serve as significant predictors for the number of shares and hits of unpublished web novels, respectively; however, the number of hits and shares of published web novels is more unpredictable. Yi-Ju Lin, Shu-Kai Hsieh |
LDK | 2 |
| 2019 | Augmenting Chinese WordNet semantic relations with contextualized embeddingsabstractConstructing semantic relations in WordNet has been a labour-intensive task, especially in a dynamic and fastchanging language environment.Combined with recent advancements of contextualized embeddings, this paper proposes the concept of morphologyguided sense vectors, which can be used to semi-automatically augment semantic relations in Chinese Wordnet (CWN).This paper (1) built sense vectors with pre-trained contextualized embedding models; (2) demonstrated the sense vectors computed were consistent with the sense distinctions made in CWN; and (3) predicted the potential semantically-related sense pairs with high accuracy by sense vectors model. Yu-Hsiang Tseng, Shu-Kai Hsieh |
GWC | 2 |
| 2018 | Fluid Annotation: A Granularity-aware Annotation Tool for Chinese Word Fluidity
Shu-Kai Hsieh, Yu-Hsiang Tseng, Chih-yao Lee, Chiung-Yu Chiang |
LREC | 1 |
| 2018 | Sinitic Wordnet: Laying the Groundwork with Chinese Varieties Written in Traditional CharactersabstractThe present work seeks to make the logographic nature of Chinese script a relevant research ground in wordnet studies.While wordnets are not so much about words as about the concepts represented in words, synset formation inevitably involves the use of orthographic and/or phonetic representations to serve as headword for a given concept.For wordnets of Chinese languages, if their synsets are mapped with each other, the connection from logographic forms to lexicalized concepts can be explored backwards to, for instance, help trace the development of cognates in different varieties of Chinese.The Sinitic Wordnet project is an attempt to construct such an integrated wordnet that aggregates three Chinese varieties that are widely spoken in Taiwan and all written in traditional Chinese characters. Chih-yao Lee, Shu-Kai Hsieh |
GWC | 2 |
| 2016 | Evaluative Pattern Extraction for Automated Text GenerationabstractGetting travel tips from the experienced bloggers and online forums has been one of the important supplements to the travel guidebook in the web society.In this paper we present a novel approach by identifying and extracting evaluative patterns, providing a different linguistically-motivated framework for automated evaluative text generation.We target at domain-specific observation in online travel blogs in Chinese.Results suggest that the semantic prosody accompanying the patterns demonstrates that online travel bloggers prefer to employ tacit pragmatic strategy in presenting their sentiment polarity in comments.The extracted patterns and their differentiation can be beneficial to identifying and characterizing evaluative language for further automated opinion summarization and macro/micro planning in natural language generation (NLG) as well. Chia-Chen Lee, Shu-Kai Hsieh |
INLG | 2 |
| 2015 | An Arguing Lexicon for Stance Classification on Short Text Comments in Chinese
Ju-Han Chuang, Shu-Kai Hsieh |
PACLIC | 2 |
| 2014 | Skillex, an action labelling efficiency score: the case for French and Mandarin
Yann Desalle, Bruno Gaume, Karine Duvignau, Hintat Cheung, Shu-Kai Hsieh, Pierre Magistry, Jean-Luc Nespoulous |
CogSci | 5 |
| 2014 | Why Chinese Web-as-Corpus is Wacky? Or: How Big Data is Killing Chinese Corpus Linguistics
Shu-Kai Hsieh |
LREC | 1 |
| 2014 | Leveraging Morpho-semantics for the Discovery of Relations in Chinese WordnetabstractSemantic relations of different types have played an important role in wordnet, and have been widely recognized in various fields.In recent years, with the growing interests of constructing semantic network in support of intelligent systems, automatic semantic relation discovery has become an urgent task.This paper aims to extract semantic relations relying on the in situ morpho-semantic structure in Chinese which can dispense of an outside source such as corpus or web data.Manual evaluation of thousands of word pairs shows that most relations can be successful predicted.We believe that it can serve as a valuable starting point in complementing with other approaches, which will hold promise for the robust lexical relations acquisition. Shu-Kai Hsieh, Yu-Yun Chang |
GWC | 1 |
| 2012 | Chinese Sentiments on the Clouds: A Preliminary Experiment on Corpus Processing and Exploration on Cloud Service
Shu-Kai Hsieh, Yu-Yun Chang, Meng-Xian Shih |
PACLIC | 1 |
| 2010 | Towards an Automatic Measurement of Verbal Lexicon Acquisition: The Case for a Young Children-versus-Adults Classification in French and Mandarin
Yann Desalle, Shu-Kai Hsieh, Bruno Gaume, Hintat Cheung |
PACLIC | 2 |
| 2010 | Graph Representation of Synonymy and Translation Resources for Crosslinguistic Modelisation of Meaning
Benoît Gaillard, Yannick Chudy, Pierre Magistry, Shu-Kai Hsieh, Emmanuel Navarro |
PACLIC | 4 |
| 2009 | Bridging the Gap between Graph Modeling and Developmental Psycholinguistics: An Experiment on Measuring Lexical Proximity in Chinese Semantic Space
Shu-Kai Hsieh, Chun-Han Chang, Ivy Kuo, Hintat Cheung, Chu-Ren Huang, Bruno Gaume |
PACLIC | 1 |
| 2008 | Constructing Taxonomy of Numerative Classifiers for Asian Languages
Kiyoaki Shirai, Takenobu Tokunaga, Chu-Ren Huang, Shu-Kai Hsieh, Tzu-Yi Kuo, Virach Sornlertlamvanich, Thatsanee Charoenporn |
IJCNLP | 4 |
| 2008 | Extracting Concrete Senses of Lexicon through Measurement of Conceptual Similarity in Ontologies
Siaw-Fong Chung, Laurent Prévot 0001, Kathleen Ahrens, Shu-Kai Hsieh, Chu-Ren Huang |
LREC | 5 |
| 2008 | Adapting International Standard for Asian Language Technologies
Takenobu Tokunaga, Dain Kaplan, Chu-Ren Huang, Shu-Kai Hsieh, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Kiyoaki Shirai, Virach Sornlertlamvanich, Thatsanee Charoenporn, Yingju Xia |
LREC | 4 |
| 2008 | KYOTO: a System for Mining, Structuring and Distributing Knowledge across Languages and Cultures
Piek Vossen, Eneko Agirre, Nicoletta Calzolari, Christiane Fellbaum, Shu-Kai Hsieh, Chu-Ren Huang, Hitoshi Isahara, Kyoko Kanzaki, Andrea Marchetti, Monica Monachini, Federico Neri, Remo Raffaelli, German Rigau, Maurizio Tesconi, Joop VanGent |
LREC | 5 |
| 2007 | Automatic Discovery of Named Entity Variants: Grammar-driven Approaches to Non-Alphabetical Transliterations
Chu-Ren Huang, Petr Simon, Shu-Kai Hsieh |
ACL | 3 |
| 2007 | Rethinking Chinese Word Segmentation: Tokenization, Character Classification, or Wordbreak Identification
Chu-Ren Huang, Petr Simon, Shu-Kai Hsieh, Laurent Prévot 0001 |
ACL | 3 |
| 2006 | When Conset Meets Synset: A Preliminary Survey of an Ontological Lexical Resource Based on Chinese Characters
Shu-Kai Hsieh, Chu-Ren Huang |
ACL | 1 |
| 2006 | GuangQunFangPu: e-Humanities Combining Textual and Botanic InformationabstractIn this paper, we propose a lexicon-driven and ontologymerging methodology of constructing diachronic domain knowledge via Sinica BOW, a bilingual ontological lexical resource based on WordNet and SUMO ontology. The main domain knowledge that we model our specialized ontology on is GuangQunFangPu, a Chinese classic literature of botany. Our studies yields promising result, we believe that the proposed research scenario will boost the on-going development of e-Humanities. Shu-Kai Hsieh, Shu-Ming Chang, Chun-Han Chang, Yi-Shuan Zhou, Chu-Ren Huang, Feng-Ju Lo, Ru-Yng Chang |
e-Science | 1 |