VLDB 2026 Research / reviewers in the wild / expert
Jason S. Chang
dblp:93/641
· DBLP profile ↗
39ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-8227-7382ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Artificial intelligence
7 papers |
Information extraction and text analysis · 51% Machine translation · 26% Language models and text generation · 23% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 90% Web and social media mining · 10% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
review-based recommendation |
0.7 | 1 | 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023 |
Security and privacy of machine learning
recommender system attack |
0.7 | 1 | 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023 |
Security and privacy of machine learning › recommender system attack
shilling attacks |
0.7 | 1 | 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
syntactic disambiguation |
0.2 | 1 | 2014 | Ambiguity Resolution for Vt-N Structures in Chinese · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › lexical semantics
multiword expression |
0.1 | 1 | 2009 | Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies · EMNLP 2009 |
Natural language and speech › Machine translation
translational equivalence |
0.1 | 1 | 2009 | Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies · EMNLP 2009 |
Natural language and speech › Machine translation
transliteration |
0.1 | 1 | 2007 | Learning to Find English to Chinese Transliterations on the Web · EMNLP-CoNLL 2007 |
Web and social media mining
web mining |
0.1 | 1 | 2007 | Learning to Find English to Chinese Transliterations on the Web · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
chinese parsing |
0.1 | 1 | 2014 | Ambiguity Resolution for Vt-N Structures in Chinese · EMNLP 2014 |
Natural language and speech › Machine translation
terminology translation |
0.1 | 1 | 2005 | Learning Source-Target Surface Patterns for Web-based Terminology Translation · ACL 2005 |
Natural language and speech › Information extraction and text analysis › phrase extraction
terminology extraction |
0.0 | 1 | 2005 | Learning Source-Target Surface Patterns for Web-based Terminology Translation · ACL 2005 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.0pre-trained language model · 2.0adversarial training · 2.0semi-supervised learning · 0.2classifier · 0.2PCFG parser · 0.2transliteration mining · 0.1web sentence acquisition · 0.1pattern-based generation · 0.1normalized correlation frequencies · 0.1language model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improve & Explain: AI-Based Writing Improvement and Informative Feedback
Wei-Chin Lee, Kai-Wen Tuan, Cheng-En Hsu, Jo-Chi Hsiao, Hai-Lun Tu, Jason S. Chang |
ICALT | 6 |
| 2025 | CleverFox: Integrating Visual Mnemonics with AI for Enhanced Language Learning
Yung-Chu Chiang, Zi-Xian Tang, Yi-Ching Luo, Jason S. Chang |
MMM (5) | 4 |
| 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review GenerationabstractReview-Based Recommender Systems (RBRS) have attracted increasing research interest due to their ability to alleviate well-known cold-start problems. RBRS utilizes reviews to construct the user and items representations. However, in this paper, we argue that such a reliance on reviews may instead expose systems to the risk of being shilled. To explore this possibility, in this paper, we propose the first generation-based model for shilling attacks against RBRSs. Specifically, we learn a fake review generator through reinforcement learning, which maliciously promotes items by forcing prediction shifts after adding generated reviews to the system. By introducing the auxiliary rewards to increase text fluency and diversity with the aid of pre-trained language models and aspect predictors, the generated reviews can be effective for shilling with high fidelity. Experimental results demonstrate that the proposed framework can successfully attack three different kinds of RBRSs on the Amazon corpus with three domains and Yelp corpus. Furthermore, human studies also show that the generated reviews are fluent and informative. Finally, equipped with Attack Review Generators (ARGs), RBRSs with adversarial training are much more robust to malicious reviews. Hung-Yun Chiang, Yi-Syuan Chen, Yun-Zhu Song, Hong-Han Shuai, Jason S. Chang |
KDD | 5 |
| 2023 | Learning to Paraphrase Sentences to Different Complexity LevelsabstractAbstract While sentence simplification is an active research topic in NLP, its adjacent tasks of sentence complexification and same-level paraphrasing are not. To train models on all three tasks, we present two new unsupervised datasets. We compare these datasets, one labeled by a weak classifier and the other by a rule-based approach, with a single supervised dataset. Using these three datasets for training, we perform extensive experiments on both multitasking and prompting strategies. Compared to other systems trained on unsupervised parallel data, models trained on our weak classifier labeled dataset achieve state-of-the-art performance on the ASSET simplification benchmark. Our models also outperform previous work on sentence-level targeting. Finally, we establish how a handful of Large Language Models perform on these tasks under a zero-shot setting. Alison Chi, Li-Kuang Chen, Yi-Chen Chang, Shu-Hui Lee, Jason S. Chang |
Trans. Assoc. Comput. Linguistics | 5 |
| 2019 | Extracting Grammatical Error Corrections from Wikipedia Revision HistoryabstractThis paper describes the process of extracting and filtering Wikipedia revision history as a resource for grammatical error correction (GEC). Edits in Wikipedia revision history vary widely, including grammatical error corrections, information supplements, format amendments, and even vandalism. To extract only GEC-related revisions, we use an automated error annotation toolkit, ERRANT1, and extend it to process large data in parallel efficiently. With error-type analysis, we can then identify GEC-related edits and omit other unrelated edits (i.e., only the correction parts are reserved). The resulting corpus is - to our knowledge - the largest publicly available corpus of parallel possibly erroneous and correct sentences with error type labels. Jhih-Jie Chen, Yi-Dong Wu, Yu-Chuan Tai, Ching-Yu Helen Yang, Hai-Lun Tu, Jason S. Chang |
IEEE BigData | 6 |
| 2019 | Word Sense Disambiguation Using Wikipedia Link GraphabstractIn this paper, we present a observation that a word with some special sense should appear in specific category of articles, and this appearance will “take” the word to its Wiki sense [1] by Wikipedia link graph. This property can help to solve word sense disambiguation problem using Wikipedia. Hai-Lun Tu, Peichen Ho, Jason S. Chang, Li-Guang Chen |
IEEE BigData | 3 |
| 2018 | Chinese Spelling Check based on Neural Machine Translation
Chiao-Wen Li, Jhih-Jie Chen, Jason S. Chang |
PACLIC | 3 |
| 2018 | SmartWrite: Extracting Chinese Lexical Grammar Patterns Using Dependency Parsing
Cheng-Cyuan Peng, Ching-Yu Helen Yang, Jhih-Jie Chen, Jason S. Chang |
PACLIC | 4 |
| 2015 | WriteAhead2: Mining Lexical Grammar Patterns for Assisted WritingabstractThis paper describes WriteAhead2, an interactive writing environment that provides lexical and grammatical suggestions for second language learners, and helps them write fluently and avoid common writing errors. The method involves learning phrase templates from dictionary examples, and extracting grammar patterns with example phrases from an academic corpus. At run-time, as the user types word after word, the actions trigger a list after list of suggestions. Each successive list contains grammar patterns and examples, most relevant to the half-baked sentence. WriteAhead2 facilitates steady, timely, and spot-on interactions between learner writers and relevant information for effective assisted writing. Preliminary experiments show that WriteAhead2 has the potential to induce better writing and improve writing skills. Jim Chang, Jason S. Chang |
HLT-NAACL | 2 |
| 2015 | Learning Sentential Patterns of Various Rhetoric Moves for Assisted Academic Writing
Jim Chang, Hsiang-Ling Hsu, Joanne Boisson, Hao-Chun Peng, Yu-Hsuan Wu, Jason S. Chang |
PACLIC | 6 |
| 2014 | Ambiguity Resolution for Vt-N Structures in ChineseabstractThe syntactic ambiguity of a transitive verb (Vt) followed by a noun (N) has long been a problem in Chinese parsing. In this paper, we propose a classifier to resolve the ambiguity of Vt-N structures. The design of the classifier is based on three important guidelines, namely, adopting linguistically motivated features, using all available resources, and easy in-tegration into a parsing model. The lin-guistically motivated features include semantic relations, context, and morpho-logical structures; and the available re-sources are treebank, thesaurus, affix da-tabase, and large corpora. We also pro-pose two learning approaches that resolve the problem of data sparseness by auto-parsing and extracting relative knowledge from large-scale unlabeled data. Our experiment results show that the Vt-N classifier outperforms the cur-rent PCFG parser. Furthermore, it can be easily and effectively integrated into the PCFG parser and general statistical pars-ing models. Evaluation of the learning approaches indicates that world knowledge facilitates Vt-N disambigua-tion through data selection and error cor-rection. 1 Yu-Ming Hsieh, Jason S. Chang, Keh-Jiann Chen |
EMNLP | 2 |
| 2014 | TakeTwo: A Word Aligner based on Self Learning
Jim Chang, Jian-Cheng Wu, Jason S. Chang |
PACLIC | 3 |
| 2013 | Translating Chinese Unknown Words by Automatically Acquired Templates
Ming-Hong Bai, Yu-Ming Hsieh, Keh-Jiann Chen, Jason S. Chang |
IJCNLP | 4 |
| 2013 | Augmentable Paraphrase Extraction Framework
Mei-hua Chen, Shih-Ting Huang, Jason S. Chang |
IJCNLP | 4 |
| 2013 | A Computer-Assisted Translation and Writing SystemabstractWe introduce a method for learning to predict text and grammatical construction in a computer-assisted translation and writing framework. In our approach, predictions are offered on the fly to help the user make appropriate lexical and grammar choices during the translation of a source text, thus improving translation quality and productivity. The method involves automatically generating general-to-specific word usage summaries (i.e., writing suggestion module), and automatically learning high-confidence word- or phrase-level translation equivalents (i.e., translation suggestion module). At runtime, the source text and its translation prefix entered by the user are broken down into n-grams to generate grammar and translation predictions, which are further combined and ranked via translation and language models. These ranked prediction candidates are iteratively and interactively displayed to the user in a pop-up menu as translation or writing hints. We present a prototype writing assistant, TransAhead , that applies the method to a human-computer collaborative environment. Automatic and human evaluations show that novice translators or language learners substantially benefit from our system in terms of translation performance (i.e., translation accuracy and productivity) and language learning (i.e., collocation usage and grammar). In general, our methodology of inline grammar and text predictions or suggestions has great potential in the field of computer-assisted translation, writing, or even language learning. Chung-Chi Huang, Mei-hua Chen, Ping-Che Yang, Jason S. Chang |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2012 | TransAhead: A Writing Assistant for CAT and CALL
Chung-Chi Huang, Ping-Che Yang, Mei-hua Chen, Hung-ting Hsieh, Ting-hui Kao, Jason S. Chang |
EACL | 6 |
| 2012 | TransAhead: A Computer-Assisted Translation and Writing Tool
Chung-Chi Huang, Ping-Che Yang, Keh-Jiann Chen, Jason S. Chang |
HLT-NAACL | 4 |
| 2011 | Using Sublexical Translations to Handle the OOV Problem in Machine TranslationabstractWe introduce a method for learning to translate out-of-vocabulary (OOV) words. The method focuses on combining sublexical/constituent translations of an OOV to generate its translation candidates. In our approach, wildcard searches are formulated based on our OOV analysis, aimed at maximizing the probability of retrieving OOVs’ sublexical translations from existing resources of Machine Translation (MT) systems. At run-time, translation candidates of the unknown words are generated from their suitable sublexical translations and ranked based on monolingual and bilingual information. We have incorporated the OOV model into a state-of-the-art machine translation system and experimental results show that our model indeed helps to ease the impact of OOVs on translation quality, especially for sentences containing more OOVs (significant improvement). Chung-Chi Huang, Ho-Ching Yen, Ping-Che Yang, Shih-Ting Huang, Jason S. Chang |
ACM Trans. Asian Lang. Inf. Process. | 5 |
| 2010 | GRASP: Grammar- and Syntax-based Pattern-Finder for Collocation and Phrase Learning
Mei-hua Chen, Chung-Chi Huang, Shih-Ting Huang, Jason S. Chang |
PACLIC | 4 |
| 2009 | Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies
Ming-Hong Bai, Jia-Ming You, Keh-Jiann Chen, Jason S. Chang |
EMNLP | 4 |
| 2009 | Learning Bilingual Linguistic Reordering Model for Statistical Machine Translation
Han-Bin Chen, Jian-Cheng Wu, Jason S. Chang |
HLT-NAACL | 3 |
| 2009 | WikiSense: Supersense Tagging of Wikipedia Named Entities Based WordNet
Joseph Z. Chang, Richard Tzong-Han Tsai, Jason S. Chang |
PACLIC | 3 |
| 2009 | Review Classification Using Semantic Features and Run-Time Weighting
Chung-Chi Huang, Meng-chiech Lee, Zhe-nan Lin, Jason S. Chang |
PACLIC | 4 |
| 2009 | Extending Bilingual WordNet via Hierarchical Word Translation Classification
Tzu-yi Nien, Tsun Ku, Chung-Chi Huang, Mei-hua Chen, Jason S. Chang |
PACLIC | 5 |
| 2008 | Bilingual Segmentation for Alignment and Translation
Chung-Chi Huang, Wei-Teh Chen, Jason S. Chang |
CICLing | 3 |
| 2008 | Improving Word Alignment by Adjusting Chinese Word Segmentation
Ming-Hong Bai, Keh-Jiann Chen, Jason S. Chang |
IJCNLP | 3 |
| 2007 | Learning to Find English to Chinese Transliterations on the Web
Jian-Cheng Wu, Jason S. Chang |
EMNLP-CoNLL | 2 |
| 2006 | FAST - An Automatic Generation System for Grammar TestsabstractThis paper introduces a method for the semi-automatic generation of grammar test items by applying Natural Language Processing (NLP) techniques. Based on manually-designed patterns, sentences gathered from the Web are transformed into tests on grammaticality. The method involves representing test writing knowledge as test patterns, acquiring authentic sentences on the Web, and applying generation strategies to transform sentences into items. At runtime, sentences are converted into two types of TOEFL-style question: multiple-choice and error detection. We also describe a prototype system FAST (Free Assessment of Structural Tests). Evaluation on a set of generated questions indicates that the proposed method performs satisfactory quality. Our methodology provides a promising approach and offers significant potential for computer assisted language learning and assessment. Chia-Yin Chen, Hsien-Chin Liou, Jason S. Chang |
ACL | 3 |
| 2006 | Computational Analysis of Move Structures in Academic AbstractsabstractThis paper introduces a method for computational analysis of move structures in abstracts of research articles. In our approach, sentences in a given abstract are analyzed and labeled with a specific move in light of various rhetorical functions. The method involves automatically gathering a large number of abstracts from the Web and building a language model of abstract moves. We also present a prototype concordancer, CARE, which exploits the move-tagged abstracts for digital learning. This system provides a promising approach to Web-based computer-assisted academic writing. Jien-Chen Wu, Yu-Chia Chang, Hsien-Chin Liou, Jason S. Chang |
ACL | 4 |
| 2006 | Extraction of transliteration pairs from parallel corpora using a statistical transliteration model
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
Inf. Sci. | 2 |
| 2006 | Alignment of bilingual named entities in parallel corpora using statistical models and multiple knowledge sourcesabstractNamed entity (NE) extraction is one of the fundamental tasks in natural language processing (NLP). Although many studies have focused on identifying NEs within monolingual documents, aligning NEs in bilingual documents has not been investigated extensively due to the complexity of the task. In this article we introduce a new approach to aligning bilingual NEs in parallel corpora by incorporating statistical models with multiple knowledge sources. In our approach, we model the process of translating an English NE phrase into a Chinese equivalent using lexical translation/transliteration probabilities for word translation and alignment probabilities for word reordering. The method involves automatically learning phrase alignment and acquiring word translations from a bilingual phrase dictionary and parallel corpora, and automatically discovering transliteration transformations from a training set of name-transliteration pairs. The method also involves language-specific knowledge functions, including handling abbreviations, recognizing Chinese personal names, and expanding acronyms. At runtime, the proposed models are applied to each source NE in a pair of bilingual sentences to generate and evaluate the target NE candidates; the source and target NEs are then aligned based on the computed probabilities. Experimental results demonstrate that the proposed approach, which integrates statistical models with extra knowledge sources, is highly feasible and offers significant improvement in performance compared to our previous work, as well as the traditional approach of IBM Model 4. Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Learning Source-Target Surface Patterns for Web-based Terminology Translation
Jian-Cheng Wu, Tracy Lin, Jason S. Chang |
ACL | 3 |
| 2005 | Web-Based Unsupervised Learning for Query Formulation in Question Answering
Yi-Chia Wang, Jian-Cheng Wu, Tyne Liang, Jason S. Chang |
IJCNLP | 4 |
| 2004 | Bilingual Sentence Alignment Based on Punctuation Statistics and Lexicon
Thomas C. Chuang, Jian-Cheng Wu, Tracy Lin, Wen-Chie Shei, Jason S. Chang |
IJCNLP | 5 |
| 2003 | A Statistical Approach to Chinese-to-English Back-Transliteration
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
PACLIC | 2 |
| 2002 | An Operator Assisted Call Routing System
Chun-Jen Lee, Jason S. Chang |
PACLIC | 2 |
| 1998 | Topical Clustering of MRD Senses Based on Information Retrieval Techniques
Jen Nan Chen, Jason S. Chang |
Comput. Linguistics | 2 |
| 1997 | An Alignment Method for Noisy Parallel Corpora based on Image Processing TechniquesabstractThis paper presents a new approach to bitext correspondence problem (BCP) of noisy bilingual corpora based on image processing (IP) techniques. By using one of several ways of estimating the lexical translation probability (LTP) between pairs of source and target words, we can turn a bitext into a discrete gray-level image. We contend that the BCP, when seen in the light, bears a striking resemblance to the line detection problem in IP. Therefore, BCPs, including sentence and word alignment, can benefit from a wealth of effective, well established IP techniques, including convolution-based filters, texture analysis and Hough transform. This paper describes a new program, PlotAlign that produces a word-level bitext map for noisy or non-literal bitext, based on these techniques. Jason S. Chang, Mathis H. M. Chen |
ACL | 1 |
| 1997 | A Class-based Approach to Word Alignment
Sue J. Ker, Jason S. Chang |
Comput. Linguistics | 2 |