Jason S. Chang

dblp:93/641 · DBLP profile ↗
← Back
39ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-8227-7382ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 100%
Artificial intelligence
7 papers
Information extraction and text analysis · 51% Machine translation · 26% Language models and text generation · 23%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 90% Web and social media mining · 10%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
review-based recommendation
0.712023
Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023
Security and privacy of machine learning
recommender system attack
0.712023
Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023
Security and privacy of machine learning › recommender system attack
shilling attacks
0.712023
Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023
Natural language and speech › Language models and text generation
text generation
0.212023
Shilling Black-box Review-based Recommender Systems through Fake Review Generation · KDD 2023
Natural language and speech › Information extraction and text analysis › syntactic parsing
syntactic disambiguation
0.212014
Ambiguity Resolution for Vt-N Structures in Chinese · EMNLP 2014
Natural language and speech › Information extraction and text analysis › lexical semantics
multiword expression
0.112009
Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies · EMNLP 2009
Natural language and speech › Machine translation
translational equivalence
0.112009
Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies · EMNLP 2009
Natural language and speech › Machine translation
transliteration
0.112007
Learning to Find English to Chinese Transliterations on the Web · EMNLP-CoNLL 2007
Web and social media mining
web mining
0.112007
Learning to Find English to Chinese Transliterations on the Web · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis › syntactic parsing
chinese parsing
0.112014
Ambiguity Resolution for Vt-N Structures in Chinese · EMNLP 2014
Natural language and speech › Machine translation
terminology translation
0.112005
Learning Source-Target Surface Patterns for Web-based Terminology Translation · ACL 2005
Natural language and speech › Information extraction and text analysis › phrase extraction
terminology extraction
0.012005
Learning Source-Target Surface Patterns for Web-based Terminology Translation · ACL 2005

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.0pre-trained language model · 2.0adversarial training · 2.0semi-supervised learning · 0.2classifier · 0.2PCFG parser · 0.2transliteration mining · 0.1web sentence acquisition · 0.1pattern-based generation · 0.1normalized correlation frequencies · 0.1language model · 0.1
YearPublicationVenuePosition
2025 Improve & Explain: AI-Based Writing Improvement and Informative Feedback
Wei-Chin Lee, Kai-Wen Tuan, Cheng-En Hsu, Jo-Chi Hsiao, Hai-Lun Tu, Jason S. Chang
ICALT6
2025 CleverFox: Integrating Visual Mnemonics with AI for Enhanced Language Learning
Yung-Chu Chiang, Zi-Xian Tang, Yi-Ching Luo, Jason S. Chang
MMM (5)4
2023 Shilling Black-box Review-based Recommender Systems through Fake Review Generation
abstract
Review-Based Recommender Systems (RBRS) have attracted increasing research interest due to their ability to alleviate well-known cold-start problems. RBRS utilizes reviews to construct the user and items representations. However, in this paper, we argue that such a reliance on reviews may instead expose systems to the risk of being shilled. To explore this possibility, in this paper, we propose the first generation-based model for shilling attacks against RBRSs. Specifically, we learn a fake review generator through reinforcement learning, which maliciously promotes items by forcing prediction shifts after adding generated reviews to the system. By introducing the auxiliary rewards to increase text fluency and diversity with the aid of pre-trained language models and aspect predictors, the generated reviews can be effective for shilling with high fidelity. Experimental results demonstrate that the proposed framework can successfully attack three different kinds of RBRSs on the Amazon corpus with three domains and Yelp corpus. Furthermore, human studies also show that the generated reviews are fluent and informative. Finally, equipped with Attack Review Generators (ARGs), RBRSs with adversarial training are much more robust to malicious reviews.
Hung-Yun Chiang, Yi-Syuan Chen, Yun-Zhu Song, Hong-Han Shuai, Jason S. Chang
KDD5
2023 Learning to Paraphrase Sentences to Different Complexity Levels
abstract
Abstract While sentence simplification is an active research topic in NLP, its adjacent tasks of sentence complexification and same-level paraphrasing are not. To train models on all three tasks, we present two new unsupervised datasets. We compare these datasets, one labeled by a weak classifier and the other by a rule-based approach, with a single supervised dataset. Using these three datasets for training, we perform extensive experiments on both multitasking and prompting strategies. Compared to other systems trained on unsupervised parallel data, models trained on our weak classifier labeled dataset achieve state-of-the-art performance on the ASSET simplification benchmark. Our models also outperform previous work on sentence-level targeting. Finally, we establish how a handful of Large Language Models perform on these tasks under a zero-shot setting.
Alison Chi, Li-Kuang Chen, Yi-Chen Chang, Shu-Hui Lee, Jason S. Chang
Trans. Assoc. Comput. Linguistics5
2019 Extracting Grammatical Error Corrections from Wikipedia Revision History
abstract
This paper describes the process of extracting and filtering Wikipedia revision history as a resource for grammatical error correction (GEC). Edits in Wikipedia revision history vary widely, including grammatical error corrections, information supplements, format amendments, and even vandalism. To extract only GEC-related revisions, we use an automated error annotation toolkit, ERRANT1, and extend it to process large data in parallel efficiently. With error-type analysis, we can then identify GEC-related edits and omit other unrelated edits (i.e., only the correction parts are reserved). The resulting corpus is - to our knowledge - the largest publicly available corpus of parallel possibly erroneous and correct sentences with error type labels.
Jhih-Jie Chen, Yi-Dong Wu, Yu-Chuan Tai, Ching-Yu Helen Yang, Hai-Lun Tu, Jason S. Chang
IEEE BigData6
2019 Word Sense Disambiguation Using Wikipedia Link Graph
abstract
In this paper, we present a observation that a word with some special sense should appear in specific category of articles, and this appearance will “take” the word to its Wiki sense [1] by Wikipedia link graph. This property can help to solve word sense disambiguation problem using Wikipedia.
Hai-Lun Tu, Peichen Ho, Jason S. Chang, Li-Guang Chen
IEEE BigData3
2018 Chinese Spelling Check based on Neural Machine Translation
Chiao-Wen Li, Jhih-Jie Chen, Jason S. Chang
PACLIC3
2018 SmartWrite: Extracting Chinese Lexical Grammar Patterns Using Dependency Parsing
Cheng-Cyuan Peng, Ching-Yu Helen Yang, Jhih-Jie Chen, Jason S. Chang
PACLIC4
2015 WriteAhead2: Mining Lexical Grammar Patterns for Assisted Writing
abstract
This paper describes WriteAhead2, an interactive writing environment that provides lexical and grammatical suggestions for second language learners, and helps them write fluently and avoid common writing errors. The method involves learning phrase templates from dictionary examples, and extracting grammar patterns with example phrases from an academic corpus. At run-time, as the user types word after word, the actions trigger a list after list of suggestions. Each successive list contains grammar patterns and examples, most relevant to the half-baked sentence. WriteAhead2 facilitates steady, timely, and spot-on interactions between learner writers and relevant information for effective assisted writing. Preliminary experiments show that WriteAhead2 has the potential to induce better writing and improve writing skills.
Jim Chang, Jason S. Chang
HLT-NAACL2
2015 Learning Sentential Patterns of Various Rhetoric Moves for Assisted Academic Writing
Jim Chang, Hsiang-Ling Hsu, Joanne Boisson, Hao-Chun Peng, Yu-Hsuan Wu, Jason S. Chang
PACLIC6
2014 Ambiguity Resolution for Vt-N Structures in Chinese
abstract
The syntactic ambiguity of a transitive verb (Vt) followed by a noun (N) has long been a problem in Chinese parsing. In this paper, we propose a classifier to resolve the ambiguity of Vt-N structures. The design of the classifier is based on three important guidelines, namely, adopting linguistically motivated features, using all available resources, and easy in-tegration into a parsing model. The lin-guistically motivated features include semantic relations, context, and morpho-logical structures; and the available re-sources are treebank, thesaurus, affix da-tabase, and large corpora. We also pro-pose two learning approaches that resolve the problem of data sparseness by auto-parsing and extracting relative knowledge from large-scale unlabeled data. Our experiment results show that the Vt-N classifier outperforms the cur-rent PCFG parser. Furthermore, it can be easily and effectively integrated into the PCFG parser and general statistical pars-ing models. Evaluation of the learning approaches indicates that world knowledge facilitates Vt-N disambigua-tion through data selection and error cor-rection. 1
Yu-Ming Hsieh, Jason S. Chang, Keh-Jiann Chen
EMNLP2
2014 TakeTwo: A Word Aligner based on Self Learning
Jim Chang, Jian-Cheng Wu, Jason S. Chang
PACLIC3
2013 Translating Chinese Unknown Words by Automatically Acquired Templates
Ming-Hong Bai, Yu-Ming Hsieh, Keh-Jiann Chen, Jason S. Chang
IJCNLP4
2013 Augmentable Paraphrase Extraction Framework
Mei-hua Chen, Shih-Ting Huang, Jason S. Chang
IJCNLP4
2013 A Computer-Assisted Translation and Writing System
abstract
We introduce a method for learning to predict text and grammatical construction in a computer-assisted translation and writing framework. In our approach, predictions are offered on the fly to help the user make appropriate lexical and grammar choices during the translation of a source text, thus improving translation quality and productivity. The method involves automatically generating general-to-specific word usage summaries (i.e., writing suggestion module), and automatically learning high-confidence word- or phrase-level translation equivalents (i.e., translation suggestion module). At runtime, the source text and its translation prefix entered by the user are broken down into n-grams to generate grammar and translation predictions, which are further combined and ranked via translation and language models. These ranked prediction candidates are iteratively and interactively displayed to the user in a pop-up menu as translation or writing hints. We present a prototype writing assistant, TransAhead , that applies the method to a human-computer collaborative environment. Automatic and human evaluations show that novice translators or language learners substantially benefit from our system in terms of translation performance (i.e., translation accuracy and productivity) and language learning (i.e., collocation usage and grammar). In general, our methodology of inline grammar and text predictions or suggestions has great potential in the field of computer-assisted translation, writing, or even language learning.
Chung-Chi Huang, Mei-hua Chen, Ping-Che Yang, Jason S. Chang
ACM Trans. Asian Lang. Inf. Process.4
2012 TransAhead: A Writing Assistant for CAT and CALL
Chung-Chi Huang, Ping-Che Yang, Mei-hua Chen, Hung-ting Hsieh, Ting-hui Kao, Jason S. Chang
EACL6
2012 TransAhead: A Computer-Assisted Translation and Writing Tool
Chung-Chi Huang, Ping-Che Yang, Keh-Jiann Chen, Jason S. Chang
HLT-NAACL4
2011 Using Sublexical Translations to Handle the OOV Problem in Machine Translation
abstract
We introduce a method for learning to translate out-of-vocabulary (OOV) words. The method focuses on combining sublexical/constituent translations of an OOV to generate its translation candidates. In our approach, wildcard searches are formulated based on our OOV analysis, aimed at maximizing the probability of retrieving OOVs’ sublexical translations from existing resources of Machine Translation (MT) systems. At run-time, translation candidates of the unknown words are generated from their suitable sublexical translations and ranked based on monolingual and bilingual information. We have incorporated the OOV model into a state-of-the-art machine translation system and experimental results show that our model indeed helps to ease the impact of OOVs on translation quality, especially for sentences containing more OOVs (significant improvement).
Chung-Chi Huang, Ho-Ching Yen, Ping-Che Yang, Shih-Ting Huang, Jason S. Chang
ACM Trans. Asian Lang. Inf. Process.5
2010 GRASP: Grammar- and Syntax-based Pattern-Finder for Collocation and Phrase Learning
Mei-hua Chen, Chung-Chi Huang, Shih-Ting Huang, Jason S. Chang
PACLIC4
2009 Acquiring Translation Equivalences of Multiword Expressions by Normalized Correlation Frequencies
Ming-Hong Bai, Jia-Ming You, Keh-Jiann Chen, Jason S. Chang
EMNLP4
2009 Learning Bilingual Linguistic Reordering Model for Statistical Machine Translation
Han-Bin Chen, Jian-Cheng Wu, Jason S. Chang
HLT-NAACL3
2009 WikiSense: Supersense Tagging of Wikipedia Named Entities Based WordNet
Joseph Z. Chang, Richard Tzong-Han Tsai, Jason S. Chang
PACLIC3
2009 Review Classification Using Semantic Features and Run-Time Weighting
Chung-Chi Huang, Meng-chiech Lee, Zhe-nan Lin, Jason S. Chang
PACLIC4
2009 Extending Bilingual WordNet via Hierarchical Word Translation Classification
Tzu-yi Nien, Tsun Ku, Chung-Chi Huang, Mei-hua Chen, Jason S. Chang
PACLIC5
2008 Bilingual Segmentation for Alignment and Translation
Chung-Chi Huang, Wei-Teh Chen, Jason S. Chang
CICLing3
2008 Improving Word Alignment by Adjusting Chinese Word Segmentation
Ming-Hong Bai, Keh-Jiann Chen, Jason S. Chang
IJCNLP3
2007 Learning to Find English to Chinese Transliterations on the Web
Jian-Cheng Wu, Jason S. Chang
EMNLP-CoNLL2
2006 FAST - An Automatic Generation System for Grammar Tests
abstract
This paper introduces a method for the semi-automatic generation of grammar test items by applying Natural Language Processing (NLP) techniques. Based on manually-designed patterns, sentences gathered from the Web are transformed into tests on grammaticality. The method involves representing test writing knowledge as test patterns, acquiring authentic sentences on the Web, and applying generation strategies to transform sentences into items. At runtime, sentences are converted into two types of TOEFL-style question: multiple-choice and error detection. We also describe a prototype system FAST (Free Assessment of Structural Tests). Evaluation on a set of generated questions indicates that the proposed method performs satisfactory quality. Our methodology provides a promising approach and offers significant potential for computer assisted language learning and assessment.
Chia-Yin Chen, Hsien-Chin Liou, Jason S. Chang
ACL3
2006 Computational Analysis of Move Structures in Academic Abstracts
abstract
This paper introduces a method for computational analysis of move structures in abstracts of research articles. In our approach, sentences in a given abstract are analyzed and labeled with a specific move in light of various rhetorical functions. The method involves automatically gathering a large number of abstracts from the Web and building a language model of abstract moves. We also present a prototype concordancer, CARE, which exploits the move-tagged abstracts for digital learning. This system provides a promising approach to Web-based computer-assisted academic writing.
Jien-Chen Wu, Yu-Chia Chang, Hsien-Chin Liou, Jason S. Chang
ACL4
2006 Extraction of transliteration pairs from parallel corpora using a statistical transliteration model
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang
Inf. Sci.2
2006 Alignment of bilingual named entities in parallel corpora using statistical models and multiple knowledge sources
abstract
Named entity (NE) extraction is one of the fundamental tasks in natural language processing (NLP). Although many studies have focused on identifying NEs within monolingual documents, aligning NEs in bilingual documents has not been investigated extensively due to the complexity of the task. In this article we introduce a new approach to aligning bilingual NEs in parallel corpora by incorporating statistical models with multiple knowledge sources. In our approach, we model the process of translating an English NE phrase into a Chinese equivalent using lexical translation/transliteration probabilities for word translation and alignment probabilities for word reordering. The method involves automatically learning phrase alignment and acquiring word translations from a bilingual phrase dictionary and parallel corpora, and automatically discovering transliteration transformations from a training set of name-transliteration pairs. The method also involves language-specific knowledge functions, including handling abbreviations, recognizing Chinese personal names, and expanding acronyms. At runtime, the proposed models are applied to each source NE in a pair of bilingual sentences to generate and evaluate the target NE candidates; the source and target NEs are then aligned based on the computed probabilities. Experimental results demonstrate that the proposed approach, which integrates statistical models with extra knowledge sources, is highly feasible and offers significant improvement in performance compared to our previous work, as well as the traditional approach of IBM Model 4.
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang
ACM Trans. Asian Lang. Inf. Process.2
2005 Learning Source-Target Surface Patterns for Web-based Terminology Translation
Jian-Cheng Wu, Tracy Lin, Jason S. Chang
ACL3
2005 Web-Based Unsupervised Learning for Query Formulation in Question Answering
Yi-Chia Wang, Jian-Cheng Wu, Tyne Liang, Jason S. Chang
IJCNLP4
2004 Bilingual Sentence Alignment Based on Punctuation Statistics and Lexicon
Thomas C. Chuang, Jian-Cheng Wu, Tracy Lin, Wen-Chie Shei, Jason S. Chang
IJCNLP5
2003 A Statistical Approach to Chinese-to-English Back-Transliteration
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang
PACLIC2
2002 An Operator Assisted Call Routing System
Chun-Jen Lee, Jason S. Chang
PACLIC2
1998 Topical Clustering of MRD Senses Based on Information Retrieval Techniques
Jen Nan Chen, Jason S. Chang
Comput. Linguistics2
1997 An Alignment Method for Noisy Parallel Corpora based on Image Processing Techniques
abstract
This paper presents a new approach to bitext correspondence problem (BCP) of noisy bilingual corpora based on image processing (IP) techniques. By using one of several ways of estimating the lexical translation probability (LTP) between pairs of source and target words, we can turn a bitext into a discrete gray-level image. We contend that the BCP, when seen in the light, bears a striking resemblance to the line detection problem in IP. Therefore, BCPs, including sentence and word alignment, can benefit from a wealth of effective, well established IP techniques, including convolution-based filters, texture analysis and Hough transform. This paper describes a new program, PlotAlign that produces a word-level bitext map for noisy or non-literal bitext, based on these techniques.
Jason S. Chang, Mathis H. M. Chen
ACL1
1997 A Class-based Approach to Word Alignment
Sue J. Ker, Jason S. Chang
Comput. Linguistics2