VLDB 2026 Research / reviewers in the wild / expert
Wei-Yun Ma
dblp:72/4128
· DBLP profile ↗
21ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0002-4094-7464ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Information extraction and text analysis · 52% Deep learning architectures and training · 23% Generative modeling · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
named entity recognition |
1.1 | 3 | 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER · AAAI 2020 GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation Extraction · ACL (1) 2019 Leveraging Linguistic Structures for Named Entity Recognition with Bidirectional Recursive Neural Networks · EMNLP 2017 |
Knowledge graphs
knowledge graph construction |
1.0 | 1 | 2026 | Efficient LLM Adaptation for Opinion Knowledge Graph Construction: Lessons from the Telecom Industry · SIGIR 2026 |
Natural language and speech › Information extraction and text analysis
text classification |
1.0 | 2 | 2025 | Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification · ICLR 2025 Speed Reading: Learning to Read ForBackward via Shuttle · EMNLP 2018 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification · ICLR 2025 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.8 | 2 | 2020 | Relation Extraction Exploiting Full Dependency Forests · AAAI 2020 GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation Extraction · ACL (1) 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.8 | 2 | 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER · AAAI 2020 Speed Reading: Learning to Read ForBackward via Shuttle · EMNLP 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network › bidirectional recurrent network
bidirectional LSTM |
0.4 | 1 | 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER · AAAI 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
dependency-based relation extraction |
0.4 | 1 | 2020 | Relation Extraction Exploiting Full Dependency Forests · AAAI 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.4 | 1 | 2019 | GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation Extraction · ACL (1) 2019 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.3 | 1 | 2018 | Speed Reading: Learning to Read ForBackward via Shuttle · EMNLP 2018 |
Natural language and speech › Machine translation
paraphrase-based translation |
0.2 | 1 | 2015 | System Combination for Machine Translation through Paraphrasing · EMNLP 2015 |
Natural language and speech › Machine translation
system combination |
0.2 | 1 | 2015 | System Combination for Machine Translation through Paraphrasing · EMNLP 2015 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER · AAAI 2020 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation |
0.1 | 1 | 2020 | Relation Extraction Exploiting Full Dependency Forests · AAAI 2020 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.1 | 1 | 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER · AAAI 2020 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.1 | 1 | 2019 | GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation Extraction · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.1 | 1 | 2018 | Speed Reading: Learning to Read ForBackward via Shuttle · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual information extraction |
0.1 | 1 | 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task · ACL/IJCNLP 2009 |
Methods — techniques the papers use, named apart from their topics
large language model adaptation · 1.0weighted loss · 0.9data weighting · 0.9dependency parsing · 0.8self-attention · 0.4differentiable parsing · 0.4cross-BiLSTM · 0.4CRF · 0.4graph convolutional network · 0.4backward reading · 0.3LSTM · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient LLM Adaptation for Opinion Knowledge Graph Construction: Lessons from the Telecom Industry
Nai-Chi Yang, Yu-Ming Hsieh, Wei-Yun Ma, Kuo-Wei Chang |
SIGIR | 3 |
| 2025 | Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationabstractSynthetic data augmentation via Large Language Models (LLMs) allows researchers to leverage additional training data, thus enhancing the performance of downstream tasks, especially when real-world data is scarce. However, the generated data can deviate from the real-world data, and this misalignment can bring about deficient results while applying the trained model to applications. Therefore, we proposed efficient weighted-loss approaches to align synthetic data with real-world distribution by emphasizing high-quality and diversified data generated by LLMs using merely a tiny amount of real-world data. We empirically assessed the effectiveness of our methods on multiple text classification tasks, and the results showed that leveraging our approaches on a BERT-level model robustly outperformed standard cross-entropy and other data weighting approaches, providing potential solutions to effectively leveraging synthetic data from any suitable data generator. Hsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma, Pu-Jen Cheng |
ICLR | 4 |
| 2024 | Automatic Construction of a Chinese Review Dataset for Aspect Sentiment Triplet Extraction via Iterative Weak SupervisionabstractAspect Sentiment Triplet Extraction (ASTE), introduced in 2020, is a task that involves the extraction of three key elements: target aspects, descriptive opinion spans, and their corresponding sentiment polarity. This process, however, faces a significant hurdle, particularly when applied to Chinese languages, due to the lack of sufficient datasets for model training, largely attributable to the arduous manual labeling process. To address this issue, we present an innovative framework that facilitates the automatic construction of ASTE via Iterative Weak Supervision, negating the need for manual labeling, aided by a discriminator to weed out subpar samples. The objective is to successively improve the quality of this raw data and generate supplementary data. The effectiveness of our approach is underscored by our results, which include the creation of a substantial Chinese review dataset. This dataset encompasses over 60,000 Google restaurant reviews in Chinese and features more than 200,000 extracted triplets. Moreover, we have also established a robust baseline model by leveraging a novel method of weak supervision. Both our dataset and model are openly accessible to the public. Chia-Wen Lu, Ching-Wen Yang, Wei-Yun Ma |
LREC/COLING | 3 |
| 2024 | Generating Attractive and Authentic Copywriting from Customer ReviewsabstractThe goal of product copywriting is to capture the interest of potential buyers by emphasizing the features of products through text descriptions.As e-commerce platforms offer a wide range of services, it's becoming essential to dynamically adjust the styles of these autogenerated descriptions.Typical approaches to copywriting generation often rely solely on specified product attributes, which may result in dull and repetitive content.To tackle this issue, we propose to generate copywriting based on customer reviews, as they provide firsthand practical experiences with products, offering a richer source of information than just product attributes.We have developed a sequenceto-sequence framework, enhanced with reinforcement learning, to produce copywriting that is attractive, authentic, and rich in information.Our framework outperforms all existing baseline and zero-shot large language models, including LLaMA-2-chat-7B and GPT-3.5 1 , in terms of both attractiveness and faithfulness.Furthermore, this work features the use of LLMs for aspect-based summaries collection and argument allure assessment.Experiments demonstrate the effectiveness of using LLMs for marketing domain corpus construction.The code and the dataset is publicly available at: https://github.com/YuXiangLin1234/ Copywriting-Generation. Yu-Xiang Lin, Wei-Yun Ma |
NAACL-HLT | 2 |
| 2020 | Relation Extraction Exploiting Full Dependency ForestsabstractDependency syntax has long been recognized as a crucial source of features for relation extraction. Previous work considers 1-best trees produced by a parser during preprocessing. However, error propagation from the out-of-domain parser may impact the relation extraction performance. We propose to leverage full dependency forests for this task, where a full dependency forest encodes all possible trees. Such representations of full dependency forests provide a differentiable connection between a parser and a relation extraction model, and thus we are also able to study adjusting the parser parameters based on end-task loss. Experiments on three datasets show that full dependency forests and parser adjustment give significant improvements over carefully designed baselines, showing state-of-the-art or competitive performances on biomedical or newswire benchmarks. Lifeng Jin, Linfeng Song, Yue Zhang 0004, Kun Xu 0005, Wei-Yun Ma, Dong Yu 0001 |
AAAI | 5 |
| 2020 | Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NERabstractBiLSTM has been prevalently used as a core module for NER in a sequence-labeling setup. State-of-the-art approaches use BiLSTM with additional resources such as gazetteers, language-modeling, or multi-task supervision to further improve NER. This paper instead takes a step back and focuses on analyzing problems of BiLSTM itself and how exactly self-attention can bring improvements. We formally show the limitation of (CRF-)BiLSTM in modeling cross-context patterns for each word – the XOR limitation. Then, we show that two types of simple cross-structures – self-attention and Cross-BiLSTM – can effectively remedy the problem. We test the practical impacts of the deficiency on real-world NER datasets, OntoNotes 5.0 and WNUT 2017, with clear and consistent improvements over the baseline, up to 8.7% on some of the multi-token entity mentions. We give in-depth analyses of the improvements across several aspects of NER, especially the identification of multi-token mentions. This study should lay a sound foundation for future improvements on sequence-labeling NER1. Peng-Hsuan Li, Tsu-Jui Fu, Wei-Yun Ma |
AAAI | 3 |
| 2020 | CA-EHN: Commonsense Analogy from E-HowNetabstractEmbedding commonsense knowledge is crucial for end-to-end models to generalize inference beyond training corpora. However, existing word analogy datasets have tended to be handcrafted, involving permutations of hundreds of words with only dozens of pre-defined relations, mostly morphological relations and named entities. In this work, we model commonsense knowledge down to word-level analogical reasoning by leveraging E-HowNet, an ontology that annotates 88K Chinese words with their structured sense definitions and English translations. We present CA-EHN, the first commonsense word analogy dataset containing 90,505 analogies covering 5,656 words and 763 relations. Experiments show that CA-EHN stands out as a great indicator of how well word representations embed commonsense knowledge. The dataset is publicly available at https://github.com/ckiplab/CA-EHN. Peng-Hsuan Li, Tsan-Yu Yang, Wei-Yun Ma |
LREC | 3 |
| 2020 | Headword-Oriented Entity Linking: A Special Entity Linking Task with Dataset and BaselineabstractIn this paper, we design headword-oriented entity linking (HEL), a specialized entity linking problem in which only the headwords of the entities are to be linked to knowledge bases; mention scopes of the entities do not need to be identified in the problem setting. This special task is motivated by the fact that in many articles referring to specific products, the complete full product names are rarely written; instead, they are often abbreviated to shorter, irregular versions or even just to their headwords, which are usually their product types, such as “stick” or “mask” in a cosmetic context. To fully design the special task, we construct a labeled cosmetic corpus as a public benchmark for this problem, and propose a product embedding model to address the task, where each product corresponds to a dense representation to encode the different information on products and their context jointly. Besides, to increase training data, we propose a special transfer learning framework in which distant supervision with heuristic patterns is first utilized, followed by supervised learning using a small amount of manually labeled data. The experimental results show that our model provides a strong benchmark performance on the special task. Mu Yang, Chi-Yen Chen, Yi-Hui Lee, Qian-hui Zeng, Wei-Yun Ma, Chen-Yang Shih, Wei-Jhih Chen |
LREC | 5 |
| 2020 | Semantic Guidance of Dialogue Generation with Reinforcement LearningabstractNeural encoder-decoder models have shown promising performance for human-computer dialogue systems over the past few years.However, due to the maximum-likelihood objective for the decoder, the generated responses are often universal and safe to the point that they lack meaningful information and are no longer relevant to the post.To address this, in this paper, we propose semantic guidance using reinforcement learning to ensure that the generated responses indeed include the given or predicted semantics and that these semantics do not appear repeatedly in the response.Synsets, which comprise sets of manually defined synonyms, are used as the form of assigned semantics.For a given/assigned/predicted synset, only one of its synonyms should appear in the generated response; this constitutes a simple but effective semantic-control mechanism.We conduct both quantitative and qualitative evaluations, which show that the generated responses are not only higher-quality but also reflect the assigned semantic controls. Cheng-Hsun Hsueh, Wei-Yun Ma |
SIGdial | 2 |
| 2019 | GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation ExtractionabstractIn this paper, we present GraphRel, an end-to-end relation extraction model which uses graph convolutional networks (GCNs) to jointly learn named entities and relations.In contrast to previous baselines, we consider the interaction between named entities and relations via a relation-weighted GCN to better extract relations.Linear and dependency structures are both used to extract both sequential and regional features of the text, and a complete word graph is further utilized to extract implicit features among all word pairs of the text.With the graph-based approach, the prediction for overlapping relations is substantially improved over previous sequential approaches.We evaluate GraphRel on two public datasets: NYT and WebNLG.Results show that GraphRel maintains high precision while increasing recall substantially.Also, GraphRel outperforms previous work by 3.2% and 5.8% (F1 score), achieving a new state-of-the-art for relation extraction. Tsu-Jui Fu, Peng-Hsuan Li, Wei-Yun Ma |
ACL (1) | 3 |
| 2018 | Speed Reading: Learning to Read ForBackward via ShuttleabstractWe present LSTM-Shuttle, which applies human speed reading techniques to natural language processing tasks for accurate and efficient comprehension.In contrast to previous work, LSTM-Shuttle not only reads shuttling forward but also goes back.Shuttling forward enables high efficiency, and going backward gives the model a chance to recover lost information, ensuring better prediction.We evaluate LSTM-Shuttle on sentiment analysis, news classification, and cloze on IMDB, Rotten Tomatoes, AG, and Children's Book Test datasets.We show that LSTM-Shuttle predicts both better and more quickly.To demonstrate how LSTM-Shuttle actually behaves, we also analyze the shuttling operation and present a case study. Tsu-Jui Fu, Wei-Yun Ma |
EMNLP | 2 |
| 2018 | Word Embedding Evaluation Datasets and Wikipedia Title Embedding for Chinese
Chi-Yen Chen, Wei-Yun Ma |
LREC | 2 |
| 2018 | Extended HowNet 2.0 - An Entity-Relation Common-Sense Representation Model
Wei-Yun Ma, Yueh-Yin Shih |
LREC | 1 |
| 2017 | Leveraging Linguistic Structures for Named Entity Recognition with Bidirectional Recursive Neural NetworksabstractIn this paper, we utilize the linguistic structures of texts to improve named entity recognition by BRNN-CNN, a special bidirectional recursive network attached with a convolutional network.Motivated by the observation that named entities are highly related to linguistic constituents, we propose a constituent-based BRNN-CNN for named entity recognition.In contrast to classical sequential labeling methods, the system first identifies which text chunks are possible named entities by whether they are linguistic constituents.Then it classifies these chunks with a constituency tree structure by recursively propagating syntactic and semantic information to each constituent node.This method surpasses current state-of-the-art on OntoNotes 5.0 with automatically generated parses. Peng-Hsuan Li, Ruo-Ping Dong, Yu-Siang Wang, Ju-Chieh Chou, Wei-Yun Ma |
EMNLP | 5 |
| 2015 | System Combination for Machine Translation through ParaphrasingabstractIn this paper, we propose a paraphrasing model to address the task of system combination for machine translation.We dynamically learn hierarchical paraphrases from target hypotheses and form a synchronous context-free grammar to guide a series of transformations of target hypotheses into fused translations.The model is able to exploit phrasal and structural system-weighted consensus and also to utilize existing information about word ordering present in the target hypotheses.In addition, to consider a diverse set of plausible fused translations, we develop a hybrid combination architecture, where we paraphrase every target hypothesis using different fusing techniques to obtain fused translations for each target, and then make the final selection among all fused translations.Our experimental results show that our approach can achieve a significant improvement over combination baselines.i h EP to denote i h E attached with related word positions, use i h e to denote a phrase within i h E , and use i h ep to denote i h e attached with related word positions.For instance, If i h E is "you buy the book", then i h EP would be "you 1 buy 2 the 3 book 4 ".If i h eis "the book", then i h ep is "the 3 book 4 ".For a given sentence i, a MT system h and a MT system k, we use a SCFG denoted by i Wei-Yun Ma, Kathy McKeown |
EMNLP | 1 |
| 2013 | Using a Supertagged Dependency Language Model to Select a Good Translation in System Combination
Wei-Yun Ma, Kathy McKeown |
HLT-NAACL | 1 |
| 2011 | System Combination for Machine Translation Based on Text-to-Text Generation
Wei-Yun Ma, Kathy McKeown |
MTSummit | 1 |
| 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman |
ACL/IJCNLP | 9 |
| 2006 | Uniform and Effective Tagging of a Heterogeneous Giga-word Corpus
Wei-Yun Ma, Chu-Ren Huang |
LREC | 1 |
| 2006 | Knowledge-Rich Approach to Automatic Grammatical Information Acquisition: Enriching Chinese Sketch Engine with a Lexical Grammar
Chu-Ren Huang, Wei-Yun Ma, Yi-Ching Wu, Chih-Ming Chiu |
PACLIC | 2 |
| 2002 | Unknown Word Extraction for Chinese Documents
Keh-Jiann Chen, Wei-Yun Ma |
COLING | 2 |