VLDB 2026 Research / reviewers in the wild / expert
Xiaochang Peng
dblp:153/9516
· DBLP profile ↗
11ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Information extraction and text analysis · 38% Trustworthy machine learning · 36% Knowledge representation and reasoning · 11% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 62% Logic in computer science · 24% Algorithms and data structures · 14% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › semantic parsing
abstract meaning representation parsing |
0.7 | 2 | 2018 | Sequence-to-sequence Models for Cache Transition Systems · ACL (1) 2018 AMR Parsing With Cache Transition Systems · AAAI 2018 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.7 | 2 | 2018 | Sequence-to-sequence Models for Cache Transition Systems · ACL (1) 2018 AMR Parsing With Cache Transition Systems · AAAI 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation |
0.6 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction |
0.6 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
transition-based parsing |
0.3 | 1 | 2018 | Sequence-to-sequence Models for Cache Transition Systems · ACL (1) 2018 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
AMR-to-text generation |
0.2 | 1 | 2016 | AMR-to-text generation as a Traveling Salesman Problem · EMNLP 2016 |
Mathematical optimization › combinatorial optimization › vehicle routing
traveling salesman problem |
0.2 | 1 | 2016 | AMR-to-text generation as a Traveling Salesman Problem · EMNLP 2016 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.2 | 1 | 2014 | Type-based MCMC for Sampling Tree Fragments from Forests · EMNLP 2014 |
Natural language and speech › Machine translation › syntax-based machine translation
synchronous context-free grammar |
0.2 | 1 | 2014 | Type-based MCMC for Sampling Tree Fragments from Forests · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
text classification |
0.2 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
abstract meaning representation |
0.1 | 1 | 2018 | Sequence-to-sequence Models for Cache Transition Systems · ACL (1) 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
semantic graph |
0.1 | 1 | 2018 | Sequence-to-sequence Models for Cache Transition Systems · ACL (1) 2018 |
Logic in computer science
transition systems |
0.1 | 1 | 2018 | AMR Parsing With Cache Transition Systems · AAAI 2018 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
graph-to-text generation |
0.1 | 1 | 2016 | AMR-to-text generation as a Traveling Salesman Problem · EMNLP 2016 |
Methods — techniques the papers use, named apart from their topics
transition-based parsing · 0.7cache transition system · 0.7select-predict pipeline · 0.6joint training · 0.6attribution algorithm · 0.6maximum entropy classifier · 0.5graph partitioning · 0.5sequence-to-sequence model · 0.3hard attention · 0.3feature embedding · 0.3traveling salesman problem solver · 0.2word alignment · 0.2type-based MCMC · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionabstractAn extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without compromising the LM’s (i.e., task model’s) task performance. Although attribution algorithms and select-predict pipelines are commonly used in rationale extraction, they both rely on certain heuristics that hinder them from satisfying all three desiderata. In light of this, we propose UNIREX, a flexible learning framework which generalizes rationale extractor optimization as follows: (1) specify architecture for a learned rationale extractor; (2) select explainability objectives (\ie faithfulness and plausibility criteria); and (3) jointly train the task model and rationale extractor on the task using selected objectives. UNIREX enables replacing prior works’ heuristic design choices with a generic learned rationale extractor in (1) and optimizing it for all three desiderata in (2)-(3). To facilitate comparison between methods w.r.t. multiple desiderata, we introduce the Normalized Relative Gain (NRG) metric. On five English text classification datasets, our best UNIREX configuration outperforms baselines by an average of 32.9% NRG. Plus, UNIREX rationale extractors’ faithfulness can even generalize to unseen datasets and tasks. Aaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 0005, Shaoliang Nie, Xiaochang Peng, Xiang Ren 0001, Hamed Firooz |
ICML | 6 |
| 2019 | Ordered Tree Decomposition for HRG Rule ExtractionabstractWe present algorithms for extracting Hyperedge Replacement Grammar (HRG) rules from a graph along with a vertex order. Our algorithms are based on finding a tree decomposition of smallest width, relative to the vertex order, and then extracting one rule for each node in this structure. The assumption of a fixed order for the vertices of the input graph makes it possible to solve the problem in polynomial time, in contrast to the fact that the problem of finding optimal tree decompositions for a graph is NP-hard. We also present polynomial-time algorithms for parsing based on our HRGs, where the input is a vertex sequence and the output is a graph structure. The intended application of our algorithms is grammar extraction and parsing for semantic representation of natural language. We apply our algorithms to data annotated with Abstract Meaning Representations and report on the characteristics of the resulting grammars. Daniel Gildea, Giorgio Satta, Xiaochang Peng |
Comput. Linguistics | 3 |
| 2019 | Neural Models of Text Normalization for Speech ApplicationsabstractMachine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 King Ave. For this task, state-of-the-art industrial systems depend heavily on hand-written language-specific grammars. We propose neural network models that treat text normalization for TTS as a sequence-to-sequence problem, in which the input is a text token in context, and the output is the verbalization of that token. We find that the most effective model, in accuracy and efficiency, is one where the sentential context is computed once and the results of that computation are combined with the computation of each token in sequence to compute the verbalization. This model allows for a great deal of flexibility in terms of representing the context, and also allows us to integrate tagging and segmentation into the process. These models perform very well overall, but occasionally they will predict wildly inappropriate verbalizations, such as reading 3 cm as three kilometers. Although rare, such verbalizations are a major issue for TTS applications. We thus use finite-state covering grammars to guide the neural models, either during training and decoding, or just during decoding, away from such “unrecoverable” errors. Such grammars can largely be learned from data. Hao Zhang 0010, Richard Sproat, Axel H. Ng, Felix Stahlberg, Xiaochang Peng, Kyle Gorman, Brian Roark |
Comput. Linguistics | 5 |
| 2018 | AMR Parsing With Cache Transition SystemsabstractIn this paper, we present a transition system that generalizes transition-based dependency parsing techniques to generateAMR graphs rather than tree structures. In addition to a buffer and a stack, we use a fixed-size cache, and allow the system to build arcs to any vertices present in the cache at the same time. The size of the cache provides a parameter that can trade off between the complexity of the graphs that can be built and the ease of predicting actions during parsing. Our results show that a cache transition system can cover almost all AMR graphs with a small cache size, and our end-to-end system achieves competitive results in comparison with other transition-based approaches for AMR parsing. Xiaochang Peng, Daniel Gildea, Giorgio Satta |
AAAI | 1 |
| 2018 | Sequence-to-sequence Models for Cache Transition SystemsabstractIn this paper, we present a sequenceto-sequence based approach for mapping natural language sentences to AMR semantic graphs.We transform the sequence to graph mapping problem to a word sequence to transition action sequence problem using a special transition system called a cache transition system.To address the sparsity issue of neural AMR parsing, we feed feature embeddings from the transition state to provide relevant local information for each decoder state.We present a monotonic hard attention model for the transition framework to handle the strictly left-to-right alignment between each transition state and the current buffer input focus.We evaluate our neural transition model on the AMR parsing task, and our parser outperforms other sequence-to-sequence approaches and achieves competitive results in comparison with the best-performing models. 1 Xiaochang Peng, Linfeng Song, Daniel Gildea, Giorgio Satta |
ACL (1) | 1 |
| 2018 | Cache Transition Systems for Graph ParsingabstractMotivated by the task of semantic parsing, we describe a transition system that generalizes standard transition-based dependency parsing techniques to generate a graph rather than a tree. Our system includes a cache with fixed size m, and we characterize the relationship between the parameter m and the class of graphs that can be produced through the graph-theoretic concept of tree decomposition. We find empirically that small cache sizes cover a high percentage of sentences in existing semantic corpora. Daniel Gildea, Giorgio Satta, Xiaochang Peng |
Comput. Linguistics | 3 |
| 2017 | Addressing the Data Sparsity Issue in Neural AMR ParsingabstractNeural attention models have achieved great success in different NLP tasks.However, they have not fulfilled their promise on the AMR parsing task due to the data sparsity issue.In this paper, we describe a sequence-to-sequence model for AMR parsing and present different ways to tackle the data sparsity problem.We show that our methods achieve significant improvement over a baseline neural attention model and our results are also competitive against state-of-the-art systems that do not use extra linguistic resources. Xiaochang Peng, Daniel Gildea, Nianwen Xue |
EACL (1) | 1 |
| 2016 | AMR-to-text generation as a Traveling Salesman ProblemabstractThe task of AMR-to-text generation is to generate grammatical text that sustains the semantic meaning for a given AMR graph.We attack the task by first partitioning the AMR graph into smaller fragments, and then generating the translation for each fragment, before finally deciding the order by solving an asymmetric generalized traveling salesman problem (AGTSP).A Maximum Entropy classifier is trained to estimate the traveling costs, and a TSP solver is used to find the optimized solution.The final model reports a BLEU score of 22.44 on the SemEval-2016 Task8 dataset. Linfeng Song, Yue Zhang 0004, Xiaochang Peng, Zhiguo Wang 0006, Daniel Gildea |
EMNLP | 3 |
| 2015 | A Synchronous Hyperedge Replacement Grammar based approach for AMR parsingabstractThis paper presents a synchronous-graphgrammar-based approach for string-to-AMR parsing.We apply Markov Chain Monte Carlo (MCMC) algorithms to learn Synchronous Hyperedge Replacement Grammar (SHRG) rules from a forest that represents likely derivations consistent with a fixed string-to-graph alignment.We make an analogy of string-to-AMR parsing to the task of phrase-based machine translation and come up with an efficient algorithm to learn graph grammars from string-graph pairs.We propose an effective approximation strategy to resolve the complexity issue of graph compositions.We also show some useful strategies to overcome existing problems in an SHRG-based parser and present preliminary results of a graph-grammar-based approach. Xiaochang Peng, Linfeng Song, Daniel Gildea |
CoNLL | 1 |
| 2014 | Type-based MCMC for Sampling Tree Fragments from ForestsabstractThis paper applies type-based Markov Chain Monte Carlo (MCMC) algorithms to the problem of learning Synchronous Context-Free Grammar (SCFG) rules from a forest that represents all possible rules consistent with a fixed word align-ment. While type-based MCMC has been shown to be effective in a number of NLP applications, our setting, where the tree structure of the sentence is itself a hid-den variable, presents a number of chal-lenges to type-based inference. We de-scribe methods for defining variable types and efficiently indexing variables in or-der to overcome these challenges. These methods lead to improvements in both log likelihood and BLEU score in our experi-ments. Xiaochang Peng, Daniel Gildea |
EMNLP | 1 |
| 2013 | Capturing Long-distance Dependencies in Sequence Models: A Case Study of Chinese Part-of-speech Tagging
Weiwei Sun 0007, Xiaochang Peng, Xiaojun Wan 0001 |
IJCNLP | 2 |