VLDB 2026 Research / reviewers in the wild / expert
Kewei Tu
dblp:22/918
· DBLP profile ↗
82ranked-venue papers
8as first author
40since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 5 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GiLT: Augmenting Transformer Language Models with Dependency GraphsabstractAugmenting Transformers with linguistic structures effectively enhances the syntactic generalization performance of language models.Previous work in this direction focuses on syntactic tree structures of languages, in particular constituency tree structures.We propose Graph-Infused Layers Transformer Language Model (GiLT) which leverages dependency graphs for augmenting Transformer language models.Unlike most previous work, GiLT does not insert extra structural tokens in language modeling; instead, it injects structural information into language modeling by modulating attention weights in the Transformer with features extracted from the dependency graph that is incrementally constructed along with token prediction.In our experiments, GiLT with semantic dependency graphs achieves better syntactic generalization while maintaining competitive perplexity in comparison with Transformer language model baselines.In addition, GiLT can be finetuned from a pretrained language model to achieve improved downstream task performance.Our code is released at https://github.com/cookie-pie-oops/GiLT-LM. Yida Zhao, Chuyan Zhou, Kewei Tu |
ACL (1) | 4 |
| 2025 | Look Both Ways and No Sink: Converting LLMs into Text Encoders without TrainingabstractRecent advancements have demonstrated the advantage of converting pretrained large language models into powerful text encoders by enabling bidirectional attention in transformer layers. However, existing methods often require extensive training on large-scale datasets, posing challenges in low-resource, domain-specific scenarios. In this work, we show that a pretrained large language model can be converted into a strong text encoder without additional training. We first conduct a comprehensive empirical study to investigate different conversion strategies and identify the impact of the attention sink phenomenon on the performance of converted encoder models. Based on our findings, we propose a novel approach that enables bidirectional attention and suppresses the attention sink phenomenon, resulting in superior performance. Extensive experiments on multiple domains demonstrate the effectiveness of our approach. Our work provides new insights into the training-free conversion of text encoders in low-resource scenarios and contributes to the advancement of domain-specific text representation generation. Our code is available at https://github.com/bigai-nlco/Look-Both-Ways-and-No-Sink. Ziyong Lin, Haoyi Wu, Kewei Tu, Zilong Zheng, Zixia Jia |
ACL (1) | 4 |
| 2025 | A Systematic Study of Compositional Syntactic Transformer Language ModelsabstractSyntactic language models (SLMs) enhance Transformers by incorporating syntactic biases through the modeling of linearized syntactic parse trees alongside surface sentences.This paper focuses on compositional SLMs that are based on constituency parse trees and contain explicit bottom-up composition of constituent representations.We identify key aspects of design choices in existing compositional SLMs and propose a unified framework encompassing both existing models and novel variants.We conduct a comprehensive empirical evaluation of all the variants in our framework across language modeling, syntactic generalization, summarization, dialogue, and inference efficiency.Based on the experimental results, we make multiple recommendations on the design of compositional SLMs.Our code is released at https://github.com/ zhaoyd1/compositional_SLMs. Yida Zhao, Hao Xve, Kewei Tu |
ACL (1) | 4 |
| 2025 | Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based InferenceabstractDespite the advancements made in Vision Large Language Models (VLLMs), like text Large Language Models (LLMs), they have limitations in addressing questions that require real-time information or are knowledgeintensive.Indiscriminately adopting Retrieval Augmented Generation (RAG) techniques is an effective yet expensive way to enable models to answer queries beyond their knowledge scopes.To mitigate the dependence on retrieval and simultaneously maintain, or even improve, the performance benefits provided by retrieval, we propose a method to detect the knowledge boundary of VLLMs, allowing for more efficient use of techniques like RAG.Specifically, we propose a method with two variants that finetune a VLLM on an automatically constructed dataset for boundary identification.Experimental results on various types of Visual Question Answering datasets show that our method successfully depicts a VLLM's knowledge boundary, based on which we are able to reduce indiscriminate retrieval while maintaining or improving the performance.In addition, we show that the knowledge boundary identified by our method for one VLLM can be used as a surrogate boundary for other VLLMs.Code will be released at https://github.com/Chord-Che n-30/VLLM-KnowledgeBoundary Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinyu Geng, Pengjun Xie, Fei Huang 0002, Kewei Tu |
EMNLP | 8 |
| 2025 | Parallel Continuous Chain-of-Thought with Jacobi IterationabstractContinuous chain-of-thought has been shown to be effective in saving reasoning tokens for large language models.By reasoning with continuous latent thought tokens, continuous CoT is able to perform implicit reasoning in a compact manner.However, the sequential dependencies between latent thought tokens spoil parallel training, leading to long training time.In this paper, we propose Parallel Continuous Chain-of-Thought (PCCoT), which performs Jacobi iteration on the latent thought tokens, updating them iteratively in parallel instead of sequentially and thus improving both training and inference efficiency of continuous CoT.Experiments demonstrate that by choosing the proper number of iterations, we are able to achieve comparable or even better performance while saving nearly 50% of the training and inference time.Moreover, PC-CoT shows better stability and robustness in the training process.Our code is available at https://github.com/whyNLP/PCCoT. Haoyi Wu, Zhihao Teng, Kewei Tu |
EMNLP | 3 |
| 2025 | EvolveSearch: An Iterative Self-Evolving Search AgentabstractDing-Chu Zhang, Yida Zhao, Jialong Wu, Liwen Zhang, Baixuan Li, Wenbiao Yin, Yong Jiang, Yu-Feng Li, Kewei Tu, Pengjun Xie, Fei Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dingchu Zhang, Yida Zhao, Jialong Wu 0007, Baixuan Li, Wenbiao Yin, Yong Jiang 0005, Kewei Tu, Pengjun Xie, Fei Huang 0002 |
EMNLP | 9 |
| 2025 | Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language ModelingabstractDespite the success of Transformers, handling longer contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention, which often requires post-training with a larger attention window, significantly increasing computational and memory costs. In this paper, we propose a novel attention mechanism based on dynamic context, Grouped Cross Attention (GCA), which can generalize to 1000 $\times$ the pre-training context length while maintaining the ability to access distant information with a constant attention window size. For a given input sequence, we split it into chunks and use each chunk to retrieve top-$k$ relevant past chunks for subsequent text generation.
Specifically, unlike most previous works that use an off-the-shelf retriever, our key innovation allows the retriever to learn how to retrieve past chunks that better minimize the auto-regressive loss of subsequent tokens in an end-to-end manner, which adapts better to causal language models.
Such a mechanism accommodates retrieved chunks with a fixed-size attention window to achieve long-range information access, significantly reducing computational and memory costs during training and inference.
Experiments show that GCA-based models achieve near-perfect accuracy in passkey retrieval for 16M context lengths, which is $1000 \times$ the training length. Zhihao Teng, Kewei Tu |
ICML | 5 |
| 2025 | Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory AccessabstractA key advantage of Recurrent Neural Networks (RNNs) over Transformers is their linear computational and space complexity enables faster training and inference for long sequences. However, RNNs are fundamentally unable to randomly access historical context, and simply integrating attention mechanisms may undermine their efficiency advantages.
To overcome this limitation, we propose \textbf{H}ierarchical \textbf{S}parse \textbf{A}ttention (HSA), a novel attention mechanism that enhances RNNs with long-range random access flexibility while preserving their merits in efficiency and length generalization. HSA divides inputs into chunks, selecting the top-$k$ chunks and hierarchically aggregates information.
The core innovation lies in learning token-to-chunk relevance based on fine-grained token-level information inside each chunk. This approach enhances the precision of chunk selection across both in-domain and out-of-domain context lengths.
To make HSA efficient, we further introduce a hardware-aligned kernel design.
By combining HSA with Mamba, we introduce RAMba, which achieves perfect accuracy in passkey retrieval across 64 million contexts despite pre-training on only 4K-length contexts, and significant improvements on various downstream tasks, with nearly constant memory footprint. These results show RAMba's huge potential in long-context modeling. Jiaqi Leng 0003, Kewei Tu |
NeurIPS | 4 |
| 2024 | Frame Semantic Role Labeling Using Arbitrary-Order Conditional Random FieldsabstractThis paper presents an approach to frame semantic role labeling (FSRL), a task in natural language processing that identifies semantic roles within a text following the theory of frame semantics. Unlike previous approaches which do not adequately model correlations and interactions amongst arguments, we propose arbitrary-order conditional random fields (CRFs) that are capable of modeling full interaction amongst an arbitrary number of arguments of a given predicate. To achieve tractable representation and inference, we apply canonical polyadic decomposition to the arbitrary-order factor in our proposed CRF and utilize mean-field variational inference for approximate inference. We further unfold our iterative inference procedure into a recurrent neural network that is connected to our neural encoder and scorer, enabling end-to-end training and inference. Finally, we also improve our model with several techniques such as span-based scoring and decoding. Our experiments show that our approach achieves state-of-the-art performance in FSRL. Chaoyi Ai, Kewei Tu |
AAAI | 2 |
| 2024 | SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingabstractLarge language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT. Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005 |
AAAI | 10 |
| 2024 | Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleabstractA syntactic language model (SLM) incrementally generates a sentence with its syntactic tree in a left-to-right manner.We present Generative Pretrained Structured Transformers (GPST), an unsupervised SLM at scale capable of being pre-trained from scratch on raw texts with high parallelism.GPST circumvents the limitations of previous SLMs such as relying on gold trees and sequential training.It consists of two components, a usual SLM supervised by a uni-directional language modeling loss, and an additional composition model, which induces syntactic parse trees and computes constituent representations, supervised by a bi-directional language modeling loss.We propose a representation surrogate to enable joint parallel training of the two models in a hard-EM fashion.We pre-train GPST on OpenWebText, a corpus with 9 billion tokens, and demonstrate the superiority of GPST over GPT-2 with a comparable size in numerous tasks covering both language understanding and language generation.Meanwhile, GPST also significantly outperforms existing unsupervised SLMs on left-to-right grammar induction, while holding a substantial acceleration on training.1 * Equal contribution, see appendix A.7 for details. Pengyu Ji, Qingyang Zhu, Kewei Tu |
ACL (1) | 5 |
| 2024 | Layer-Condensed KV Cache for Efficient Inference of Large Language ModelsabstractHuge memory consumption has been a major bottleneck for deploying high-throughput large language models in real-world applications.In addition to the large number of parameters, the key-value (KV) cache for the attention mechanism in the transformer architecture consumes a significant amount of memory, especially when the number of layers is large for deep language models.In this paper, we propose a novel method that only computes and caches the KVs of a small number of layers, thus significantly saving memory consumption and improving inference throughput.Our experiments on large language models show that our method achieves up to 26× higher throughput than standard transformers and competitive performance in language modeling and downstream tasks.In addition, our method is orthogonal to existing transformer memory-saving techniques, so it is straightforward to integrate them with our model, achieving further improvement in inference efficiency.Our code is available at https://github.com/whyNLP/LCKV. Haoyi Wu, Kewei Tu |
ACL (1) | 2 |
| 2024 | Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language ModelsabstractSyntactic Transformer language models aim to achieve better generalization through simultaneously modeling syntax trees and sentences.While prior work has been focusing on adding constituency-based structures to Transformers, we introduce Dependency Transformer Grammars (DTGs), a new class of Transformer language model with explicit dependency-based inductive bias.DTGs simulate dependency transition systems with constrained attention patterns by modifying attention masks, incorporate the stack information through relative positional encoding, and augment dependency arc representation with a combination of token embeddings and operation embeddings.When trained on a dataset of sentences annotated with dependency trees, DTGs achieve better generalization while maintaining comparable perplexity with Transformer language model baselines.DTGs also outperform recent constituencybased models, showing that dependency can better guide Transformer language models. Yida Zhao, Chao Lou, Kewei Tu |
ACL (1) | 3 |
| 2024 | Augmenting Transformers with Recursively Composed Multi-grained RepresentationsabstractWe present ReCAT, a recursive composition augmented Transformer that is able to explicitly model hierarchical syntactic structures of raw texts without relying on gold trees during both learning and inference.
Existing research along this line restricts data to follow a hierarchical tree structure and thus lacks inter-span communications.
To overcome the problem, we propose a novel contextual inside-outside (CIO) layer that learns contextualized representations of spans through bottom-up and top-down passes, where a bottom-up pass forms representations of high-level spans by composing low-level spans, while a top-down pass combines information inside and outside a span. By stacking several CIO layers between the embedding layer and the attention layers in Transformer, the ReCAT model can perform both deep intra-span and deep inter-span interactions, and thus generate multi-grained representations fully contextualized with other spans.
Moreover, the CIO layers can be jointly pre-trained with Transformers, making ReCAT enjoy scaling ability, strong performance, and interpretability at the same time. We conduct experiments on various sentence-level and span-level tasks. Evaluation results indicate that ReCAT can significantly outperform vanilla Transformer models on all span-level tasks and recursive models on natural language inference tasks. More interestingly, the hierarchical structures induced by ReCAT exhibit strong consistency with human-annotated syntactic trees, indicating good interpretability brought by the CIO layers. Qingyang Zhu, Kewei Tu |
ICLR | 3 |
| 2023 | Modeling Instance Interactions for Joint Information Extraction with Neural High-Order Conditional Random FieldabstractPrior works on joint Information Extraction (IE) typically model instance (e.g., event triggers, entities, roles, relations) interactions by representation enhancement, type dependencies scoring, or global decoding.We find that the previous models generally consider binary type dependency scoring of a pair of instances, and leverage local search such as beam search to approximate global solutions.To better integrate cross-instance interactions, in this work, we introduce a joint IE framework (CRFIE) that formulates joint IE as a high-order Conditional Random Field.Specifically, we design binary factors and ternary factors to directly model interactions between not only a pair of instances but also triplets.Then, these factors are utilized to jointly predict labels of all instances.To address the intractability problem of exact high-order inference, we incorporate a high-order neural decoder that is unfolded from a mean-field variational inference method, which achieves consistent learning and inference.The experimental results show that our approach achieves consistent improvements on three IE tasks compared with our baseline and prior work. Zixia Jia, Zhaohui Yan 0001, Wenjuan Han, Zilong Zheng, Kewei Tu |
ACL (1) | 5 |
| 2023 | Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity TypingabstractUltra-fine entity typing (UFET) predicts extremely free-formed types (e.g., president, politician) of a given entity mention (e.g., Joe Biden) in context.State-of-the-art (SOTA) methods use the cross-encoder (CE) based architecture.CE concatenates a mention (and its context) with each type and feeds the pair into a pretrained language model (PLM) to score their relevance.It brings deeper interaction between the mention and the type to reach better performance but has to perform N (the type set size) forward passes to infer all the types of a single mention.CE is therefore very slow in inference when the type set is large (e.g., N = 10k for UFET).To this end, we propose to perform entity typing in a recall-expand-filter manner.The recall and expansion stages prune the large type set and generate K (typically much smaller than N ) most relevant type candidates for each mention.At the filter stage, we use a novel model called MCCE to concurrently encode and score all these K candidates in only one forward pass to obtain the final type prediction.We investigate different model options for each stage and conduct extensive experiments to compare each option, experiments show that our method reaches SOTA performance on UFET and is thousands of times faster than the CE-based architecture.We also found our method is very effective in fine-grained (130 types) and coarse-grained (9 types) entity typing. Chengyue Jiang, Wenyang Hui, Yong Jiang 0005, Xiaobin Wang, Pengjun Xie, Kewei Tu |
ACL (1) | 6 |
| 2023 | Do PLMs Know and Understand Ontological Knowledge?abstractOntological knowledge, which comprises classes and properties and their relationships, is integral to world knowledge.It is significant to explore whether Pretrained Language Models (PLMs) know and understand such knowledge.However, existing PLM-probing studies focus mainly on factual knowledge, lacking a systematic probing of ontological knowledge.In this paper, we focus on probing whether PLMs store ontological knowledge and have a semantic understanding of the knowledge rather than rote memorization of the surface form.To probe whether PLMs know ontological knowledge, we investigate how well PLMs memorize: (1) types of entities; (2) hierarchical relationships among classes and properties, e.g., Person is a subclass of Animal and Member of Sports Team is a subproperty of Member of ; (3) domain and range constraints of properties, e.g., the subject of Member of Sports Team should be a Person and the object should be a Sports Team.To further probe whether PLMs truly understand ontological knowledge beyond memorization, we comprehensively study whether they can reliably perform logical reasoning with given knowledge according to ontological entailment rules.Our probing results show that PLMs can memorize certain ontological knowledge and utilize implicit knowledge in reasoning.However, both the memorizing and reasoning performances are less than perfect, indicating incomplete knowledge and understanding. Weiqi Wu, Chengyue Jiang, Yong Jiang 0005, Pengjun Xie, Kewei Tu |
ACL (1) | 5 |
| 2023 | Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span SelectionabstractWe present a simple and unified approach for both continuous and discontinuous constituency parsing via autoregressive span selection.Constituency parsing aims to produce a set of non-crossing spans so that they can form a constituency parse tree.We sort gold spans in a predefined order and train a pointer network to autoregressively select spans by that order.To deal with a discontinuous span, we consecutively select its subspans from left to right, label all but the last subspans with a special discontinuous label, and label the last subspan with the whole discontinuous span's label.We use a simple heuristic to output valid trees from selected spans so that our approach is able to predict all possible continuous and discontinuous constituency trees without sacrificing data coverage and without the need to use expensive chart-based parsing algorithms.Extensive experiments show that our model achieves stateof-the-art or competitive performance on all benchmarks of continuous and discontinuous constituency parsing . 1 Kewei Tu |
ACL (1) | 2 |
| 2023 | COMBO: A Complete Benchmark for Open KG CanonicalizationabstractOpen knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text.The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need to be canonicalized.Existing datasets for open KG canonicalization only provide gold entitylevel canonicalization for noun phrases.In this paper, we present COMBO, a Complete Benchmark for Open KG canonicalization.Compared with existing datasets, we additionally provide gold canonicalization for relation phrases, gold ontology-level canonicalization for noun phrases, as well as source sentences from which triples are extracted.We also propose metrics for evaluating each type of canonicalization.On the COMBO dataset, we empirically compare previously proposed canonicalization methods as well as a few simple baseline methods based on pretrained language models.We find that properly encoding the phrases in a triple using pretrained language models results in better relation canonicalization and ontology-level canonicalization of the noun phrase.We release our dataset, baselines, and evaluation scripts at Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu |
EACL | 6 |
| 2023 | Joint Entity and Relation Extraction with Span Pruning and Hypergraph Neural NetworksabstractEntity and Relation Extraction (ERE) is an important task in information extraction.Recent marker-based pipeline models achieve state-ofthe-art performance, but still suffer from the error propagation issue.Also, most of current ERE models do not take into account higherorder interactions between multiple entities and relations, while higher-order modeling could be beneficial.In this work, we propose Hyper-Graph neural network for ERE (HGERE), which is built upon the PL-marker (a state-of-the-art marker-based pipleline model).To alleviate error propagation,we use a high-recall pruner mechanism to transfer the burden of entity identification and labeling from the NER module to the joint module of our model.For higher-order modeling, we build a hypergraph, where nodes are entities (provided by the span pruner) and relations thereof, and hyperedges encode interactions between two different relations or between a relation and its associated subject and object entities.We then run a hypergraph neural network for higher-order inference by applying message passing over the built hypergraph.Experiments on three widely used benchmarks (ACE2004, ACE2005 and SciERC) for ERE task show significant improvements over the previous state-of-the-art PL-marker. 1 Zhaohui Yan 0001, Wei Liu 0131, Kewei Tu |
EMNLP | 4 |
| 2023 | Using Interpretation Methods for Model EnhancementabstractIn the age of neural natural language processing, there are plenty of works trying to derive interpretations of neural models.Intuitively, when gold rationales exist during training, one can additionally train the model to match its interpretation with the rationales.However, this intuitive idea has not been fully explored.In this paper, we propose a framework of utilizing interpretation methods and gold rationales to enhance models.Our framework is very general in the sense that it can incorporate various interpretation methods.Previously proposed gradient-based methods can be shown as an instance of our framework.We also propose two novel instances utilizing two other types of interpretation methods, erasure/replace-based and extractor-based methods, for model enhancement.We conduct comprehensive experiments on a variety of tasks.Experimental results show that our framework is effective especially in low-resource settings in enhancing models with various interpretation methods, and our two newly-proposed methods outperform gradient-based methods in most settings.Code is available at https://github. com/Chord-Chen-30/UIMER. Chengyue Jiang, Kewei Tu |
EMNLP | 3 |
| 2023 | AMR Parsing with Causal Hierarchical Attention and PointersabstractTranslation-based AMR parsers have recently gained popularity due to their simplicity and effectiveness.They predict linearized graphs as free texts, avoiding explicit structure modeling.However, this simplicity neglects structural locality in AMR graphs and introduces unnecessary tokens to represent coreferences.In this paper, we introduce new target forms of AMR parsing and a novel model, CHAP, which is equipped with causal hierarchical attention and the pointer mechanism, enabling the integration of structures into the Transformer decoder.We empirically explore various alternative modeling options.Experiments show that our model outperforms baseline models on four out of five benchmarks in the setting of no additional data. Chao Lou, Kewei Tu |
EMNLP | 2 |
| 2023 | A Multi-Grained Self-Interpretable Symbolic-Neural Model For Single/Multi-Labeled Text Classification
Xinyu Kong, Kewei Tu |
ICLR | 3 |
| 2022 | Span-Based Semantic Role Labeling with Argument Pruning and Second-Order InferenceabstractWe study graph-based approaches to span-based semantic role labeling. This task is difficult due to the need to enumerate all possible predicate-argument pairs and the high degree of imbalance between positive and negative samples. Based on these difficulties, high-order inference that considers interactions between multiple arguments and predicates is often deemed beneficial but has rarely been used in span-based semantic role labeling. Because even for second-order inference, there are already O(n^5) parts for a sentence of length n, and exact high-order inference is intractable. In this paper, we propose a framework consisting of two networks: a predicate-agnostic argument pruning network that reduces the number of candidate arguments to O(n), and a semantic role labeling network with an optional second-order decoder that is unfolded from an approximate inference algorithm. Our experiments show that our framework achieves significant and consistent improvement over previous approaches. Zixia Jia, Zhaohui Yan 0001, Haoyi Wu, Kewei Tu |
AAAI | 4 |
| 2022 | Nested Named Entity Recognition as Latent Lexicalized Constituency ParsingabstractNested named entity recognition (NER) has been receiving increasing attention.Recently, Fu et al. (2020) adapt a span-based constituency parser to tackle nested NER.They treat nested entities as partially-observed constituency trees and propose the masked inside algorithm for partial marginalization.However, their method cannot leverage entity heads, which have been shown useful in entity mention detection and entity typing.In this work, we resort to more expressive structures, lexicalized constituency trees in which constituents are annotated by headwords, to model nested entities.We leverage the Eisner-Satta algorithm to perform partial marginalization and inference efficiently.In addition, we propose to use (1) a two-stage strategy (2) a head regularization loss and (3) a head-aware labeling loss in order to enhance the performance.We make a thorough ablation study to investigate the functionality of each component.Experimentally, our method achieves the state-ofthe-art performance on ACE2004, ACE2005 and NNE, and competitive performance on GENIA, and meanwhile has a fast inference speed.Our code will be publicly available at: github.com/LouChao98/nner_as_parsing. Chao Lou, Kewei Tu |
ACL (1) | 3 |
| 2022 | Headed-Span-Based Projective Dependency ParsingabstractWe propose a new method for projective dependency parsing based on headed spans.In a projective dependency tree, the largest subtree rooted at each word covers a contiguous sequence (i.e., a span) in the surface order.We call such a span marked by a root word headed span.A projective dependency tree can be represented as a collection of headed spans.We decompose the score of a dependency tree into the scores of the headed spans and design a novel O(n 3 ) dynamic programming algorithm to enable global training and exact inference.Our model achieves state-of-the-art or competitive results on PTB, CTB, and UD 1 . Kewei Tu |
ACL (1) | 2 |
| 2022 | Bottom-Up Constituency Parsing and Nested Named Entity Recognition with Pointer NetworksabstractConstituency parsing and nested named entity recognition (NER) are similar tasks since they both aim to predict a collection of nested and non-crossing spans.In this work, we cast nested NER to constituency parsing and propose a novel pointing mechanism for bottomup parsing to tackle both tasks.The key idea is based on the observation that if we traverse a constituency tree in post-order, i.e., visiting a parent after its children, then two consecutively visited spans would share a boundary.Our model tracks the shared boundaries and predicts the next boundary at each step by leveraging a pointer network.As a result, it needs only linear steps to parse and thus is efficient.It also maintains a parsing configuration for structural consistency, i.e., always outputting valid trees.Experimentally, our model achieves the state-of-the-art performance on PTB among all BERT-based models (96.01 F1 score) and competitive performance on CTB7 in constituency parsing; and it also achieves strong performance on three benchmark datasets of nested NER: ACE2004, ACE2005, and GENIA 1 . Kewei Tu |
ACL (1) | 2 |
| 2022 | Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random FieldabstractUltra-fine entity typing (UFET) aims to predict a wide range of type phrases that correctly describe the categories of a given entity mention in a sentence.Most recent works infer each entity type independently, ignoring the correlations between types, e.g., when an entity is inferred as a president, it should also be a politician and a leader.To this end, we use an undirected graphical model called pairwise conditional random field (PCRF) to formulate the UFET problem, in which the type variables are not only unarily influenced by the input but also pairwisely relate to all the other type variables.We use various modern backbones for entity typing to compute unary potentials, and derive pairwise potentials from type phrase representations that both capture prior semantic information and facilitate accelerated inference.We use mean-field variational inference for efficient type inference on very large type sets and unfold it as a neural network module to enable end-to-end training.Experiments on UFET show that the Neural-PCRF consistently outperforms its backbones with little cost and results in a competitive performance against crossencoder based SOTA while being thousands of times faster.We also find Neural-PCRF effective on a widely used fine-grained entity typing dataset with a smaller type set.We pack Neural-PCRF as a network module that can be plugged onto multi-label type classifiers with ease and release it in github.com/modelscope/ adaseq/examples/NPCRF. Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu |
EMNLP | 5 |
| 2022 | ITA: Image-Text Alignments for Multi-Modal Named Entity RecognitionabstractXinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Kewei Tu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xinyu Wang 0013, Min Gui, Yong Jiang 0005, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Kewei Tu |
NAACL-HLT | 8 |
| 2022 | Dynamic Programming in Rank Space: Scaling Structured Inference with Low-Rank HMMs and PCFGsabstractHidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) are widely used structured models, both of which can be represented as factor graph grammars (FGGs), a powerful formalism capable of describing a wide range of models.Recent research found it beneficial to use large state spaces for HMMs and PCFGs.However, inference with large state spaces is computationally demanding, especially for PCFGs.To tackle this challenge, we leverage tensor rank decomposition (aka.CPD) to decrease inference computational complexities for a subset of FGGs subsuming HMMs and PCFGs.We apply CPD on the factors of an FGG and then construct a new FGG defined in the rank space.Inference with the new FGG produces the same result but has a lower time complexity when the rank size is smaller than the state size.We conduct experiments on HMM language modeling and unsupervised PCFG parsing, showing better performance than previous work.Our code is publicly available at https://github.com/ VPeterV/RankSpace-Models. Wei Liu 0131, Kewei Tu |
NAACL-HLT | 3 |
| 2022 | Improving Constituent Representation with Hypertree Neural NetworksabstractMany natural language processing tasks involve text spans and thus high-quality span representations are needed to enhance neural approaches to these tasks.Most existing methods of span representation are based on simple derivations (such as max-pooling) from word representations and do not utilize compositional structures of natural language.In this paper, we aim to improve representations of constituent spans using a novel hypertree neural networks (HTNN) that is structured with constituency parse trees.Each node in the HTNN represents a constituent of the input sentence and each hyperedge represents a composition of smaller child constituents into a larger parent constituent.In each update iteration of the HTNN, the representation of each constituent is computed based on all the hyperedges connected to it, thus incorporating both bottom-up and top-down compositional information.We conduct comprehensive experiments to evaluate HTNNs against other span representation models and the results show the effectiveness of HTNN. Hao Zhou 0044, Gongshen Liu, Kewei Tu |
NAACL-HLT | 3 |
| 2021 | Multi-View Cross-Lingual Structured Prediction with Minimum SupervisionabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 7 |
| 2021 | Risk Minimization for Zero-shot Sequence LabelingabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 7 |
| 2021 | Improving Named Entity Recognition by External Context Retrieving and Cooperative LearningabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 7 |
| 2021 | Automated Concatenation of Embeddings for Structured PredictionabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 7 |
| 2021 | Structural Knowledge Distillation: Tractably Distilling Information for Structured PredictorabstractXinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Zhaohui Yan 0001, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 9 |
| 2021 | Neural Bi-Lexicalized PCFG InductionabstractSonglin Yang, Yanpeng Zhao, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yanpeng Zhao, Kewei Tu |
ACL/IJCNLP (1) | 3 |
| 2021 | Adapting Unsupervised Syntactic Parsing Methodology for Discourse Dependency ParsingabstractLiwen Zhang, Ge Wang, Wenjuan Han, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ge Wang 0005, Wenjuan Han, Kewei Tu |
ACL/IJCNLP (1) | 4 |
| 2021 | Neuralizing Regular Expressions for Slot FillingabstractNeural models and symbolic rules such as regular expressions have their respective merits and weaknesses.In this paper, we study the integration of the two approaches for the slot filling task by converting regular expressions into neural networks.Specifically, we first convert regular expressions into a special form of finite-state transducers, then unfold its approximate inference algorithm as a bidirectional recurrent neural model that performs slot filling via sequence labeling.Experimental results show that our model has superior zero-shot and few-shot performance and stays competitive when there are sufficient training data. Chengyue Jiang, Zijian Jin, Kewei Tu |
EMNLP (1) | 3 |
| 2021 | PCFGs Can Do Better: Inducing Probabilistic Context-Free Grammars with Many SymbolsabstractProbabilistic context-free grammars (PCFGs) with neural parameterization have been shown to be effective in unsupervised phrasestructure grammar induction.However, due to the cubic computational complexity of PCFG representation and parsing, previous approaches cannot scale up to a relatively large number of (nonterminal and preterminal) symbols.In this work, we present a new parameterization form of PCFGs based on tensor decomposition, which has at most quadratic computational complexity in the symbol number and therefore allows us to use a much larger number of symbols.We further use neural parameterization for the new form to improve unsupervised parsing performance.We evaluate our model across ten languages and empirically demonstrate the effectiveness of using more symbols. Yanpeng Zhao, Kewei Tu |
NAACL-HLT | 3 |
| 2020 | Semi-Supervised Semantic Dependency Parsing Using CRF AutoencodersabstractSemantic dependency parsing, which aims to find rich bi-lexical relationships, allows words to have multiple dependency heads, resulting in graph-structured representations.We propose an approach to semi-supervised learning of semantic dependency parsers based on the CRF autoencoder framework.Our encoder is a discriminative neural semantic dependency parser that predicts the latent parse graph of the input sentence.Our decoder is a generative neural model that reconstructs the input sentence conditioned on the latent parse graph.Our model is arc-factored and therefore parsing and learning are both tractable.Experiments show our model achieves significant and consistent improvement over the supervised baseline. Zixia Jia, Youmi Ma, Jiong Cai, Kewei Tu |
ACL | 4 |
| 2020 | An Empirical Comparison of Unsupervised Constituency Parsing MethodsabstractUnsupervised constituency parsing aims to learn a constituency parser from a training corpus without parse tree annotations.While many methods have been proposed to tackle the problem, including statistical and neural methods, their experimental results are often not directly comparable due to discrepancies in datasets, data preprocessing, lexicalization, and evaluation metrics.In this paper, we first examine experimental settings used in previous work and propose to standardize the settings for better comparability between methods.We then empirically compare several existing methods, including decade-old and newly proposed ones, under the standardized settings on English and Japanese, two languages with different branching tendencies.We find that recent models do not show a clear advantage over decade-old models in our experiments.We hope our work can provide new insights into existing methods and facilitate future empirical evaluation of unsupervised constituency parsing. Jiong Cai, Yong Jiang 0005, Kewei Tu |
ACL | 5 |
| 2020 | Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationabstractOpen-domain dialogue generation has gained increasing attention in Natural Language Processing.Its evaluation requires a holistic means.Human ratings are deemed as the gold standard.As human evaluation is inefficient and costly, an automated substitute is highly desirable.In this paper, we propose holistic evaluation metrics that capture different aspects of open-domain dialogues.Our metrics consist of (1) GPT-2 based context coherence between sentences in a dialogue, (2) GPT-2 based fluency in phrasing, (3) n-gram based diversity in responses to augmented queries, and (4) textual-entailment-inference based logical self-consistency.The empirical validity of our metrics is demonstrated by strong correlations with human judgments.We open source the code and relevant materials.1 Bo Pang 0004, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Kewei Tu |
ACL | 6 |
| 2020 | Structure-Level Knowledge Distillation For Multilingual Sequence LabelingabstractMultilingual sequence labeling is a task of predicting label sequences using a single unified model for multiple languages.Compared with relying on multiple monolingual models, using a multilingual model has the benefit of a smaller model size, easier in online serving, and generalizability to low-resource languages.However, current multilingual models still underperform individual monolingual models significantly due to model capacity limitations.In this paper, we propose to reduce the gap between monolingual models and the unified multilingual model by distilling the structural knowledge of several monolingual models (teachers) to the unified multilingual model (student).We propose two novel KD methods based on structure-level information:(1) approximately minimizes the distance between the student's and the teachers' structurelevel probability distributions, (2) aggregates the structure-level knowledge to local distributions and minimizes the distance between two local probability distributions.Our experiments on 4 multilingual tasks with 25 datasets show that our approaches outperform several strong baselines and have stronger zero-shot generalizability than both the baseline model and teacher models. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Fei Huang 0002, Kewei Tu |
ACL | 6 |
| 2020 | A Survey of Unsupervised Dependency ParsingabstractSyntactic dependency parsing is an important task in natural language processing.Unsupervised dependency parsing aims to learn a dependency parser from sentences that have no annotation of their correct parse trees.Despite its difficulty, unsupervised parsing is an interesting research direction because of its capability of utilizing almost unlimited unannotated text data.It also serves as the basis for other research in low-resource parsing.In this paper, we survey existing approaches to unsupervised dependency parsing, identify two major classes of approaches, and discuss recent trends.We hope that our survey can provide insights for researchers and facilitate future research on this topic. Wenjuan Han, Yong Jiang 0005, Hwee Tou Ng, Kewei Tu |
COLING | 4 |
| 2020 | Deep Inside-outside Recursive Autoencoder with All-span ObjectiveabstractDeep inside-outside recursive autoencoder (DIORA) is a neural-based model designed for unsupervised constituency parsing.During its forward computation, it provides phrase and contextual representations for all spans in the input sentence.By utilizing the contextual representation of each leaf-level span, the span of length 1, to reconstruct the word inside the span, the model is trained without labeled data.In this work, we extend the training objective of DIORA by making use of all spans instead of only leaf-level spans.We test our new training objective on datasets of two languages: English and Japanese, and empirically show that our method achieves improvement in parsing accuracy over the original DIORA. Ruyue Hong, Jiong Cai, Kewei Tu |
COLING | 3 |
| 2020 | Semi-Supervised Dependency Parsing with Arc-Factored Variational AutoencodingabstractMannual annotation for dependency parsing is both labourious and time costly, resulting in the difficulty to learn practical dependency parsers for many languages due to the lack of labelled training corpora.To compensate for the scarcity of labelled data, semi-supervised dependency parsing methods are developed to utilize unlabelled data in the training procedure of dependency parsers.In previous work, the autoencoder framework is a prevalent approach for the utilization of unlabelled data.In this framework, training sentences are reconstructed from a decoder conditioned on dependency trees predicted by an encoder.The tree structure requirement brings challenges for both the encoder and the decoder.Sophisticated techniques are employed to tackle these challenges at the expense of model complexity and approximations in encoding and decoding.In this paper, we propose a model based on the variational autoencoder framework.By relaxing the tree constraint in both the encoder and the decoder during training, we make the learning of our model fully arc-factored and thus circumvent the challenges brought by the tree constraint.We evaluate our model on datasets across several languages and the results demonstrate the advantage of our model over previous approaches in both parsing accuracy and speed. Ge Wang 0005, Kewei Tu |
COLING | 2 |
| 2020 | Second-Order Unsupervised Neural Dependency ParsingabstractMost of the unsupervised dependency parsers are based on first-order probabilistic generative models that only consider local parent-child information.Inspired by second-order supervised dependency parsing, we proposed a second-order extension of unsupervised neural dependency models that incorporate grandparent-child or sibling information.We also propose novel design of the neural parameterization and optimization methods of the dependency models.In secondorder models, the number of grammar rules grows cubically with the increase of vocabulary size, making it difficult to train lexicalized models that may contain thousands of words.To circumvent this problem while still benefiting from both second-order parsing and lexicalization, we use the agreement-based learning framework to jointly train a second-order unlexicalized model and a first-order lexicalized model.Experiments on multiple datasets show the effectiveness of our second-order models compared with recent state-of-the-art methods.Our joint model achieves a 10% improvement over the previous state-of-the-art parser on the full WSJ test set. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
COLING | 4 |
| 2020 | Adversarial Attack and Defense of Structured Prediction ModelsabstractBuilding an effective adversarial attacker and elaborating on countermeasures for adversarial attacks for natural language processing (NLP) have attracted a lot of research in recent years.However, most of the existing approaches focus on classification problems.In this paper, we investigate attacks and defenses for structured prediction tasks in NLP.Besides the difficulty of perturbing discrete words and the sentence fluency problem faced by attackers in any NLP tasks, there is a specific challenge to attackers of structured prediction models: the structured output of structured prediction models is sensitive to small perturbations in the input.To address these problems, we propose a novel and unified framework that learns to attack a structured prediction model using a sequence-to-sequence model with feedbacks from multiple reference models of the same structured prediction task.Based on the proposed attack, we further reinforce the victim model with adversarial training, making its prediction more robust and accurate.We evaluate the proposed framework in dependency parsing and part-of-speech tagging.Automatic and human evaluations show that our proposed framework succeeds in both attacking state-of-the-art structured prediction models and boosting them with adversarial training. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP (1) | 4 |
| 2020 | Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural NetworksabstractNeural networks can achieve impressive performance on many natural language processing applications, but they typically need large labeled data for training and are not easily interpretable.On the other hand, symbolic rules such as regular expressions are interpretable, require no training, and often achieve decent accuracy; but rules cannot benefit from labeled data when available and hence underperform neural networks in rich-resource scenarios.In this paper, we propose a type of recurrent neural networks called FA-RNNs that combine the advantages of neural networks and regular expression rules.An FA-RNN can be converted from regular expressions and deployed in zero-shot and cold-start scenarios.It can also utilize labeled data for training to achieve improved prediction accuracy.After training, an FA-RNN often remains interpretable and can be converted back into regular expressions.We apply FA-RNNs to text classification and observe that FA-RNNs significantly outperform previous neural approaches in both zeroshot and low-resource settings and remain very competitive in rich-resource settings. Chengyue Jiang, Yinggong Zhao, Shanbo Chu, Libin Shen, Kewei Tu |
EMNLP (1) | 5 |
| 2020 | AIN: Fast and Accurate Sequence Labeling with Approximate Inference NetworkabstractThe linear-chain Conditional Random Field (CRF) model is one of the most widely-used neural sequence labeling approaches.Exact probabilistic inference algorithms such as the forward-backward and Viterbi algorithms are typically applied in training and prediction stages of the CRF model.However, these algorithms require sequential computation that makes parallelization impossible.In this paper, we propose to employ a parallelizable approximate variational inference algorithm for the CRF model.Based on this algorithm, we design an approximate inference network that can be connected with the encoder of the neural CRF model to form an end-to-end network, which is amenable to parallelization for faster training and prediction.The empirical results show that our proposed approaches achieve a 12.7-fold improvement in decoding speed with long sentences and a competitive accuracy compared with the traditional CRF approach. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
EMNLP (1) | 7 |
| 2019 | Bidirectional Transition-Based Dependency ParsingabstractTransition-based dependency parsing is a fast and effective approach for dependency parsing. Traditionally, a transitionbased dependency parser processes an input sentence and predicts a sequence of parsing actions in a left-to-right manner. During this process, an early prediction error may negatively impact the prediction of subsequent actions. In this paper, we propose a simple framework for bidirectional transitionbased parsing. During training, we learn a left-to-right parser and a right-to-left parser separately. To parse a sentence, we perform joint decoding with the two parsers. We propose three joint decoding algorithms that are based on joint scoring, dual decomposition, and dynamic oracle respectively. Empirical results show that our methods lead to competitive parsing accuracy and our method based on dynamic oracle consistently achieves the best performance. Yunzhe Yuan, Yong Jiang 0005, Kewei Tu |
AAAI | 3 |
| 2019 | Enhancing Unsupervised Generative Dependency Parser with Contextual InformationabstractMost of the unsupervised dependency parsers are based on probabilistic generative models that learn the joint distribution of the given sentence and its parse.Probabilistic generative models usually explicit decompose the desired dependency tree into factorized grammar rules, which lack the global features of the entire sentence.In this paper, we propose a novel probabilistic model called discriminative neural dependency model with valence (D-NDMV) that generates a sentence and its parse from a continuous latent representation, which encodes global contextual information of the generated sentence.We propose two approaches to model the latent representation: the first deterministically summarizes the representation from the sentence and the second probabilistically models the representation conditioned on the sentence.Our approach can be regarded as a new type of autoencoder model to unsupervised dependency parsing that combines the benefits of both generative and discriminative techniques.In particular, our approach breaks the context-free independence assumption in previous generative approaches and therefore becomes more expressive.Our extensive experimental results on seventeen datasets from various sources show that our approach achieves competitive accuracy compared with both generative and discriminative state-of-the-art unsupervised dependency parsers. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
ACL (1) | 3 |
| 2019 | Second-Order Semantic Dependency Parsing with End-to-End Neural NetworksabstractSemantic dependency parsing aims to identify semantic relationships between words in a sentence that form a graph.In this paper, we propose a second-order semantic dependency parser, which takes into consideration not only individual dependency edges but also interactions between pairs of edges.We show that second-order parsing can be approximated using mean field (MF) variational inference or loopy belief propagation (LBP).We can unfold both algorithms as recurrent layers of a neural network and therefore can train the parser in an end-to-end manner.Our experiments show that our approach achieves stateof-the-art performance. Xinyu Wang 0013, Jingxian Huang, Kewei Tu |
ACL (1) | 3 |
| 2019 | Latent Variable Sentiment GrammarabstractNeural models have been investigated for sentiment classification over constituent trees.They learn phrase composition automatically by encoding tree structures but do not explicitly model sentiment composition, which requires to encode sentiment class labels.To this end, we investigate two formalisms with deep sentiment representations that capture sentiment subtype expressions by latent variables and Gaussian mixture vectors, respectively.Experiments on Stanford Sentiment Treebank (SST) show the effectiveness of sentiment grammar over vanilla neural encoders.Using ELMo embeddings, our method gives the best results on this benchmark. Kewei Tu, Yue Zhang 0004 |
ACL (1) | 2 |
| 2019 | Multilingual Grammar Induction with Continuous Language IdentificationabstractWenjuan Han, Ge Wang, Yong Jiang, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wenjuan Han, Ge Wang 0005, Yong Jiang 0005, Kewei Tu |
EMNLP/IJCNLP (1) | 4 |
| 2019 | A Regularization-based Framework for Bilingual Grammar InductionabstractYong Jiang, Wenjuan Han, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Lexicalized Neural Unsupervised Dependency Parsing
Wenjuan Han, Yong Jiang 0005, Kewei Tu |
Neurocomputing | 3 |
| 2019 | Learning and evaluation of latent dependency forest models
Yong Jiang 0005, Kewei Tu |
Neural Comput. Appl. | 3 |
| 2018 | Maximum A Posteriori Inference in Sum-Product NetworksabstractSum-product networks (SPNs) are a class of probabilistic graphical models that allow tractable marginal inference. However, the maximum a posteriori (MAP) inference in SPNs is NP-hard. We investigate MAP inference in SPNs from both theoretical and algorithmic perspectives. For the theoretical part, we reduce general MAP inference to its special case without evidence and hidden variables; we also show that it is NP-hard to approximate the MAP problem to 2nε for fixed 0 ≤ ε < 1, where n is the input size. For the algorithmic part, we first present an exact MAP solver that runs reasonably fast and could handle SPNs with up to 1k variables and 150k arcs in our experiments. We then present a new approximate MAP solver with a good balance between speed and accuracy, and our comprehensive experiments on real-world datasets show that it has better overall performance than existing approximate solvers. Jun Mei, Yong Jiang 0005, Kewei Tu |
AAAI | 3 |
| 2018 | Gaussian Mixture Latent Vector GrammarsabstractWe introduce Latent Vector Grammars (LVeGs), a new framework that extends latent variable grammars such that each nonterminal symbol is associated with a continuous vector space representing the set of (infinitely many) subtypes of the nonterminal.We show that previous models such as latent variable grammars and compositional vector grammars can be interpreted as special cases of LVeGs.We then present Gaussian Mixture LVeGs (GM-LVeGs), a new special case of LVeGs that uses Gaussian mixtures to formulate the weights of production rules over subtypes of nonterminals.A major advantage of using Gaussian mixtures is that the partition function and the expectations of subtype rules can be computed using an extension of the inside-outside algorithm, which enables efficient inference and learning.We apply GM-LVeGs to part-of-speech tagging and constituency parsing and show that GM-LVeGs can achieve competitive accuracies.Our code is available at https://github.com/zhaoyanpeng/lveg. Yanpeng Zhao, Kewei Tu |
ACL (1) | 3 |
| 2018 | QA4IE: A Question Answering Based Framework for Information Extraction
Hao Zhou 0044, Yanru Qu, Weinan Zhang 0001, Suoheng Li, Shu Rong, Dongyu Ru, Lihua Qian, Kewei Tu, Yong Yu 0001 |
ISWC (1) | 9 |
| 2017 | Latent Dependency Forest ModelsabstractProbabilistic modeling is one of the foundations of modern machine learning and artificial intelligence. In this paper, we propose a novel type of probabilistic models named latent dependency forest models (LDFMs). A LDFM models the dependencies between random variables with a forest structure that can change dynamically based on the variable values. It is therefore capable of modeling context-specific independence. We parameterize a LDFM using a first-order non-projective dependency grammar. Learning LDFMs from data can be formulated purely as a parameter learning problem, and hence the difficult problem of model structure learning is circumvented. Our experimental results show that LDFMs are competitive with existing probabilistic models. Shanbo Chu, Yong Jiang 0005, Kewei Tu |
AAAI | 3 |
| 2017 | CRF Autoencoder for Unsupervised Dependency ParsingabstractUnsupervised dependency parsing, which tries to discover linguistic dependency structures from unannotated data, is a very challenging task.Almost all previous work on this task focuses on learning generative models.In this paper, we develop an unsupervised dependency parsing model based on the CRF autoencoder.The encoder part of our model is discriminative and globally normalized which allows us to use rich features as well as universal linguistic priors.We propose an exact algorithm for parsing as well as a tractable learning algorithm.We evaluated the performance of our model on eight multilingual treebanks and found that our model achieved comparable performance with state-of-the-art approaches. Jiong Cai, Yong Jiang 0005, Kewei Tu |
EMNLP | 3 |
| 2017 | Dependency Grammar Induction with Neural Lexicalization and Big Training DataabstractWe study the impact of big models (in terms of the degree of lexicalization) and big data (in terms of the training corpus size) on dependency grammar induction.We experimented with L-DMV, a lexicalized version of Dependency Model with Valence (Klein and Manning, 2004) and L-NDMV, our lexicalized extension of the Neural Dependency Model with Valence (Jiang et al., 2016).We find that L-DMV only benefits from very small degrees of lexicalization and moderate sizes of training corpora.L-NDMV can benefit from big training data and lexicalization of greater degrees, especially when enhanced with good model initialization, and it achieves a result that is competitive with the current state-of-the-art. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP | 3 |
| 2017 | Combining Generative and Discriminative Approaches to Unsupervised Dependency Parsing via Dual DecompositionabstractUnsupervised dependency parsing aims to learn a dependency parser from unannotated sentences.Existing work focuses on either learning generative models using the expectation-maximization algorithm and its variants, or learning discriminative models using the discriminative clustering algorithm.In this paper, we propose a new learning strategy that learns a generative model and a discriminative model jointly based on the dual decomposition method.Our method is simple and general, yet effective to capture the advantages of both models and improve their learning results.We tested our method on the UD treebank and achieved a state-ofthe-art performance on thirty languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 3 |
| 2017 | Semi-supervised Structured Prediction with Neural CRF AutoencoderabstractIn this paper we propose an end-to-end neural CRF autoencoder (NCRF-AE) model for semi-supervised learning of sequential structured prediction problems. Our NCRF-AE consists of two parts: an encoder which is a CRF model enhanced by deep neural networks, and a decoder which is a generative model trying to reconstruct the input. Our model has a unified structure with different loss functions for labeled and unlabeled data with shared parameters. We developed a variation of the EM algorithm for optimizing both the encoder and the decoder simultaneously by decoupling their parameters. Our Experimental results over the Part-of-Speech (POS) tagging task on eight different languages, show that our model can outperform competitive systems in both supervised and semi-supervised scenarios. Xiao Zhang 0017, Yong Jiang 0005, Kewei Tu, Dan Goldwasser |
EMNLP | 4 |
| 2017 | Structured Attentions for Visual Question AnsweringabstractVisual attention, which assigns weights to image regions according to their relevance to a question, is considered as an indispensable part by most Visual Question Answering models. Although the questions may involve complex rela- tions among multiple regions, few attention models can ef- fectively encode such cross-region relations. In this paper, we demonstrate the importance of encoding such relations by showing the limited effective receptive field of ResNet on two datasets, and propose to model the visual attention as a multivariate distribution over a grid-structured Con- ditional Random Field on image regions. We demonstrate how to convert the iterative inference algorithms, Mean Field and Loopy Belief Propagation, as recurrent layers of an end-to-end neural network. We empirically evalu- ated our model on 3 datasets, in which it surpasses the best baseline model of the newly released CLEVR dataset [13] by 9.5%, and the best published model on the VQA dataset [3] by 1.25%. Source code is available at https://github.com/zhuchen03/vqa-sva. Yanpeng Zhao, Shuaiyi Huang, Kewei Tu |
ICCV | 4 |
| 2017 | Learning Bayesian network structures under incremental construction curricula
Yanpeng Zhao, Yetian Chen, Kewei Tu, Jin Tian 0001 |
Neurocomputing | 3 |
| 2016 | Unsupervised Neural Dependency ParsingabstractUnsupervised dependency parsing aims to learn a dependency grammar from text annotated with only POS tags.Various features and inductive biases are often used to incorporate prior knowledge into learning.One useful type of prior information is that there exist correlations between the parameters of grammar rules involving different POS tags.Previous work employed manually designed features or special prior distributions to encode such information.In this paper, we propose a novel approach to unsupervised dependency parsing that uses a neural model to predict grammar rule probabilities based on distributed representation of POS tags.The distributed representation is automatically learned from data and captures the correlations between POS tags.Our experiments show that our approach outperforms previous approaches utilizing POS correlations and is competitive with recent state-of-the-art approaches on nine different languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 3 |
| 2016 | Context-Dependent Sense EmbeddingabstractWord embedding has been widely studied and proven helpful in solving many natural language processing tasks.However, the ambiguity of natural language is always a problem on learning high quality word embeddings.A possible solution is sense embedding which trains embedding for each sense of words instead of each word.Some recent work on sense embedding uses context clustering methods to determine the senses of words, which is heuristic in nature.Other work creates a probabilistic model and performs word sense disambiguation and sense embedding iteratively.However, most of the previous work has the problems of learning sense embeddings based on imperfect word embeddings as well as ignoring the dependency between sense choices of neighboring words.In this paper, we propose a novel probabilistic model for sense embedding that is not based on problematic word embedding of polysemous words and takes into account the dependency between sense choices.Based on our model, we derive a dynamic programming inference algorithm and an Expectation-Maximization style unsupervised learning algorithm.The empirical studies show that our model outperforms the state-of-the-art model on a word sense induction task by a 13% relative gain. Kewei Tu, Yong Yu 0001 |
EMNLP | 2 |
| 2016 | Modified Dirichlet Distribution: Allowing Negative Parameters to Induce Stronger SparsityabstractThe Dirichlet distribution (Dir) is one of the most widely used prior distributions in statistical approaches to natural language processing. The parameters of Dir are required to be positive, which significantly limits its strength as a sparsity prior. In this paper, we propose a simple modification to the Dirichlet distribution that allows the parameters to be negative. Our modified Dirichlet distribution (mDir) not only induces much stronger sparsity, but also simultaneously performs smoothing. mDir is still conjugate to the multinomial distribution, which simplifies posterior inference. We introduce two simple and efficient algorithms for finding the mode of mDir. Our experiments on learning Gaussian mixtures and unsupervised dependency parsing demonstrate the advantage of mDir over Dir. © 2016 Association for Computational Linguistics Kewei Tu |
EMNLP | 1 |
| 2016 | Stochastic and-or Grammars: A Unified Framework and Logic Perspective
Kewei Tu |
IJCAI | 1 |
| 2015 | Curriculum Learning of Bayesian Network Structures
Yanpeng Zhao, Yetian Chen, Kewei Tu, Jin Tian 0001 |
ACML | 3 |
| 2013 | Unsupervised Structure Learning of Stochastic And-Or GrammarsabstractStochastic And-Or grammars compactly represent both compositionality and reconfigurability and have been used to model different types of data such as images and events. We present a unified formalization of stochastic And-Or grammars that is agnostic to the type of the data being modeled, and propose an unsupervised approach to learning the structures as well as the parameters of such grammars. Starting from a trivial initial grammar, our approach iteratively induces compositions and reconfigurations in a unified manner and optimizes the posterior probability of the grammar. In our empirical evaluation, we applied our approach to learning event grammars and image grammars and achieved comparable or better performance than previous approaches. Kewei Tu, Maria Pavlovskaia, Song-Chun Zhu |
NIPS | 1 |
| 2012 | Unambiguity Regularization for Unsupervised Learning of Probabilistic Grammars
Kewei Tu, Vasant G. Honavar |
EMNLP-CoNLL | 1 |
| 2011 | On the Utility of Curricula in Unsupervised Learning of Probabilistic GrammarsabstractSection 2 gives more details of the experimental settings and results. Section 3 discusses the related work. Kewei Tu, Vasant G. Honavar |
IJCAI | 1 |
| 2011 | Exemplar-based Robust Coherent BiclusteringabstractThe biclustering, co-clustering, or subspace clustering problem involves simultaneously grouping the rows and columns of a data matrix to uncover biclusters or sub-matrices of the data matrix that optimize a desired objective function. In coherent biclustering, the objective function contains a coherence measure of the biclusters. We introduce a novel formulation of the coherent biclustering problem and use it to derive two algorithms. The first algorithm is based on loopy message passing; and the second relies on a greedy strategy yielding an algorithm that is significantly faster than the first. A distinguishing feature of these algorithms is that they identify an exemplar or a prototypical member of each bicluster. We note the interference from background elements in biclustering, and offer a means to circumvent such interference using additional regularization. Our experiments with synthetic as well as real-world datasets show that our algorithms are competitive with the current state-of-the-art algorithms for finding coherent biclusters. Kewei Tu, Xixiu Ouyang, Dingyi Han, Vasant G. Honavar |
SDM | 1 |
| 2005 | CMC: Combining Multiple Schema-Matching Strategies Based on Credibility Prediction
Kewei Tu, Yong Yu 0001 |
DASFAA | 1 |
| 2005 | Towards Imaging Large-Scale Ontologies for Quick Understanding and Analysis
Kewei Tu, Miao Xiong, Lei Zhang 0007, Yong Yu 0001 |
ISWC | 1 |
| 2005 | An Approach to RDF(S) Query, Manipulation and Inference on Databases
Yong Yu 0001, Kewei Tu, Lei Zhang 0007 |
WAIM | 3 |
| 2004 | ORIENT: Integrate Ontology Engineering into Industry Tooling Environment
Lei Zhang 0007, Yong Yu 0001, Kewei Tu, MingChuan Guo, Guo Tong Xie, Zhong Su |
ISWC | 5 |