Xuezhe Ma

dblp:127/0230 · DBLP profile ↗
← Back
36ranked-venue papers
14as first author
19since 2021 · last 2026
0000-0001-7582-1653ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 14 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MegaCoin: Enhancing Medium-Grained Color Perception for Vision-Language Models
abstract
In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction. However, despite advances in multimodal modeling, there remains a significant lack of specialized datasets that rigorously evaluate a model's capacity to discern subtle color variations and spatial context---critical elements for situational comprehension and reliable deployment across real-world applications. Toward that goal, we curate MegaCoin, a high-quality, human-labeled dataset based on \emph{real} images with various contextual attributes. MegaCoin consists of two parts: MegaCoin-Instruct, which serves as a supervised fine-tuning (SFT) dataset for VLMs; and MegaCoin-Bench, an annotated test set that can be used as a stand-alone QA dataset. MegaCoin provides three annotated features for 220,000 real images: foreground color, background color, and description of an object's physical environment, constituting 660k human annotations. In addition, MegaCoin can be applied to benchmark domain generalization (DG) algorithms. We explore benchmarking DG methods in the linear probing setup for VLM and show some new insights. Last but not least, we show that VLMs, including GPT-4o, have subpar color recognition capabilities, and fine-tuning with MegaCoin can result in improved performance on visual evaluation tasks. In certain cases, MegaCoin fine-tuned small-scale open-source models such as LLaVA and Bunny can outperform closed-source GPT-4o. We hope the utilities of MegaCoin can shed light on the directions VLMs can improve and provide a more complex platform for domain generalization algorithms.
Ming-Chang Chiu, Shicheng Wen, Xuezhe Ma
AAAI4
2025 Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
abstract
We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy
ICLR8
2025 LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
abstract
Nan Xu, Xuezhe Ma. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xuezhe Ma
NAACL (Long Papers)2
2024 MIDDAG: Where Does Our News Go? Investigating Information Diffusion via Community-Level Information Pathways
abstract
We present MIDDAG, an intuitive, interactive system that visualizes the information propagation paths on social media triggered by COVID-19-related news articles accompanied by comprehensive insights including user/community susceptibility level, as well as events and popular opinions raised by the crowd while propagating the information. Besides discovering information flow patterns among users, we construct communities among users and develop the propagation forecasting capability, enabling tracing and understanding of how information is disseminated at a higher level. A demo video and more are available at https://info-pathways.github.io.
Mingyu Derek Ma, Alexander K. Taylor 0002, Nuan Wen, Po-Nien Kung, Wenna Qin, Shicheng Wen, Azure Zhou, Diyi Yang, Xuezhe Ma, Nanyun Peng 0001, Wei Wang 0010
AAAI10
2024 Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
abstract
The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities.
Xuezhe Ma, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang 0025, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou
NeurIPS1
2023 RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation
abstract
Endowing chatbots with a consistent persona is essential to an engaging conversation, yet it remains an unresolved challenge.In this work, we propose a new retrieval-enhanced approach for personalized response generation.Specifically, we design a hierarchical transformer retriever trained on dialogue domain data to perform personalized retrieval and a context-aware prefix encoder that fuses the retrieved information to the decoder more effectively.Extensive experiments on a real-world dataset demonstrate the effectiveness of our model at generating more fluent and personalized responses.We quantitatively evaluate our model's performance under a suite of human and automatic metrics and find it to be superior compared to state-of-the-art baselines on English Reddit conversations. 1
Hyundong Cho, Marjorie Freedman, Xuezhe Ma, Jonathan May
ACL (1)4
2023 Challenges in Context-Aware Neural Machine Translation
abstract
Context-aware neural machine translation, a paradigm that involves leveraging information beyond sentence-level context to resolve intersentential discourse dependencies and improve document-level translation quality, has given rise to a number of recent techniques.However, despite well-reasoned intuitions, most context-aware translation models yield only modest improvements over sentence-level systems.In this work, we investigate and present several core challenges, relating to discourse phenomena, context usage, model architectures, and document-level evaluation, that impede progress within the field.To address these problems, we propose a more realistic setting for document-level translation, called paragraphto-paragraph (PARA2PARA) translation, and collect a new dataset of Chinese-English novels to promote future research.1 † Equal contribution. 1 We release the paper's code and dataset here: https: //github.com/Linghao-Jin/canmt-challenges.
Linghao Jin, Jacqueline He, Jonathan May, Xuezhe Ma
EMNLP4
2023 Evaluating Large Language Models on Controlled Generation Tasks
abstract
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Rahul Gupta 0001, John Wieting, Nanyun Peng 0001, Xuezhe Ma
EMNLP9
2023 Look-back Decoding for Open-Ended Text Generation
abstract
Given a prefix (context), open-ended generation aims to decode texts that are coherent, which do not abruptly drift from previous topics, and informative, which do not suffer from undesired repetitions.In this paper, we propose Look-back , an improved decoding algorithm that leverages the Kullback-Leibler divergence to track the distribution distance between current and historical decoding steps.Thus Lookback can automatically predict potential repetitive phrase and topic drift, and remove tokens that may cause the failure modes, restricting the next token probability distribution within a plausible distance to the history.We perform decoding experiments on document continuation and story generation, and demonstrate that Look-back is able to generate more fluent and coherent text, outperforming other strong decoding methods significantly in both automatic and human evaluations 1 .
Chunting Zhou, Asli Celikyilmaz, Xuezhe Ma
EMNLP4
2023 Better May Not Be Fairer: A Study on Subgroup Discrepancy in Image Classification
abstract
In this paper, we provide 20,000 non-trivial human annotations on popular datasets as a first step to bridge gap to studying how natural semantic spurious features affect image classification, as prior works often study datasets mixing low-level features due to limitations in accessing realistic datasets. We investigate how natural background colors play a role as spurious features by annotating the test sets of CIFAR10 and CIFAR100 into subgroups based on the background color of each image. We name our datasets CIFAR10-B and CIFAR100-B1and integrate them with CIFAR-Cs.We find that overall human-level accuracy does not guarantee consistent subgroup performances, and the phenomenon remains even on models pre-trained on ImageNet or after data augmentation (DA). To alleviate this issue, we propose FlowAug, a semantic DA that leverages decoupled semantic representations captured by a pre-trained generative flow. Experimental results show that FlowAug achieves more consistent subgroup results than other types of DA methods on CIFAR10/100 and on CIFAR10/100-C. Additionally, it shows better generalization performance.Furthermore, we propose a generic metric, MacroStd, for studying model robustness to spurious correlations, where we take a macro average on the weighted standard deviations across different classes. We show MacroStd being more predictive of better performances; per our metric, FlowAug demonstrates improvements on subgroup discrepancy. Although this metric is proposed to study our curated datasets, it applies to all datasets that have subgroups or subclasses. Lastly, we also show superior out-of-distribution results on CIFAR10.1.
Ming-Chang Chiu, Xuezhe Ma
ICCV3
2023 Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, Luke Zettlemoyer
ICLR1
2023 LIMA: Less Is More for Alignment
abstract
Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences. We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling. LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history. Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data. In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback. Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output.
Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy
NeurIPS7
2022 Improving Stability of Fine-Tuning Pretrained Language Models via Component-Wise Gradient Norm Clipping
abstract
Fine-tuning over large pretrained language models (PLMs) has established many stateof-the-art results.Despite its superior performance, such fine-tuning can be unstable, resulting in significant variance in performance and potential risks for practical applications.Previous works have attributed such instability to the catastrophic forgetting problem in the top layers of PLMs, which indicates iteratively fine-tuning layers in top-down manner is a promising solution.In this paper, we first point out that this method does not always work out due to different convergence speeds of different layers/modules.Inspired by this observation, we propose a simple componentwise gradient norm clipping method to adjust the convergence speed for different components.Experiment results demonstrate that our method achieves consistent improvements in terms of generalization performance, convergence speed and training stability.The codebase can be found at https://github.com/ yangalan123/FineTuningStability.
Chenghao Yang 0001, Xuezhe Ma
EMNLP2
2022 Towards a Unified View of Parameter-Efficient Transfer Learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig
ICLR3
2021 AESOP: Paraphrase Generation with Adaptive Syntactic Control
abstract
We propose to control paraphrase generation through carefully chosen target syntactic structures to generate more proper and higher quality paraphrases. Our model, AESOP, leverages a pretrained language model and adds deliberately chosen syntactical control via a retrieval-based selection module to generate fluent paraphrases. Experiments show that AESOP achieves state-of-the-art performances on semantic preservation and syntactic conformation on two benchmark datasets with ground-truth syntactic control from human-annotated exemplars. Moreover, with the retrieval-based target syntax selection module, AESOP generates paraphrases with even better qualities than the current best model using human-annotated target syntactic parses according to human evaluation. We further demonstrate the effectiveness of AESOP to improve classification models' robustness to syntactic perturbation by data augmentation on two GLUE tasks.
Jiao Sun, Xuezhe Ma, Nanyun Peng 0001
EMNLP (1)2
2021 Decoupling Global and Local Representations via Invertible Generative Flows
Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy
ICLR1
2021 Examining and Combating Spurious Features under Distribution Shift
abstract
A central goal of machine learning is to learn robust representations that capture the fundamental relationship between inputs and output labels. However, minimizing training errors over finite or biased datasets results in models latching on to spurious correlations between the training input/output pairs that are not fundamental to the problem at hand. In this paper, we define and analyze robust and spurious representations using the information-theoretic concept of minimal sufficient statistics. We prove that even when there is only bias of the input distribution (i.e. covariate shift), models can still pick up spurious features from their training data. Group distributionally robust optimization (DRO) provides an effective tool to alleviate covariate shift by minimizing the worst-case training losses over a set of pre-defined groups. Inspired by our analysis, we demonstrate that group DRO can fail when groups do not directly account for various spurious correlations that occur in the data. To address this, we further propose to minimize the worst-case losses over a more flexible set of distributions that are defined on the joint distribution of groups and instances, instead of treating each group as a whole at optimization time. Through extensive experiments on one image and two language tasks, we show that our model is significantly more robust than comparable baselines under various partitions.
Chunting Zhou, Xuezhe Ma, Paul Michel, Graham Neubig
ICML2
2021 Personalized Response Generation via Generative Split Memory Network
abstract
Despite the impressive successes of generation and dialogue systems, how to endow a text generation system with particular personality traits to deliver more personalized responses remains under-investigated.In this work, we look at how to generate personalized responses for questions on Reddit by utilizing personalized user profiles and posting histories.Specifically, we release an open-domain single-turn dialog dataset made up of 1.5M conversation pairs together with 300k profiles of users and related comments.We then propose a memory network to generate personalized responses in dialogue that utilizes a novel mechanism of splitting memories: one for user profile meta attributes and the other for user-generated information like comment histories.Experimental results show the quantitative and qualitative improvements of our simple split memory network model over the state-of-the-art response generation baselines.The dataset and code are available here.
Yuwei Wu 0003, Xuezhe Ma, Diyi Yang
NAACL-HLT2
2021 Luna: Linear Unified Nested Attention
abstract
The quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Luna, a linear unified nested attention mechanism that approximates softmax attention with two nested linear attention functions, yielding only linear (as opposed to quadratic) time and space complexity. Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length. Then, the packed sequence is unpacked using the second attention function. As compared to a more traditional attention mechanism, Luna introduces an additional sequence with a fixed length as input and an additional corresponding output, which allows Luna to perform attention operation linearly, while also storing adequate contextual information. We perform extensive evaluations on three benchmarks of sequence modeling tasks: long-context sequence modelling, neural machine translation and masked language modeling for large-scale pretraining. Competitive or even better experimental results demonstrate both the effectiveness and efficiency of Luna compared to a variety of strong baseline methods including the full-rank attention and other efficient sparse and dense attention methods.
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma 0001, Luke Zettlemoyer
NeurIPS1
2020 A Two-Step Approach for Implicit Event Argument Detection
abstract
In this work, we explore the implicit event argument detection task, which studies event arguments beyond sentence boundaries.The addition of cross-sentence argument candidates imposes great challenges for modeling.To reduce the number of candidates, we adopt a two-step approach, decomposing the problem into two sub-problems: argument head-word detection and head-to-span expansion.Evaluated on the recent RAMS dataset (Ebner et al., 2020), our model achieves overall better performance than a strong sequence labeling baseline.We further provide detailed error analysis, presenting where the model mainly makes errors and indicating directions for future improvements.It remains a challenge to detect implicit arguments, calling for more future work of document-level modeling for this task.
Zhisong Zhang, Xiang Kong, Zhengzhong Liu 0001, Xuezhe Ma, Eduard H. Hovy
ACL4
2019 Choosing Transfer Languages for Cross-Lingual Learning
abstract
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig
ACL (1)10
2019 An Empirical Investigation of Structured Output Modeling for Graph-based Neural Dependency Parsing
abstract
In this paper, we investigate the aspect of structured output modeling for the state-ofthe-art graph-based neural dependency parser (Dozat and Manning, 2017).With evaluations on 14 treebanks, we empirically show that global output-structured models can generally obtain better performance, especially on the metric of sentence-level Complete Match.However, probably because neural models already learn good global views of the inputs, the improvement brought by structured output modeling is modest.
Zhisong Zhang, Xuezhe Ma, Eduard H. Hovy
ACL (1)2
2019 Cross-Lingual Dependency Parsing with Unlabeled Auxiliary Languages
abstract
Cross-lingual transfer learning has become an important weapon to battle the unavailability of annotated resources for low-resource languages.One of the fundamental techniques to transfer across languages is learning language-agnostic representations, in the form of word embeddings or contextual encodings.In this work, we propose to leverage unannotated sentences from auxiliary languages to help learning language-agnostic representations.Specifically, we explore adversarial training for learning contextual encoders that produce invariant representations across languages to facilitate cross-lingual transfer.We conduct experiments on cross-lingual dependency parsing where we train a dependency parser on a source language and transfer it to a wide range of target languages.Experiments on 28 target languages demonstrate that adversarial training significantly improves the overall transfer performances under several different settings.We conduct a careful analysis to evaluate the language-agnostic representations resulted from adversarial training.
Wasi Uddin Ahmad, Zhisong Zhang, Xuezhe Ma, Kai-Wei Chang 0001, Nanyun Peng 0001
CoNLL3
2019 FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow
abstract
Xuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xuezhe Ma, Chunting Zhou, Xian Li 0003, Graham Neubig, Eduard H. Hovy
EMNLP/IJCNLP (1)1
2019 Handling Syntactic Divergence in Low-resource Machine Translation
abstract
Chunting Zhou, Xuezhe Ma, Junjie Hu, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chunting Zhou, Xuezhe Ma, Junjie Hu 0001, Graham Neubig
EMNLP/IJCNLP (1)2
2019 MAE: Mutual Posterior-Divergence Regularization for Variational AutoEncoders
Xuezhe Ma, Chunting Zhou, Eduard H. Hovy
ICLR (Poster)1
2019 MaCow: Masked Convolutional Generative Flow
abstract
Flow-based generative models, conceptually attractive due to tractability of both the exact log-likelihood computation and latent-variable inference, and efficiency of both training and sampling, has led to a number of impressive empirical successes and spawned many advanced variants and theoretical investigations. Despite their computational efficiency, the density estimation performance of flow-based generative models significantly falls behind those of state-of-the-art autoregressive models. In this work, we introduce masked convolutional generative flow (MaCow), a simple yet effective architecture of generative flow using masked convolution. By restricting the local connectivity in a small kernel, MaCow enjoys the properties of fast and stable training, and efficient sampling, while achieving significant improvements over Glow for density estimation on standard image benchmarks, considerably narrowing the gap to autoregressive models.
Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy
NeurIPS1
2018 Stack-Pointer Networks for Dependency Parsing
abstract
We introduce a novel architecture for dependency parsing: stack-pointer networks (STACKPTR).Combining pointer networks (Vinyals et al., 2015) with an internal stack, the proposed model first reads and encodes the whole sentence, then builds the dependency tree top-down (from root-to-leaf) in a depth-first fashion.The stack tracks the status of the depthfirst search and the pointer networks select one child for the word at the top of the stack at each step.The STACKPTR parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the leftto-right restriction in classical transitionbased parsers.Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with O(n 2 ) time complexity.We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-theart performance on 21 of them.
Xuezhe Ma, Zecong Hu, Jingzhou Liu, Nanyun Peng 0001, Graham Neubig, Eduard H. Hovy
ACL (1)1
2017 An Interpretable Knowledge Transfer Model for Knowledge Base Completion
abstract
Knowledge bases are important resources for a variety of natural language processing tasks but suffer from incompleteness.We propose a novel embedding model, ITransF, to perform knowledge base completion.Equipped with a sparse attention mechanism, ITransF discovers hidden concepts of relations and transfer statistical strength through the sharing of concepts.Moreover, the learned associations between relations and concepts, which are represented by sparse attention vectors, can be interpreted easily.We evaluate ITransF on two benchmark datasets-WN18 and FB15k for knowledge base completion and obtains improvements on both the mean rank and Hits@10 metrics, over all baselines that do not use additional information.
Qizhe Xie, Xuezhe Ma, Zihang Dai, Eduard H. Hovy
ACL (1)2
2017 Dropout with Expectation-linear Regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, Eduard H. Hovy
ICLR (Poster)1
2017 Neural Probabilistic Model for Non-projective MST Parsing
abstract
In this paper, we propose a probabilistic parsing model that defines a proper conditional probability distribution over non-projective dependency trees for a given sentence, using neural representations as inputs. The neural network architecture is based on bi-directional LSTMCNNs, which automatically benefits from both word- and character-level representations, by using a combination of bidirectional LSTMs and CNNs. On top of the neural network, we introduce a probabilistic structured layer, defining a conditional log-linear model over non-projective trees. By exploiting Kirchhoff’s Matrix-Tree Theorem (Tutte, 1984), the partition functions and marginals can be computed efficiently, leading to a straightforward end-to-end model training procedure via back-propagation. We evaluate our model on 17 different datasets, across 14 different languages. Our parser achieves state-of-the-art parsing performance on nine datasets.
Xuezhe Ma, Eduard H. Hovy
IJCNLP(1)1
2016 Harnessing Deep Neural Networks with Logic Rules
abstract
Combining deep neural networks with structured logic rules is desirable to harness flexibility and reduce uninterpretability of the neural models.We propose a general framework capable of enhancing various types of neural networks (e.g., CNNs and RNNs) with declarative first-order logic rules.Specifically, we develop an iterative distillation method that transfers the structured information of logic rules into the weights of neural networks.We deploy the framework on a CNN for sentiment analysis, and an RNN for named entity recognition.With a few highly intuitive rules, we obtain substantial improvements and achieve state-of-the-art or comparable results to previous best-performing systems.
Zhiting Hu, Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy, Eric P. Xing
ACL (1)2
2016 End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
abstract
State-of-the-art sequence labeling systems traditionally require large amounts of taskspecific knowledge in the form of handcrafted features and data pre-processing.In this paper, we introduce a novel neutral network architecture that benefits from both word-and character-level representations automatically, by using combination of bidirectional LSTM, CNN and CRF.Our system is truly end-to-end, requiring no feature engineering or data preprocessing, thus making it applicable to a wide range of sequence labeling tasks.We evaluate our system on two data sets for two sequence labeling tasks -Penn Treebank WSJ corpus for part-of-speech (POS) tagging and CoNLL 2003 corpus for named entity recognition (NER).We obtain state-of-the-art performance on both datasets -97.55% accuracy for POS tagging and 91.21% F1 for NER.
Xuezhe Ma, Eduard H. Hovy
ACL (1)1
2016 Unsupervised Ranking Model for Entity Coreference Resolution
abstract
Coreference resolution is one of the first stages in deep language understanding and its importance has been well recognized in the natural language processing community. In this paper, we propose a generative, unsupervised ranking model for entity coreference resolution by introducing resolution mode variables. Our unsupervised system achieves 58.44% F1 score of the CoNLL metric on the English data from the CoNLL-2012 shared task (Pradhan et al., 2012), outperforming the Stanford deterministic system (Lee et al., 2013) by 3.01%.
Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy
HLT-NAACL1
2015 Efficient Inner-to-outer Greedy Algorithm for Higher-order Labeled Dependency Parsing
abstract
Many NLP systems use dependency parsers as critical components.Jonit learning parsers usually achieve better parsing accuracies than two-stage methods.However, classical joint parsing algorithms significantly increase computational complexity, which makes joint learning impractical.In this paper, we proposed an efficient dependency parsing algorithm that is capable of capturing multiple edge-label features, while maintaining low computational complexity.We evaluate our parser on 14 different languages.Our parser consistently obtains more accurate results than three baseline systems and three popular, off-the-shelf parsers.
Xuezhe Ma, Eduard H. Hovy
EMNLP1
2014 Unsupervised Dependency Parsing with Transferring Distribution via Parallel Guidance and Entropy Regularization
abstract
We present a novel approach for induc-ing unsupervised dependency parsers for languages that have no labeled training data, but have translated text in a resource-rich language. We train probabilistic pars-ing models for resource-poor languages by transferring cross-lingual knowledge from resource-rich language with entropy reg-ularization. Our method can be used as a purely monolingual dependency parser, requiring no human translations for the test data, thus making it applicable to a wide range of resource-poor languages. We perform experiments on three Data sets — Version 1.0 and version 2.0 of Google Universal Dependency Treebanks and Treebanks from CoNLL shared-tasks, across ten languages. We obtain state-of-the art performance of all the three data sets when compared with previously studied unsupervised and projected pars-ing systems. 1
Xuezhe Ma
ACL (1)1