Yangfeng Ji

dblp:94/8323 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0002-7793-486XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Consistency of LLMs to Comparative Statements in Mathematical Reasoning Tasks
Aidan San, Daniel Juyoung Son, Xiaodong Liu 0003, Yangfeng Ji
LREC4
2025 Unsupervised Concept Vector Extraction for Bias Control in LLMs
abstract
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases.Various strategies have been proposed to mitigate these biases, but most work studies biases as a black-box problem without considering how concepts are represented within the model.We adapt techniques from representation engineering to study how the concept of "gender" is represented within LLMs.We introduce a new method that extracts concept representations via probability weighting without labeled data and efficiently selects a steering vector for measuring and manipulating the model's representation.We develop a projection-based method that enables precise steering of model predictions and demonstrate its effectiveness in mitigating gender bias in LLMs and show that it also generalizes to racial bias. 1
Hannah Cyberey, Yangfeng Ji, David Evans 0001
EMNLP2
2025 The Good, the Bad, and the Debatable: A Survey on the Impacts of Data for In-Context Learning
abstract
In-context learning is an emergent learning paradigm that enables an LLM to learn an unseen task by seeing a number of demonstrations in the context window.The quality of the demonstrations is of paramount importance as 1) context window size limitations restrict the number of demonstrations that can be presented to the model, and 2) the model must identify the task and potentially learn new, unseen input-output mappings from the limited demonstration set.An increasing body of work has also shown the sensitivity of predictions to perturbations on the demonstration set.Given this importance, this work presents a survey on the current literature pertaining to the relationship between data and in-context learning.We present our survey in three parts: the "good" -qualities that are desirable when selecting demonstrations, the "bad" -qualities of demonstrations that can negatively impact the model, as well as issues that can arise in presenting demonstrations, and the "debatable" -qualities of demonstrations with mixed results or factors modulating data impacts.
Stephanie Schoch, Yangfeng Ji
EMNLP2
2025 SelectFormer in Data Markets: Privacy-Preserving and Efficient Data Selection for Transformers with Multi-Party Computation
abstract
Critical to a free data market is $ \textit{private data selection}$, i.e. the model owner selects and then appraises training data from the data owner before both parties commit to a transaction. To keep the data and model private, this process shall evaluate the target model to be trained over Multi-Party Computation (MPC). While prior work suggests that evaluating Transformer-based models over MPC is prohibitively expensive, this paper makes it practical for the purpose of data selection. Our contributions are three: (1) a new pipeline for private data selection over MPC; (2) emulating high-dimensional nonlinear operators with low-dimension MLPs, which are trained on a small sample of the data of interest; (3) scheduling MPC in a parallel, multiphase fashion. We evaluate our method on diverse Transformer models and NLP/CV benchmarks. Compared to directly evaluating the target model over MPC, our method reduces the delay from thousands of hours to tens of hours, while only seeing around 0.20% accuracy degradation from training with the selected data.
Xu Ouyang, Felix Xiaozhu Lin, Yangfeng Ji
ICLR3
2025 Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding
abstract
Computational story understanding is a crucial but under-explored area of AI, hampered by a lack of suitable datasets. To address this, we collect, preprocess and publicly release SYMON (Synopses of Movie Narratives), a new video-language dataset containing 5,193 human-narrated, short movie summary videos sourced from YouTube. SYMON features naturalistic storytelling videos for human audiences made by human creators. Compared to existing movie story datasets, the videos in SYMON are shorter yet provide higher coverage of key story events, making it ideal for computational story understanding. We establish benchmarks on story video-text alignment and story video narration generation, demonstrating significant performance improvements when models are trained on SYMON. These results underscore the value of SYMON for advancing research in vision-language story understanding and generation.
Qin Chao, Yangfeng Ji, Boyang Li 0001
ICME3
2025 IrrMap: A Large-Scale Comprehensive Dataset for Irrigation Method Mapping
abstract
We introduce IrrMap, the first large-scale dataset (1.1 million patches) for irrigation method mapping across regions. IrrMap consists of multi-resolution satellite imagery from LandSat and Sentinel, along with key auxiliary data such as crop type, land use, and vegetation indices. The dataset spans 1,668,899 farms and 11,443,492 acres across multiple western U.S. states from 2013 to 2023, providing a rich and diverse foundation for irrigation analysis and ensuring geospatial alignment and quality control. The dataset is ML-ready, with standardized 224×224 GeoTIFF patches, the multiple input modalities, carefully chosen train-test-split data, and accompanying dataloaders for seamless deep learning model training and benchmarking in irrigation mapping. The dataset is also accompanied by a complete pipeline for dataset generation, enabling researchers to extend IrrMap to new regions for irrigation data collection or adapt it with minimal effort for other similar applications in agricultural and geospatial analysis. We also analyze the irrigation method distribution across crop groups, spatial irrigation patterns (using Shannon diversity indices), and irrigated area variations for both LandSat and Sentinel, providing insights into regional and resolution-based differences. To promote further exploration, we openly release IrrMap, along with the derived datasets, benchmark models, and pipeline code, through a GitHub repository: https://github.com/Nibir088/IrrMap and Data repository: https://huggingface.co/Nibir/IrrMap, providing comprehensive documentation and implementation details.
Nibir Chandra Mandal, Oishee Bintey Hoque, Abhijin Adiga, Samarth Swarup, Mandy L. Wilson, Lu Feng 0001, Yangfeng Ji, Miaomiao Zhang 0002, Geoffrey C. Fox, Madhav V. Marathe
KDD (2)7
2025 In-Context Learning (and Unlearning) of Length Biases
abstract
Stephanie Schoch, Yangfeng Ji. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Stephanie Schoch, Yangfeng Ji
NAACL (Long Papers)2
2023 Improving Interpretability via Explicit Word Interaction Graph Layer
abstract
Recent NLP literature has seen growing interest in improving model interpretability. Along this direction, we propose a trainable neural network layer that learns a global interaction graph between words and then selects more informative words using the learned word interactions. Our layer, we call WIGRAPH, can plug into any neural network-based NLP text classifiers right after its word embedding layer. Across multiple SOTA NLP models and various NLP datasets, we demonstrate that adding the WIGRAPH layer substantially improves NLP models' interpretability and enhances models' prediction performance at the same time.
Arshdeep Sekhon, Aman Shrivastava, Zhe Wang 0025, Yangfeng Ji, Yanjun Qi
AAAI5
2023 REV: Information-Theoretic Evaluation of Free-Text Rationales
abstract
Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, Swabha Swayamdipta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Faeze Brahman, Xiang Ren 0001, Yangfeng Ji, Yejin Choi 0001, Swabha Swayamdipta
ACL (1)4
2023 Efficient NLP Model Finetuning via Multistage Data Filtering
abstract
As model finetuning is central to the modern NLP, we set to maximize its efficiency. Motivated by redundancy in training examples and the sheer sizes of pretrained models, we exploit a key opportunity: training only on important data. To this end, we set to filter training examples in a streaming fashion, in tandem with training the target model. Our key techniques are two: (1) automatically determine a training loss threshold for skipping backward training passes; (2) run a meta predictor for further skipping forward training passes. We integrate the above techniques in a holistic, three-stage training pro- cess. On a diverse set of benchmarks, our method reduces the required training examples by up to 5.3× and training time by up to 6.8×, while only seeing minor accuracy degradation. Our method is effective even for training one epoch, where each training example is encountered only once. It is simple to implement and is compatible with the existing finetuning techniques. Code is available at: https://github.com/xo28/efficient-NLP-multistage-training
Xu Ouyang, Shahina Mohd Azam Ansari, Felix Xiaozhu Lin, Yangfeng Ji
IJCAI4
2022 Adversarial Training for Improving Model Robustness? Look at Both Prediction and Interpretation
abstract
Neural language models show vulnerability to adversarial examples which are semantically similar to their original counterparts with a few words replaced by their synonyms. A common way to improve model robustness is adversarial training which follows two steps—collecting adversarial examples by attacking a target model, and fine-tuning the model on the augmented dataset with these adversarial examples. The objective of traditional adversarial training is to make a model produce the same correct predictions on an original/adversarial example pair. However, the consistency between model decision-makings on two similar texts is ignored. We argue that a robust model should behave consistently on original/adversarial example pairs, that is making the same predictions (what) based on the same reasons (how) which can be reflected by consistent interpretations. In this work, we propose a novel feature-level adversarial training method named FLAT. FLAT aims at improving model robustness in terms of both predictions and interpretations. FLAT incorporates variational word masks in neural networks to learn global word importance and play as a bottleneck teaching the model to make predictions based on important words. FLAT explicitly shoots at the vulnerability problem caused by the mismatch between model understandings on the replaced words and their synonyms in original/adversarial example pairs by regularizing the corresponding global word importance scores. Experiments show the effectiveness of FLAT in improving the robustness with respect to both predictions and interpretations of four neural network models (LSTM, CNN, BERT, and DeBERTa) to two adversarial attacks on four text classification tasks. The models trained via FLAT also show better robustness than baseline models on unforeseen adversarial examples across different attacks.
Yangfeng Ji
AAAI2
2022 Balanced Adversarial Training: Balancing Tradeoffs between Fickleness and Obstinacy in NLP Models
abstract
Traditional (fickle) adversarial examples involve finding a small perturbation that does not change an input's true label but confuses the classifier into outputting a different prediction.Conversely, obstinate adversarial examples occur when an adversary finds a small perturbation that preserves the classifier's prediction but changes the true label of an input.Adversarial training and certified robust training have shown some effectiveness in improving the robustness of machine learnt models to fickle adversarial examples.We show that standard adversarial training methods focused on reducing vulnerability to fickle adversarial examples may make a model more vulnerable to obstinate adversarial examples, with experiments for both natural language inference and paraphrase identification tasks.To counter this phenomenon, we introduce Balanced Adversarial Training, which incorporates contrastive learning to increase robustness against both fickle and obstinate adversarial examples.
Hannah Cyberey, Yangfeng Ji, David Evans 0001
EMNLP2
2022 FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act Flows
abstract
Despite recent progress in open-domain dialogue evaluation, how to develop automatic metrics remains an open problem.We explore the potential of dialogue evaluation featuring dialog act information, which was hardly explicitly modeled in previous methods.However, defined at the utterance level in general, dialog act is of coarse granularity, as an utterance can contain multiple segments possessing different functions.Hence, we propose segment act, an extension of dialog act from utterance level to segment level, and crowdsource a largescale dataset for it.To utilize segment act flows, sequences of segment acts, for evaluation, we develop the first consensus-based dialogue evaluation framework, FlowEval.This framework provides a reference-free approach for dialog evaluation by finding pseudo-references.Extensive experiments against strong baselines on three benchmark datasets demonstrate the effectiveness and other desirable characteristics of our FlowEval, pointing out a potential path for better dialogue evaluation.
Jianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji, Dong Yu 0001, Michael R. Lyu, Liwei Wang 0009
EMNLP4
2022 CS-Shapley: Class-wise Shapley Values for Data Valuation in Classification
abstract
Data valuation, or the valuation of individual datum contributions, has seen growing interest in machine learning due to its demonstrable efficacy for tasks such as noisy label detection. In particular, due to the desirable axiomatic properties, several Shapley value approximations have been proposed. In these methods, the value function is usually defined as the predictive accuracy over the entire development set. However, this limits the ability to differentiate between training instances that are helpful or harmful to their own classes. Intuitively, instances that harm their own classes may be noisy or mislabeled and should receive a lower valuation than helpful instances. In this work, we propose CS-Shapley, a Shapley value with a new value function that discriminates between training instances’ in-class and out-of-class contributions. Our theoretical analysis shows the proposed value function is (essentially) the unique function that satisfies two desirable properties for evaluating data values in classification. Further, our experiments on two benchmark evaluation tasks (data removal and noisy label detection) and four classifiers demonstrate the effectiveness of CS-Shapley over existing methods. Lastly, we evaluate the “transferability” of data values estimated from one classifier to others, and our results suggest Shapley-based data valuation is transferable for application across different models.
Stephanie Schoch, Yangfeng Ji
NeurIPS3
2021 HittER: Hierarchical Transformers for Knowledge Graph Embeddings
abstract
This paper examines the challenging problem of learning representations of entities and relations in a complex multi-relational knowledge graph.We propose HittER, a Hierarchical Transformer model to jointly learn Entityrelation composition and Relational contextualization based on a source entity's neighborhood.Our proposed model consists of two different Transformer blocks: the bottom block extracts features of each entity-relation pair in the local neighborhood of the source entity and the top block aggregates the relational information from outputs of the bottom block.We further design a masked entity prediction task to balance information from the relational context and the source entity itself.Experimental results show that HittER achieves new stateof-the-art results on multiple link prediction datasets.We additionally propose a simple approach to integrate HittER into BERT and demonstrate its effectiveness on two Freebase factoid question answering datasets.
Sanxing Chen, Xiaodong Liu 0003, Jianfeng Gao 0001, Jian Jiao 0007, Ruofei Zhang, Yangfeng Ji
EMNLP (1)6
2021 Contextualizing Variation in Text Style Transfer Datasets
abstract
Text style transfer involves rewriting the content of a source sentence in a target style.Despite there being a number of style tasks with available data, there has been limited systematic discussion of how text style datasets relate to each other.This understanding, however, is likely to have implications for selecting multiple data sources for model training.While it is prudent to consider inherent stylistic properties when determining these relationships, we also must consider how a style is realized in a particular dataset.In this paper, we conduct several empirical analyses of existing text style datasets.Based on our results, we propose a categorization of stylistic and dataset properties to consider when utilizing or comparing text style datasets.
Stephanie Schoch, Wanyu Du, Yangfeng Ji
INLG3
2021 Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
abstract
Hanjie Chen, Song Feng, Jatin Ganhotra, Hui Wan, Chulaka Gunasekara, Sachindra Joshi, Yangfeng Ji. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Song Feng 0002, Jatin Ganhotra, Hui Wan 0001, R. Chulaka Gunasekara, Sachindra Joshi, Yangfeng Ji
NAACL-HLT7
2020 Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
abstract
Generating explanations for neural networks has become crucial for their applications in real-world with respect to reliability and trustworthiness.In natural language processing, existing methods usually provide important features which are words or phrases selected from an input text as an explanation, but ignore the interactions between them.It poses challenges for humans to interpret an explanation and connect it to model prediction.In this work, we build hierarchical explanations by detecting feature interactions.Such explanations visualize how words and phrases are combined at different levels of the hierarchy, which can help users understand the decision-making of blackbox models.The proposed method is evaluated with three neural text classifiers (LSTM, CNN, and BERT) on two benchmark datasets, via both automatic and human evaluations.Experiments show the effectiveness of the proposed method in providing explanations that are both faithful to models and interpretable to humans.
Guangtao Zheng, Yangfeng Ji
ACL3
2020 A Tale of Two Linkings: Dynamically Gating between Schema Linking and Structural Linking for Text-to-SQL Parsing
abstract
In Text-to-SQL semantic parsing, selecting the correct entities (tables and columns) for the generated SQL query is both crucial and challenging; the parser is required to connect the natural language (NL) question and the SQL query to the structured knowledge in the database.We formulate two linking processes to address this challenge: schema linking which links explicit NL mentions to the database and structural linking which links the entities in the output SQL with their structural relationships in the database schema.Intuitively, the effectiveness of these two linking processes changes based on the entity being generated, thus we propose to dynamically choose between them using a gating mechanism.Integrating the proposed method with two graph neural network-based semantic parsers together with BERT representations demonstrates substantial gains in parsing accuracy on the challenging Spider dataset.Analyses show that our proposed method helps to enhance the structure of the model output when generating complicated SQL queries and offers more explainable predictions.
Sanxing Chen, Aidan San, Xiaodong Liu 0003, Yangfeng Ji
COLING4
2020 Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
abstract
To build an interpretable neural text classifier, most of the prior work has focused on designing inherently interpretable models or finding faithful explanations.A new line of work on improving model interpretability has just started, and many existing methods require either prior information or human annotations as additional inputs in training.To address this limitation, we propose the variational word mask (VMASK) method to automatically learn task-specific important words and reduce irrelevant information on classification, which ultimately improves the interpretability of model predictions.The proposed method is evaluated with three neural text classifiers (CNN, LSTM, and BERT) on seven benchmark text classification datasets.Experiments show the effectiveness of VMASK in improving both model prediction accuracy and interpretability.
Yangfeng Ji
EMNLP (1)2
2019 An Empirical Comparison on Imitation Learning and Reinforcement Learning for Paraphrase Generation
abstract
Wanyu Du, Yangfeng Ji. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wanyu Du, Yangfeng Ji
EMNLP/IJCNLP (1)2
2018 Creative Writing with a Machine in the Loop: Case Studies on Slogans and Stories
abstract
As the quality of natural language generated by artificial intelligence systems improves, writing interfaces can support interventions beyond grammar-checking and spell-checking, such as suggesting content to spark new ideas. To explore the possibility of machine-in-the-loop creative writing, we performed two case studies using two system prototypes, one for short story writing and one for slogan writing. Participants in our studies were asked to write with a machine in the loop or alone (control condition). They assessed their writing and experience through surveys and an open-ended interview. We collected additional assessments of the writing from Amazon Mechanical Turk crowdworkers. Our findings indicate that participants found the process fun and helpful and could envision use cases for future systems. At the same time, machine suggestions do not necessarily lead to better written artifacts. We therefore suggest novel natural language models and design choices that may better support creative writing.
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, Noah A. Smith
IUI4
2018 Neural Text Generation in Stories Using Entity Representations as Context
abstract
Elizabeth Clark, Yangfeng Ji, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Elizabeth Clark, Yangfeng Ji, Noah A. Smith
NAACL-HLT2
2017 Neural Discourse Structure for Text Categorization
abstract
We show that discourse structure, as defined by Rhetorical Structure Theory and provided by an existing discourse parser, benefits text categorization.Our approach uses a recursive neural network and a newly proposed attention mechanism to compute a representation of the text that focuses on salient content, from the perspective of both RST and the task.Experiments consider variants of the approach and illustrate its strengths and weaknesses.
Yangfeng Ji, Noah A. Smith
ACL (1)1
2017 Dynamic Entity Representations in Neural Language Models
abstract
Understanding a long document requires tracking how entities are introduced and evolve over time.We present a new type of language model, ENTITYNLM, that can explicitly model entities, dynamically update their representations, and contextually generate their mentions.Our model is generative and flexible; it can model an arbitrary number of entities in context while generating each entity mention at an arbitrary length.In addition, it can be used for several different tasks such as language modeling, coreference resolution, and entity prediction.Experimental results with all these tasks demonstrate that our model consistently outperforms strong baselines and prior work.
Yangfeng Ji, Chenhao Tan, Sebastian Martschat, Yejin Choi 0001, Noah A. Smith
EMNLP1
2016 A Latent Variable Recurrent Neural Network for Discourse-Driven Language Models
abstract
This paper presents a novel latent variable recurrent neural network architecture for jointly modeling sequences of words and (possibly latent) discourse relations between adjacent sentences.A recurrent neural network generates individual words, thus reaping the benefits of discriminatively-trained vector representations.The discourse relations are represented with a latent variable, which can be predicted or marginalized, depending on the task.The resulting model can therefore employ a training objective that includes not only discourse relation classification, but also word prediction.As a result, it outperforms state-ofthe-art alternatives for two tasks: implicit discourse relation classification in the Penn Discourse Treebank, and dialog act classification in the Switchboard corpus.Furthermore, by marginalizing over latent discourse relations at test time, we obtain a discourse informed language model, which improves over a strong LSTM baseline.
Yangfeng Ji, Gholamreza Haffari, Jacob Eisenstein
HLT-NAACL1
2015 Better Document-level Sentiment Analysis from RST Discourse Parsing
abstract
Discourse structure is the hidden link between surface features and document-level properties, such as sentiment polarity.We show that the discourse analyses produced by Rhetorical Structure Theory (RST) parsers can improve document-level sentiment analysis, via composition of local information up the discourse tree.First, we show that reweighting discourse units according to their position in a dependency representation of the rhetorical structure can yield substantial improvements on lexicon-based sentiment analysis.Next, we present a recursive neural network over the RST structure, which offers significant improvements over classificationbased methods.
Parminder Bhatia, Yangfeng Ji, Jacob Eisenstein
EMNLP2
2015 Closing the Gap: Domain Adaptation from Explicit to Implicit Discourse Relations
abstract
Many discourse relations are explicitly marked with discourse connectives, and these examples could potentially serve as a plentiful source of training data for recognizing implicit discourse relations.However, there are important linguistic differences between explicit and implicit discourse relations, which limit the accuracy of such an approach.We account for these differences by applying techniques from domain adaptation, treating implicitly and explicitly-marked discourse relations as separate domains.The distribution of surface features varies across these two domains, so we apply a marginalized denoising autoencoder to induce a dense, domain-general representation.The label distribution is also domain-specific, so we apply a resampling technique that is similar to instance weighting.In combination with a set of automatically-labeled data, these improvements eliminate more than 80% of the transfer loss incurred by training an implicit discourse relation classifier on explicitly-marked discourse relations.
Yangfeng Ji, Jacob Eisenstein
EMNLP1
2015 A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
abstract
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao 0001, William B. Dolan
HLT-NAACL5
2015 One Vector is Not Enough: Entity-Augmented Distributed Semantics for Discourse Relations
abstract
Discourse relations bind smaller linguistic units into coherent texts. Automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked arguments. A more subtle challenge is that it is not enough to represent the meaning of each argument of a discourse relation, because the relation may depend on links between lowerlevel components, such as entity mentions. Our solution computes distributed meaning representations for each discourse argument by composition up the syntactic parse tree. We also perform a downward compositional pass to capture the meaning of coreferent entity mentions. Implicit discourse relations are then predicted from these two representations, obtaining substantial improvements on the Penn Discourse Treebank.
Yangfeng Ji, Jacob Eisenstein
Trans. Assoc. Comput. Linguistics1
2014 Representation Learning for Text-level Discourse Parsing
abstract
Text-level discourse parsing is notoriously difficult, as distinctions between discourse relations require subtle semantic judg-ments that are not easily captured using standard features. In this paper, we present a representation learning approach, in which we transform surface features into a latent space that facilitates RST dis-course parsing. By combining the machin-ery of large-margin transition-based struc-tured prediction with representation learn-ing, our method jointly learns to parse dis-course while at the same time learning a discourse-driven projection of surface fea-tures. The resulting shift-reduce discourse parser obtains substantial improvements over the previous state-of-the-art in pre-dicting relations and nuclearity on the RST Treebank. 1
Yangfeng Ji, Jacob Eisenstein
ACL (1)1
2014 A variational Bayesian model for user intent detection
abstract
Intent detectors in state-of-the-art spoken language understanding systems are often trained with a small number of manually annotated examples collected from the application domain. Search query logs provide a large number of unlabeled queries that would be beneficial to improve such supervised classification. Furthermore, the contents of user queries as well as the clicked URLs provide information about user's intent. In this paper, we propose a variational Bayesian approach for modeling latent intents of user queries and clicked URLs when available. We use this model to enhance supervised intent classification of user queries from conversational interactions. Experiments were run with large volumes of search queries and show significant improvements over state-of-the-art systems.
Yangfeng Ji, Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür
ICASSP1
2014 Extracting Lexically Divergent Paraphrases from Twitter
abstract
We present MultiP (Multi-instance Learning Paraphrase Model), a new model suited to identify paraphrases within the short messages on Twitter. We jointly model paraphrase relations between word and sentence pairs and assume only sentence-level annotations during learning. Using this principled latent variable model alone, we achieve the performance competitive with a state-of-the-art method which combines a latent space model with a feature-based supervised classifier. Our model also captures lexically divergent paraphrases that differ from yet complement previous methods; combining our model with previous work significantly outperforms the state-of-the-art. In addition, we present a novel annotation methodology that has allowed us to crowdsource a paraphrase corpus from Twitter. We make this new dataset available to the research community.
Wei Xu 0004, Alan Ritter, Chris Callison-Burch, William B. Dolan, Yangfeng Ji
Trans. Assoc. Comput. Linguistics5
2013 Discriminative Improvements to Distributional Sentence Similarity
abstract
Matrix and tensor factorization have been applied to a number of semantic relatedness tasks, including paraphrase identification.The key idea is that similarity in the latent space implies semantic relatedness.We describe three ways in which labeled data can improve the accuracy of these approaches on paraphrase classification.First, we design a new discriminative term-weighting metric called TF-KLD, which outperforms TF-IDF.Next, we show that using the latent representation from matrix factorization as features in a classification algorithm substantially improves accuracy.Finally, we combine latent features with fine-grained n-gram overlap features, yielding performance that is 3% more accurate than the prior state-of-the-art.
Yangfeng Ji, Jacob Eisenstein
EMNLP1
2010 CDP Mixture Models for Data Clustering
abstract
In Dirichlet process (DP) mixture models, the number of components is implicitly determined by the sampling parameters of Dirichlet process. However, this kind of models usually produces lots of small mixture components when modeling real-world data, especially high-dimensional data. In this paper, we propose a new class of Dirichlet process mixture models with some constrained principles, named constrained Dirichlet process (CDP) mixture models. Based on general DP mixture models, we add a resampling step to obtain latent parameters. In this way, CDP mixture models can suppress noise and generate the compact patterns of the data. Experimental results on data clustering show the remarkable performance of the CDP mixture models.
Yangfeng Ji, Tong Lin 0002, Hongbin Zha
ICPR1
2009 Mahalanobis Distance Based Non-negative Sparse Representation for Face Recognition
abstract
Sparse representation for machine learning has been exploited in past years. Several sparse representation based classification algorithms have been developed for some applications, for example, face recognition. In this paper, we propose an improved sparse representation based classification algorithm. Firstly, for a discriminative representation, a non-negative constraint of sparse coefficient is added to sparse representation problem. Secondly, Mahalanobis distance is employed instead of Euclidean distance to measure the similarity between original data and reconstructed data. The proposed classification algorithm for face recognition has been evaluated under varying illumination and pose using standard face databases. The experimental results demonstrate that the performance of our algorithm is better than that of the up-to-date face recognition algorithm based on sparse representation.
Yangfeng Ji, Tong Lin 0002, Hongbin Zha
ICMLA1