EDBT 2026 Demo / reviewers in the wild / expert
Lifeng Jin
dblp:66/7607
· DBLP profile ↗
31ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMsabstractLarge language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs. Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2025 | Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language ModelsabstractLarge language models (LLMs) exhibit impressive natural language capabilities but suffer from hallucination – generating content ungrounded in the realities of training data. Recent work has focused on decoding techniques to improve factuality in decoding by leveraging LLMs’ hierarchical representation of factual knowledge, manipulating the predicted distributions at inference time. Current state-of-the-art approaches refine decoding by contrasting logits from a lower layer with the final layer to exploit information related factuality within the model forward procedure. However, such methods often assume the final layer is most reliable one and the lower layer selection process depends on it. In this work, we first propose logit extrapolation of critical token probabilities beyond the last layer for more accurate contrasting. We additionally employ layer-wise entropy-guided lower layer selection, decoupling the selection process from the final layer. Experiments demonstrate strong performance - surpassing state-of-the-art on multiple different datasets by large margins. Analyses show different kinds of prompts respond to different selection strategies. Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, Dong Yu 0001 |
COLING | 2 |
| 2024 | Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationabstractXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen Meng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Baolin Peng, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen M. Meng |
ACL (1) | 5 |
| 2024 | A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue GenerationabstractKnowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion. Xiangci Li, Linfeng Song, Lifeng Jin, Haitao Mi, Jessica Ouyang 0001, Dong Yu 0001 |
LREC/COLING | 3 |
| 2024 | The Trickle-down Impact of Reward Inconsistency on RLHFabstractStandard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and
appropriately adapt their reward assignments
--- and their impact on the downstream RLHF model.
In this paper, we visit a series of research questions relevant to RM inconsistency:
(1) How can we measure the consistency of reward models?
(2) How consistent are the existing RMs and how can we improve them?
(3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training?
We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM.
Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency
through extrapolation during the RM training and inference stage, respectively.
We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process. Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001 |
ICLR | 4 |
| 2024 | Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingabstractDespite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Self-correction and self-learning emerge as viable solutions, employing strategies that allow LLMs to refine their outputs and learn from self-assessed rewards. Yet, the efficacy of LLMs in self-refining its response, particularly in complex reasoning and planning task, remains dubious. In this paper, we introduce AlphaLLM for the self-improvements of LLMs, which integrates Monte Carlo Tree Search (MCTS) with LLMs to establish a self-improving loop, thereby enhancing the capabilities of LLMs without additional annotations. Drawing inspiration from the success of AlphaGo, AlphaLLM addresses the unique challenges of combining MCTS with LLM for self-improvement, including data scarcity, the vastness search spaces of language tasks, and the subjective nature of feedback in language tasks. AlphaLLM is comprised of prompt synthesis component, an efficient MCTS approach tailored for language tasks, and a trio of critic models for precise feedback. Our experimental results in mathematical reasoning tasks demonstrate that AlphaLLM significantly enhances the performance of LLMs without additional annotations, showing the potential for self-improvement in LLMs. Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu 0001, Haitao Mi, Dong Yu 0001 |
NeurIPS | 4 |
| 2024 | A randomized prospective study of a hybrid rule- and data-driven virtual patientabstractAbstract Randomized prospective studies represent the gold standard for experimental design. In this paper, we present a randomized prospective study to validate the benefits of combining rule-based and data-driven natural language understanding methods in a virtual patient dialogue system. The system uses a rule-based pattern matching approach together with a machine learning (ML) approach in the form of a text-based convolutional neural network, combining the two methods with a simple logistic regression model to choose between their predictions for each dialogue turn. In an earlier, retrospective study, the hybrid system yielded a nearly 50% error reduction on our initial data, in part due to the differential performance between the two methods as a function of label frequency. Given these gains, and considering that our hybrid approach is unique among virtual patient systems, we compare the hybrid system to the rule-based system by itself in a randomized prospective study. We evaluate 110 unique medical student subjects interacting with the system over 5,296 conversation turns, to verify whether similar gains are observed in a deployed system. This prospective study broadly confirms the findings from the earlier one but also highlights important deficits in our training data. The hybrid approach still improves over either rule-based or ML approaches individually, even handling unseen classes with some success. However, we observe that live subjects ask more out-of-scope questions than expected. To better handle such questions, we investigate several modifications to the system combination component. These show significant overall accuracy improvements and modest F1 improvements on out-of-scope queries in an offline evaluation. We provide further analysis to characterize the difficulty of the out-of-scope problem that we have identified, as well as to suggest future improvements over the baseline we establish here. Adam Stiff, Michael White 0001, Eric Fosler-Lussier, Lifeng Jin, Evan Jaffe, Douglas Danforth |
Nat. Lang. Eng. | 4 |
| 2023 | SafeConv: Explaining and Correcting Conversational Unsafe BehaviorabstractOne of the main challenges open-domain endto-end dialogue systems, or chatbots, face is the prevalence of unsafe behavior, such as toxic languages and harmful suggestions.However, existing dialogue datasets do not provide enough annotation to explain and correct such unsafe behavior.In this work, we construct a new dataset called SAFECONV for the research of conversational safety: (1) Besides the utterancelevel safety labels, SAFECONV also provides unsafe spans in an utterance, information able to indicate which words contribute to the detected unsafe behavior; (2) SAFECONV provides safe alternative responses to continue the conversation when unsafe behavior detected, guiding the conversation to a gentle trajectory.By virtue of the comprehensive annotation of SAFECONV, we benchmark three powerful models for the mitigation of conversational unsafe behavior, including a checker to detect unsafe utterances, a tagger to extract unsafe spans, and a rewriter to convert an unsafe response to a safe version.Moreover, we explore the huge benefits brought by combining the models for explaining the emergence of unsafe behavior and detoxifying chatbots.Experiments show that the detected unsafe behavior could be well explained with unsafe spans and popular chatbots could be detoxified by a huge extent.The dataset is available at https://github.com/mianzhang/SafeConv. Lifeng Jin, Linfeng Song, Haitao Mi, Wenliang Chen, Dong Yu 0001 |
ACL (1) | 2 |
| 2023 | How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation MethodabstractWenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
EACL | 2 |
| 2023 | Friend-training: Learning from Models of Different but Related TasksabstractCurrent self-training methods such as standard self-training, co-training, tri-training, and others often focus on improving model performance on a single task, utilizing differences in input features, model architectures, and training processes.However, many tasks in natural language processing are about different but related aspects of language, and models trained for one task can be great teachers for other related tasks.In this work, we propose friendtraining, a cross-task self-training framework, where models trained to do different tasks are used in an iterative training, pseudo-labeling, and retraining process to help each other for better selection of pseudo-labels.With two dialogue understanding tasks, conversational semantic role labeling and dialogue rewriting, chosen for a case study, we show that the models trained with the friend-training framework achieve the best performance compared to strong baselines. Lifeng Jin, Linfeng Song, Haitao Mi, Xiabing Zhou, Dong Yu 0001 |
EACL | 2 |
| 2023 | Discover, Explain, Improve: An Automatic Slice Detection Benchmark for Natural Language ProcessingabstractAbstract Pretrained natural language processing (NLP) models have achieved high overall performance, but they still make systematic errors. Instead of manual error analysis, research on slice detection models (SDMs), which automatically identify underperforming groups of datapoints, has caught escalated attention in Computer Vision for both understanding model behaviors and providing insights for future model training and designing. However, little research on SDMs and quantitative evaluation of their effectiveness have been conducted on NLP tasks. Our paper fills the gap by proposing a benchmark named “Discover, Explain, Improve (DEIm)” for classification NLP tasks along with a new SDM Edisa. Edisa discovers coherent and underperforming groups of datapoints; DEIm then unites them under human-understandable concepts and provides comprehensive evaluation tasks and corresponding quantitative metrics. The evaluation in DEIm shows that Edisa can accurately select error-prone datapoints with informative semantic features that summarize error patterns. Detecting difficult datapoints directly boosts model performance without tuning any original model parameters, showing that discovered slices are actionable for users.1 Wenyue Hua, Lifeng Jin, Linfeng Song, Haitao Mi, Dong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | OpenFact: Factuality Enhanced Open Knowledge ExtractionabstractAbstract We focus on the factuality property during the extraction of an OpenIE corpus named OpenFact, which contains more than 12 million high-quality knowledge triplets. We break down the factuality property into two important aspects—expressiveness and groundedness—and we propose a comprehensive framework to handle both aspects. To enhance expressiveness, we formulate each knowledge piece in OpenFact based on a semantic frame. We also design templates, extra constraints, and adopt human efforts so that most OpenFact triplets contain enough details. For groundedness, we require the main arguments of each triplet to contain linked Wikidata1 entities. A human evaluation suggests that the OpenFact triplets are much more accurate and contain denser information compared to OPIEC-Linked (Gashteovski et al., 2019), one recent high-quality OpenIE corpus grounded to Wikidata. Further experiments on knowledge base completion and knowledge base question answering show the effectiveness of OpenFact over OPIEC-Linked as supplementary knowledge to Wikidata as the major KG. Linfeng Song, Ante Wang, Xiaoman Pan, Hongming Zhang 0009, Dian Yu 0001, Lifeng Jin, Haitao Mi, Jinsong Su, Yue Zhang 0004, Dong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2023 | D$^{2}$PSG: Multi-Party Dialogue Discourse Parsing as Sequence GenerationabstractConversational discourse analysis aims to extract the interactions between dialogue turns, which is crucial for modeling complex multi-party dialogues. As the benchmarks are still limited in size and human annotations are costly, the current standard approaches apply pretrained language models, but they still require randomly initialized classifiers to make predictions. These classifiers usually require massive data to work smoothly with the pretrained encoder, causing severe data hunger issue. We propose two convenient strategies to formulate this task as a sequence generation problem, where classifier decisions are carefully converted into sequence of tokens. We then adopt a pretrained T5 1 model to solve this task so that no parameters are randomly initialized. We also leverage the descriptions of the discourse relations to help model understand their meanings. Experiments on two popular benchmarks show that our approach outperforms previous state-of-the-art models by a large margin, and it is also more robust in zero-shot and few-shot settings. Ante Wang, Linfeng Song, Lifeng Jin, Junfeng Yao, Haitao Mi, Chen Lin 0001, Jinsong Su, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Hierarchical Context Tagging for Utterance RewritingabstractUtterance rewriting aims to recover coreferences and omitted information from the latest turn of a multi-turn dialogue. Recently, methods that tag rather than linearly generate sequences have proven stronger in both in- and out-of-domain rewriting settings. This is due to a tagger's smaller search space as it can only copy tokens from the dialogue context. However, these methods may suffer from low coverage when phrases that must be added to a source utterance cannot be covered by a single context span. This can occur in languages like English that introduce tokens such as prepositions into the rewrite for grammaticality. We propose a hierarchical context tagger (HCT) that mitigates this issue by predicting slotted rules (e.g., "besides _") whose slots are later filled with context spans. HCT (i) tags the source string with token-level edit actions and slotted rules and (ii) fills in the resulting rule slots with spans from the dialogue context. This rule tagging allows HCT to add out-of-context tokens and multiple spans at once; we further cluster the rules to truncate the long tail of the rule distribution. Experiments on several benchmarks show that HCT can outperform state-of-the-art rewriting systems by ~2 BLEU points. Lisa Jin, Linfeng Song, Lifeng Jin, Dong Yu 0001, Daniel Gildea |
AAAI | 3 |
| 2022 | Salience Allocation as Guidance for Abstractive SummarizationabstractFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 4 |
| 2022 | Learning a Grammar Inducer from Massive Uncurated Instructional VideosabstractVideo-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text.While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence.Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition.Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work.Experiments show that our model trained only on large-scale YouTube data with no textvideo alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues.Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data. Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Haitao Mi, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001 |
EMNLP | 3 |
| 2021 | Instance-adaptive training with noise-robust losses against noisy labelsabstractIn order to alleviate the huge demand for annotated datasets for different tasks, many recent natural language processing datasets have adopted automated pipelines for fast-tracking usable data.However, model training with such datasets poses a challenge because popular optimization objectives are not robust to label noise induced in the annotation generation process.Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they generally use a single dataset-wise hyperparamter to control the strength of noise resistance.This work proposes novel instance-adaptive training frameworks to change dataset-wise hyperparameters of noise resistance in such losses to be instance-specific.Such instance-specific noise resistance hyperparameters are predicted by special instance-level label quality predictors, which are trained along with the main models.Experiments on noisy and corrupted NLP datasets show that proposed instance-adaptive training frameworks help increase the noiserobustness provided by such losses, promoting the use of the frameworks and associated losses in training NLP models with noisy data. Lifeng Jin, Linfeng Song, Kun Xu 0005, Dong Yu 0001 |
EMNLP (1) | 1 |
| 2021 | Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense InventoriesabstractWord Sense Disambiguation (WSD) aims to automatically identify the exact meaning of one word according to its context.Existing supervised models struggle to make correct predictions on rare word senses due to limited training data and can only select the best definition sentence from one predefined word sense inventory (e.g., WordNet).To address the data sparsity problem and generalize the model to be independent of one predefined inventory, we propose a gloss alignment algorithm that can align definition sentences (glosses) with the same meaning from different sense inventories to collect rich lexical knowledge.We then train a model to identify semantic equivalence between a target word in context and one of its glosses using these aligned inventories, which exhibits strong transfer capability to many WSD tasks 1 .Experiments on benchmark datasets show that the proposed method improves predictions on both frequent and rare word senses, outperforming prior work by 1.2% on the All-Words WSD Task and 4.3% on the Low-Shot WSD Task.Evaluation on WiC Task also indicates that our method can better capture word meanings in context. Wenlin Yao, Xiaoman Pan, Lifeng Jin, Jianshu Chen, Dian Yu 0001, Dong Yu 0001 |
EMNLP (1) | 3 |
| 2021 | Video-aided Unsupervised Grammar InductionabstractSongyang Zhang, Linfeng Song, Lifeng Jin, Kun Xu, Dong Yu, Jiebo Luo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001 |
NAACL-HLT | 3 |
| 2021 | Distant Finetuning with Discourse Relations for Stance Classification
Lifeng Jin, Kun Xu 0005, Linfeng Song, Dong Yu 0001 |
NLPCC (2) | 1 |
| 2021 | Sunway supercomputer architecture towards exascale computing: analysis and practice
Jiangang Gao, Fang Zheng 0015, Fengbin Qi, Yajun Ding, Hongsheng Lu, Wangquan He, Hongmei Wei, Lifeng Jin, Daoyong Gong, Honghui Sun, Hongtao You |
Sci. China Inf. Sci. | 9 |
| 2021 | Depth-Bounded Statistical PCFG Induction as a Model of Human Grammar AcquisitionabstractAbstract This article describes a simple PCFG induction model with a fixed category domain that predicts a large majority of attested constituent boundaries, and predicts labels consistent with nearly half of attested constituent labels on a standard evaluation data set of child-directed speech. The article then explores the idea that the difference between simple grammars exhibited by child learners and fully recursive grammars exhibited by adult learners may be an effect of increasing working memory capacity, where the shallow grammars are constrained images of the recursive grammars. An implementation of these memory bounds as limits on center embedding in a depth-specific transform of a recursive grammar yields a significant improvement over an equivalent but unbounded baseline, suggesting that this arrangement may indeed confer a learning advantage. Lifeng Jin, Lane Schwartz, Finale Doshi-Velez, Timothy A. Miller, William Schuler |
Comput. Linguistics | 1 |
| 2020 | Relation Extraction Exploiting Full Dependency ForestsabstractDependency syntax has long been recognized as a crucial source of features for relation extraction. Previous work considers 1-best trees produced by a parser during preprocessing. However, error propagation from the out-of-domain parser may impact the relation extraction performance. We propose to leverage full dependency forests for this task, where a full dependency forest encodes all possible trees. Such representations of full dependency forests provide a differentiable connection between a parser and a relation extraction model, and thus we are also able to study adjusting the parser parameters based on end-task loss. Experiments on three datasets show that full dependency forests and parser adjustment give significant improvements over carefully designed baselines, showing state-of-the-art or competitive performances on biomedical or newswire benchmarks. Lifeng Jin, Linfeng Song, Yue Zhang 0004, Kun Xu 0005, Wei-Yun Ma, Dong Yu 0001 |
AAAI | 1 |
| 2019 | Unsupervised Learning of PCFGs with Normalizing FlowabstractUnsupervised PCFG inducers hypothesize sets of compact context-free rules as explanations for sentences.These models not only provide tools for low-resource languages, but also play an important role in modeling language acquisition (Bannard et al., 2009;Abend et al., 2017).However, current PCFG induction models, using word tokens as input, are unable to incorporate semantics and morphology into induction, and may encounter issues of sparse vocabulary when facing morphologically rich languages.This paper describes a neural PCFG inducer which employs context embeddings (Peters et al., 2018) in a normalizing flow model (Dinh et al., 2015) to extend PCFG induction to use semantic and morphological information 1 .Linguistically motivated similarity penalty and categorical distance constraints are imposed on the inducer as regularization.Experiments show that the PCFG induction model with normalizing flow produces grammars with state-of-the-art accuracy on a variety of different languages.Ablation further shows a positive effect of normalizing flow, context embeddings and proposed regularizers. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, Lane Schwartz, William Schuler |
ACL (1) | 1 |
| 2019 | Variance of Average Surprisal: A Better Predictor for Quality of Grammar from Unsupervised PCFG InductionabstractIn unsupervised grammar induction, data likelihood is known to be only weakly correlated with parsing accuracy, especially at convergence after multiple runs.In order to find a better indicator for quality of induced grammars, this paper correlates several linguistically-and psycholinguisticallymotivated predictors to parsing accuracy on a large multilingual grammar induction evaluation data set.Results show that variance of average surprisal (VAS) better correlates with parsing accuracy than data likelihood, and that using VAS instead of data likelihood for model selection provides a significant accuracy boost.Further evidence shows VAS to be a better candidate than data likelihood for predicting word order typology classification.Analyses show that VAS seems to separate content words from function words in natural language grammars, and to better arrange words with different frequencies into separate classes that are more consistent with linguistic theory. Lifeng Jin, William Schuler |
ACL (1) | 1 |
| 2019 | Design and Analysis of Radiometric Calibration Mission in-orbit for Environment and Disasters Monitoring SatelliteabstractWith the rapid improvement of optical payload radiometric calibration accuracy, it is necessary to carry out in-orbit radiometric calibration to improve the calibration accuracy and image quality. The design and analysis of in-orbit radiometric calibration with moon, sun and side-slither is presented. The side-slither radiometric calibration mission based on along track scanning is proposed to obtain 60 angles polarization Stokes parameters with the calibration target of deep convective cloud. The satellite attitude stability for in-orbit radiometric calibration is analyzed and meets the requirements of optical image quality. Dexin Sun, Xuebin Liu, Lifeng Jin, Zhaoguang Bai, Huan Yin, Qipeng Cao |
IGARSS | 4 |
| 2018 | Depth-bounding is effective: Improvements and Evaluation of Unsupervised PCFG InductionabstractThere have been several recent attempts to improve the accuracy of grammar induction systems by bounding the recursive complexity of the induction model (Ponvert et al., 2011;Noji and Johnson, 2016;Shain et al., 2016;Jin et al., 2018).Modern depth-bounded grammar inducers have been shown to be more accurate than early unbounded PCFG inducers, but this technique has never been compared against unbounded induction within the same system, in part because most previous depthbounding models are built around sequence models, the complexity of which grows exponentially with the maximum allowed depth.The present work instead applies depth bounds within a chart-based Bayesian PCFG inducer (Johnson et al., 2007b), where bounding can be switched on and off, and then samples trees with and without bounding.1 Results show that depth-bounding is indeed significantly effective in limiting the search space of the inducer and thereby increasing the accuracy of the resulting parsing model.Moreover, parsing results on English, Chinese and German show that this bounded model with a new inference technique is able to produce parse trees more accurately than or competitively with state-ofthe-art constituency-based grammar induction models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
EMNLP | 1 |
| 2018 | Unsupervised Grammar Induction with Depth-bounded PCFGabstractThere has been recent interest in applying cognitively- or empirically-motivated bounds on recursion depth to limit the search space of grammar induction models (Ponvert et al., 2011; Noji and Johnson, 2016; Shain et al., 2016). This work extends this depth-bounding approach to probabilistic context-free grammar induction (DB-PCFG), which has a smaller parameter space than hierarchical sequence models, and therefore more fully exploits the space reductions of depth-bounding. Results for this model on grammar acquisition from transcribed child-directed speech and newswire text exceed or are competitive with those of other models when evaluated on parse accuracy. Moreover, grammars acquired from this model demonstrate a consistent use of category labels, something which has not been demonstrated by other acquisition models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
Trans. Assoc. Comput. Linguistics | 1 |
| 2016 | Memory-Bounded Left-Corner Unsupervised Grammar Induction on Child-Directed InputabstractThis paper presents a new memory-bounded left-corner parsing model for unsupervised raw-text syntax induction, using unsupervised hierarchical hidden Markov models (UHHMM). We deploy this algorithm to shed light on the extent to which human language learners can discover hierarchical syntax through distributional statistics alone, by modeling two widely-accepted features of human language acquisition and sentence processing that have not been simultaneously modeled by any existing grammar induction algorithm: (1) a left-corner parsing strategy and (2) limited working memory capacity. To model realistic input to human language learners, we evaluate our system on a corpus of child-directed speech rather than typical newswire corpora. Results beat or closely match those of three competing systems. Cory Shain, William Bryce, Lifeng Jin, Victoria Krakovna, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
COLING | 3 |
| 2015 | The Overall Markedness of Discourse RelationsabstractDiscourse relations can be categorized as continuous or discontinuous in the hypothesis of continuity (Murray, 1997), with continuous relations expressing normal succession of events in discourse such as temporal, spatial or causal.Asr and Demberg (2013) propose a markedness measure to test the prediction that discontinuous relations may have more unambiguous connectives, but restrict the markedness calculation to relations with explicit connectives only.This paper extends their measure to explicit and implicit relations and shows that results from this extension better fit the continuity hypothesis predictions both for the English Penn Discourse (Prasad et al., 2008) and the Chinese Discourse (Zhou and Xue, 2015) Treebanks. Lifeng Jin, Marie-Catherine de Marneffe |
EMNLP | 1 |
| 2015 | A Comparison of Word Similarity Performance Using Explanatory and Non-explanatory TextsabstractVectorial representations of words derived from large current events datasets have been shown to perform well on word similarity tasks.This paper shows vectorial representations derived from substantially smaller explanatory text datasets such as English Wikipedia and Simple English Wikipedia preserve enough lexical semantic information to make these kinds of category judgments with equal or better accuracy. Lifeng Jin, William Schuler |
HLT-NAACL | 1 |