EDBT 2026 Demo / reviewers in the wild / expert
Vincent Ng 0001
dblp:67/3142
· DBLP profile ↗
156ranked-venue papers
17as first author
50since 2021 · last 2026
0000-0001-8237-429XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 129 · 17 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 23 · 13 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHASE: Contextual History for Adaptive and Simple Exploitation in Large Language Model JailbreakingabstractWe propose Contextual History for Adaptive and Simple Exploitation (CHASE), a novel multi-turn method for Large Language Model (LLM) jailbreaking. Rather than directly attack an LLM that may be difficult to jailbreak, CHASE first collects jailbroken histories from an easy-to-jailbreak LLM and then transfers them to the target LLM. Through this history transfer process, CHASE misleads the target LLM into thinking that it is responsible for producing the jailbroken histories and increases the chances of successful jailbreaking by prompting it to continue the conversation. Extensive evaluations on mainstream LLMs show that CHASE consistently achieves higher attack success rates and demands fewer computational resources compared to existing methods. Zhiqiang Hao, Chuanyi Li, Xiao Fu 0005, Shangqi Wang, Jiao Yin 0007, Jidong Ge, Bin Luo 0003, Vincent Ng 0001 |
AAAI | 11 |
| 2026 | System L: Toward System 2-Style Legal ReasoningabstractDual-system theory distinguishes between fast, intuitive System 1 and slow, deliberative System 2. While this dichotomy describes many forms of reasoning, it oversimplifies the reality of expert legal reasoning. Legal reasoning is not merely a process of slow, logical deliberation. It is intrinsically normative, embedding precedent analysis, statutory interpretation, policy balancing, and social values. This paper envisions a reasoning architecture for legal reasoning, System L (Legal System 2), which extends traditional System 2 by integrating domain-specific normative frameworks in a structured manner. Using the IRAC (Issue–Rule–Application–Conclusion) structure as a backbone model, System L represents a blueprint for the next generation of cognitive and AI systems capable of human-like legal reasoning. Chuanyi Li, Yi Feng 0005, Vincent Ng 0001 |
AAAI | 3 |
| 2026 | Legal Judgment Prediction: A Reflection on the State of the ArtabstractAutomatic legal judgment prediction (LJP) has recently received increasing attention in the natural language processing community because of its practical values in the real world.Significant progress has been achieved on LJP in the past decade.However, most existing LJP research primarily focuses on developing methods that achieve better performance on standard evaluation datasets, with limited emphasis on the long-term advancement of the field beyond improving evaluation metrics.In this position paper, we reflect on the state of the art in LJP research, and explore issues that should motivate researchers to think beyond merely enhancing performance metrics, with the ultimate goal of sparking discussions among LJP researchers about the future trajectory of the field. Yi Feng 0005, Chuanyi Li, Vincent Ng 0001 |
ACL (1) | 3 |
| 2026 | Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the ArtabstractDespite the recent progress made in crossprompt essay scoring, there is little analysis of what makes a state-of-the-art cross-prompt scorer work well.To this end, we present an empirical analysis of how the key components of a cross-prompt scorer interact with each other and impact its overall performance.In addition, we examine for the first time the application of transductive learning to cross-prompt scoring, which represents an important starting point for providing a practical way to improve cross-prompt scorers for use in the rarelystudied classroom setting without the need for additional labeled training data. 1 Shengjie Li 0002, Vincent Ng 0001 |
ACL (1) | 2 |
| 2026 | Improving Legal Judgment Prediction via Quantitative ReasoningabstractLegal Judgment Prediction (LJP) focuses on predicting judgment results based on the facts of cases. While State-of-the-Art (SOTA) methods have shown impressive performance in law article prediction and charge prediction, they still exhibit weaknesses in prison term prediction. One major reason is that existing models fail to mimic human legal quantitative reasoning to understand monetary features in case facts. Consequently, they do not rigorously quantify the severity of the crime, which is essential for prison term prediction. In this article, we explore and explain how to leverage monetary features to improve LJP via quantitative reasoning. Specifically, we propose QR-LJP, a quantitative reasoning-based LJP model, to integrate legal reasoning knowledge into the prediction process. QR-LJP first employs a curated LLM to extract monetary values from case facts and uses legal quantitative reasoning logic to determine the total crime amount, serving as the quantitative measure of the crime’s severity. This measure is subsequently used to make judgment predictions. We evaluate our model on the real-world dataset CAIL-2018. Experimental results demonstrate that our model outperforms current SOTAs, highlighting the effectiveness of legal quantitative reasoning. Moreover, applying our quantitative reasoning strategy to existing SOTA methods yields significant improvements, especially in macro-F1 scores. Zhu Han 0001, Yi Feng 0005, Chuanyi Li, Zhiwei Fei, Xuxing Ding, Jidong Ge, Vincent Ng 0001 |
ACM Trans. Knowl. Discov. Data | 8 |
| 2026 | YourCoLo: Leveraging One-to-Many Relationships and Inter-Code Connections for User Review-Based Code LocalizationabstractIn an era where mobile devices are ubiquitous, digital distribution platforms such as the Google Play Store have become integral to our daily lives, hosting millions of applications and serving billions of users. Users can leave reviews to provide developers with valuable feedback, including requests for new features and reports of issues. These user reviews play a crucial role in software development, testing, and maintenance by informing developers about user needs and potential problems, which motivates us to revisit a key problem: given user reviews, how can we automatically identify the relevant code snippets from software codebases to assist developers in addressing the reviews? Existing practices to address this problem typically involve calculating the similarity between user reviews and code snippets. However, we identify three key limitations. First, although existing methods show promising results on individual projects, their high performance cannot be generalized across projects. Second, the state-of-the-art approach models the problem as a one-to-one relationship between a user review and code snippets, ignoring the one-to-many relationship that often exists. Third, the state-of-the-art approach focuses solely on the direct relationship between reviews and code snippets, overlooking the interconnections among code snippets themselves, which contain valuable information that can aid in accurately identifying relevant code. To address these limitations and advance the state of the art, we propose YourCoLo , a novel approach that fully leverages contextual information, one-to-many relationships, and inter-code connections. Specifically, YourCoLo is powered by three novel designs: (1) a prompt-enhanced mechanism to incorporate rich project-level context into code localization, (2) a new loss function designed to handle the one-to-many relationships between user reviews and multiple relevant code snippets, and (3) a ranking strategy that considers interconnections among related code snippets. Our experimental evaluation shows that YourCoLo substantially outperforms state-of-the-art models, surpassing CodeBERT, CodeLlama, and GraphCodeBERT by 18.3, 9.3, and 7.7 percentage points at the method level and by 18.4, 7.7, and 7.0 percentage points at the file level (in terms of mean reciprocal rank). In addition, YourCoLo also achieves improvements of 8.8 percentage points and 6.8 percentage points in mean average precision (MAP) at the method and file levels, respectively, compared to the state-of-the-art method. These results underscore YourCoLo ’s effectiveness and its potential to guide developers more accurately toward the code snippets most pertinent to user feedback. Changan Niu, Zhou Yang 0003, Chuanyi Li, Yi Feng 0005, Jidong Ge, Bin Luo 0003, David Lo 0001, Vincent Ng 0001 |
ACM Trans. Softw. Eng. Methodol. | 9 |
| 2026 | IMPACT: Identifying and Classifying Multiple Sourced and Categorized Self-Admitted Technical DebtsabstractSelf-Admitted Technical Debt (SATD) refers to sub-optimal solutions deliberately introduced to accelerate the software development process, often at the expense of software maintainability and sustainability. Therefore, timely identification and repayment of the SATD is critical for the software system. As exploration deepens, it is found that effectively prioritizing the repayment of SATD with more significant impacts on software quality requires not only identifying SATD but also further classifying it. However, existing SATD identification and classification approaches face the following challenges: (1) SATDs originate from diverse sources. Code comments are a widespread source, but recent research has revealed that SATDs can originate from other sources, such as pull requests, issues, and commit messages. Nonetheless, existing approaches primarily target code comments, lacking the capability to analyze SATDs from other sources effectively. (2) SATDs fall into diverse categories. Nonetheless, existing SATD classification approaches fail to address all SATD categories comprehensively and show inadequate performance. (3) Imbalance of existing SATD datasets. Real-world SATD data are scarce, making dataset collection challenging. Moreover, SATD distribution across different sources is uneven, further complicating the construction of high-quality datasets. To alleviate these challenges, this article presents an SATD identification and classification framework named IMPACT . First, IMPACT employs ChatGPT to construct an augmented dataset. Subsequently, it utilizes a pipeline with two fine-tuned language models of different parameter sizes to identify and classify SATD separately. To evaluate the effectiveness of IMPACT, we compare it with three state-of-the-art SATD classification methods and its two foundation models. Experimental results demonstrate that IMPACT outperforms state-of-the-art methods by a large margin, and even surpasses its foundation model GLM-4-9B-Chat. It achieves the optimal average F1 score of 0.697 on the source of pull requests, the most challenging data source. Moreover, experiments on the cross-project test set show that IMPACT demonstrates strong generalizability on unseen project data. Zhixin Yin, Yaopeng Yang, Chuanyi Li, Zongwen Shen, Jidong Ge, Wenkang Zhong, Bin Luo 0003, Vincent Ng 0001 |
ACM Trans. Softw. Eng. Methodol. | 9 |
| 2026 | P-NPR: Practical Neural Program Repair via Learning to Ensemble
Zhongqiang Pan, Chuanyi Li, Wenkang Zhong, Bin Luo 0003, Vincent Ng 0001 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Understanding AdvertisementsabstractWhile AI systems are capable of reading texts and seeing images, they typically perceive surface information explicitly conveyed with limited abilities to comprehend hidden messages (e.g., a double-edged remark). We propose the novel task of advertisement understanding: given an advertisement, which can be a text, an image, or a video, the goal is to identify the persuasion strategies used and determine the (possibly hidden) messages conveyed. Efforts on this task could enhance machine comprehension capabilities, and provide users with increased situation awareness w.r.t. the advertised message and thus possibly enable mindful decision making. We believe that this task presents long-term challenges to AI researchers and that successful understanding of ads could bring machine understanding one important step closer to human understanding. Yi Feng 0005, Chuanyi Li, Vincent Ng 0001 |
AAAI | 3 |
| 2025 | MemeQA: Holistic Evaluation for Meme UnderstandingabstractKhoi P. N. Nguyen, Terrence Li, Derek Lou Zhou, Gabriel Xiong, Pranav Balu, Nandhan Alahari, Alan Huang, Tanush Chauhan, Harshavardhan Bala, Emre Guzelordu, Affan Kashfi, Aaron Xu, Suyesh Shrestha, Megan Vu, Jerry Wang, Vincent Ng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Khoi P. N. Nguyen, Terrence Li, Derek Lou Zhou, Gabriel Xiong, Pranav Balu, Nandhan Alahari, Alan Huang, Tanush Chauhan, Harshavardhan Bala, Emre Guzelordu, Affan Kashfi, Aaron Xu, Suyesh Shrestha, Megan Kim Vu, Jerry Yining Wang, Vincent Ng 0001 |
ACL (1) | 16 |
| 2025 | InternLM-Law: An Open-Sourced Chinese Legal Large Language ModelabstractWe introduce InternLM-Law, a large language model (LLM) tailored for addressing diverse legal tasks related to Chinese laws. These tasks range from responding to standard legal questions (e.g., legal exercises in textbooks) to analyzing complex real-world legal situations. Our work contributes to Chinese Legal NLP research by (1) conducting one of the most extensive evaluations of state-of-the-art general-purpose and legal-specific LLMs to date that involves an automatic evaluation on the 20 legal NLP tasks in LawBench, a human evaluation on a challenging version of the Legal Consultation task, and an automatic evaluation of a model’s ability to handle very long legal texts; (2) presenting a methodology for training a Chinese legal LLM that offers superior performance to all of its counterparts in our extensive evaluation; and (3) facilitating future research in this area by making all of our code and model publicly available at https://github.com/InternLM/InternLM-Law. Zhiwei Fei, Songyang Zhang 0001, Xiaoyu Shen 0001, Xiao Wang 0042, Jidong Ge, Vincent Ng 0001 |
COLING | 7 |
| 2025 | Multimodal Neural Machine Translation: A Survey of the State of the ArtabstractMultimodal neural machine translation (MNMT) has received increasing attention due to its widespread applications in various fields such as cross-border e-commerce and cross-border social media platforms.The task aims to integrate other modalities, such as the visual modality, with textual data to enhance translation performance.We survey the major milestones in MNMT research, providing a comprehensive overview of relevant datasets and recent methodologies, and discussing key challenges and promising research directions. Yi Feng 0005, Chuanyi Li, Jiatong He, Vincent Ng 0001 |
EMNLP | 5 |
| 2025 | Graph-Based Multi-Trait Essay ScoringabstractWhile virtually all existing work on Automated Essay Scoring (AES) models an essay as a word sequence, we put forward the novel view that an essay can be modeled as a graph and subsequently propose GAT-AES 1 , a graph-attention network approach to AES.GAT-AES models the interactions among essay traits in a principled manner by (1) representing each essay trait as a trait node in the graph and connecting each pair of trait nodes with directed edges, and (2) allowing neighboring nodes to influence each other by using a convolutional operator to update node representations.Unlike competing approaches, which can only model one-hop dependencies, GAT-AES allows us to easily model multi-hop dependencies.Experimental results demonstrate that GAT-AES achieves the best multi-trait scoring results to date on the ASAP++ dataset.Further analysis shows that GAT-AES outperforms not only alternative graph neural networks but also approaches that use trait-attention mechanisms to model trait dependencies. Shengjie Li 0002, Vincent Ng 0001 |
EMNLP | 2 |
| 2025 | LawShift: Benchmarking Legal Judgment Prediction Under Statute ShiftsabstractLegal Judgment Prediction (LJP) seeks to predict case outcomes given available case information, offering practical value for both legal professionals and laypersons. However, a key limitation of existing LJP models is their limited adaptability to statutory revisions. Current SOTA models are neither designed nor evaluated for statutory revisions. To bridge this gap, we introduce LawShift, a benchmark dataset for evaluating LJP under statutory revisions. Covering 31 fine-grained change types, LawShift enables systematic assessment of SOTA models' ability to handle legal changes. We evaluate five representative SOTA models on LawShift, uncovering significant limitations in their response to legal updates. Our findings show that model architecture plays a critical role in adaptability, offering actionable insights and guiding future research on LJP in dynamic legal contexts. Zhuo Han, Yi Feng 0005, Wanhong Huang 0003, Xuxing Ding, Chuanyi Li, Jidong Ge, Vincent Ng 0001 |
NeurIPS | 8 |
| 2025 | Benchmarking and Categorizing the Performance of Neural Program Repair Systems for JavaabstractRecent years have seen a rise in Neural Program Repair (NPR) systems in the software engineering community, which adopt advanced deep learning techniques to automatically fix bugs. Having a comprehensive understanding of existing systems can facilitate new improvements in this area and provide practical instructions for users. However, we observe two potential weaknesses in the current evaluation of NPR systems: ① published systems are trained with varying data, and ② NPR systems are roughly evaluated through the number of totally fixed bugs. Questions such as what types of bugs are repairable for current systems cannot be answered yet. Consequently, researchers cannot make target improvements in this area and users have no idea of the real affair of existing systems. In this article, we perform a systematic evaluation of the existing nine state-of-the-art NPR systems. To perform a fair and detailed comparison, we (1) build a new benchmark and framework that supports training and validating the nine systems with unified data and (2) evaluate re-trained systems with detailed performance analysis, especially on the effectiveness and the efficiency. We believe our benchmark tool and evaluation results could offer practitioners the real affairs of current NPR systems and the implications of further facilitating the improvements of NPR. Wenkang Zhong, Chuanyi Li, Kui Liu 0001, Jidong Ge, Bin Luo 0003, Tegawendé F. Bissyandé, Vincent Ng 0001 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2024 | Conundrums in Cross-Prompt Automated Essay Scoring: Making Sense of the State of the ArtabstractCross-prompt automated essay scoring (AES), an under-investigated but challenging task that has gained increasing popularity in the AES community, aims to train an AES system that can generalize well to prompts that are unseen during model training.While recentlydeveloped cross-prompt AES models have combined essay representations that are learned via sophisticated neural architectures with socalled prompt-independent features, an intriguing question is: are complex neural models needed to achieve state-of-the-art results?We answer this question by abandoning sophisticated neural architectures and developing a purely feature-based approach to cross-prompt AES that adopts a simple neural architecture.Experiments on the ASAP dataset demonstrate that our simple approach to cross-prompt AES can achieve state-of-the-art results. Shengjie Li 0002, Vincent Ng 0001 |
ACL (1) | 2 |
| 2024 | Legal Case Retrieval: A Survey of the State of the ArtabstractRecent years have seen increasing attention on Legal Case Retrieval (LCR), a key task in the area of Legal AI that concerns the retrieval of cases from a large legal database of historical cases that are similar to a given query.This paper presents a survey of the major milestones made in LCR research, targeting researchers who are finding their way into the field and seek a brief account of the relevant datasets and the recent neural models and their performances. Yi Feng 0005, Chuanyi Li, Vincent Ng 0001 |
ACL (1) | 3 |
| 2024 | Universal Anaphora: The First Three YearsabstractThe aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. Although several papers on aspects of the initiative have appeared, no overall description of the initiative’s goals, proposals and achievements has been published yet except as an online draft. This paper aims to fill this gap, as well as to discuss its progress so far. Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng 0001, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Daniel Zeman |
LREC/COLING | 3 |
| 2024 | Automated Essay Scoring: A Reflection on the State of the ArtabstractWhile steady progress has been made on the task of automated essay scoring (AES) in the past decade, much of the recent work in this area has focused on developing models that beat existing models on a standard evaluation dataset.While improving performance numbers remains an important goal in the short term, such a focus is not necessarily beneficial for the long-term development of the field.We reflect on the state of the art in AES research, discussing issues that we believe can encourage researchers to think bigger than improving performance numbers, with the ultimate goal of triggering discussion among AES researchers on how we should move forward. Shengjie Li 0002, Vincent Ng 0001 |
EMNLP | 2 |
| 2024 | LawBench: Benchmarking Legal Knowledge of Large Language ModelsabstractZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiwei Fei, Xiaoyu Shen 0001, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang 0001, Kai Chen 0026, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng 0001 |
EMNLP | 12 |
| 2024 | Computational Meme Understanding: A SurveyabstractComputational Meme Understanding, which concerns the automated comprehension of memes, has garnered interest over the last four years and is facing both substantial opportunities and challenges.We survey this emerging area of research by first introducing a comprehensive taxonomy for memes along three dimensions -forms, functions, and topics.Next, we present three key tasks in Computational Meme Understanding, namely, classification, interpretation, and explanation, and conduct a comprehensive review of existing datasets and models, discussing their limitations.Finally, we highlight the key challenges and recommend avenues for future work. 1 Khoi P. N. Nguyen, Vincent Ng 0001 |
EMNLP | 2 |
| 2024 | FAIR: Flow Type-Aware Pre-Training of Compiler Intermediate RepresentationsabstractWhile the majority of existing pre-trained models from code learn source code features such as code tokens and abstract syntax trees, there are some other works that focus on learning from compiler intermediate representations (IRs). Existing IR-based models typically utilize IR features such as instructions, control and data flow graphs (CDFGs), call graphs, etc. However, these methods confuse variable nodes and instruction nodes in a CDFG and fail to distinguish different types of flows, and the neural networks they use fail to capture long-distance dependencies and have over-smoothing and over-squashing problems. To address these weaknesses, we propose FAIR, a Flow type-Aware pre-trained model for IR that involves employing (1) a novel input representation of IR programs; (2) Graph Transformer to address over-smoothing, over-squashing and long-dependencies problems; and (3) five pre-training tasks that we specifically propose to enable FAIR to learn the semantics of IR tokens, flow type information, and the overall representation of IR. Experimental results show that FAIR can achieve state-of-the-art results on four code-related downstream tasks. Changan Niu, Chuanyi Li, Vincent Ng 0001, David Lo 0001, Bin Luo 0003 |
ICSE | 3 |
| 2024 | Practical Program Repair via Preference-based Ensemble StrategyabstractTo date, over 40 Automated Program Repair (APR) tools have been designed with varying bug-fixing strategies, which have been demonstrated to have complementary performance in terms of being effective for different bug classes. Intuitively, it should be feasible to improve the overall bug-fixing performance of APR via assembling existing tools. Unfortunately, simply invoking all available APR tools for a given bug can result in unacceptable costs on APR execution as well as on patch validation (via expensive testing). Therefore, while assembling existing tools is appealing, it requires an efficient strategy to reconcile the need to fix more bugs and the requirements for practicality. In light of this problem, we propose a Preference-based Ensemble Program Repair framework (P-EPR), which seeks to effectively rank APR tools for repairing different bugs. P-EPR is the first non-learning-based APR ensemble method that is novel in its exploitation of repair patterns as a major source of knowledge for ranking APR tools and its reliance on a dynamic update strategy that enables it to immediately exploit and benefit from newly derived repair results. Experimental results show that P-EPR outperforms existing strategies significantly both in flexibility and effectiveness. Wenkang Zhong, Chuanyi Li, Kui Liu 0001, Tongtong Xu, Jidong Ge, Tegawendé F. Bissyandé, Bin Luo 0003, Vincent Ng 0001 |
ICSE | 8 |
| 2024 | Automated Essay Scoring: Recent Successes and Future Directions
Shengjie Li 0002, Vincent Ng 0001 |
IJCAI | 2 |
| 2024 | ICLE++: Modeling Fine-Grained Traits for Holistic Essay ScoringabstractThe majority of the recently-developed models for automated essay scoring (AES) are evaluated solely on the ASAP corpus.However, ASAP is not without its limitations.For instance, it is not clear whether models trained on ASAP can generalize well when evaluated on other corpora.In light of these limitations, we introduce ICLE++, a corpus of persuasive student essays annotated with both holistic scores and trait-specific scores.Not only can ICLE++ be used to test the generalizability of AES models trained on ASAP, but it can also facilitate the evaluation of models developed for newer AES problems such as multi-trait scoring and cross-prompt scoring.We believe that ICLE++, which represents a culmination of our longterm effort in annotating the essays in the ICLE corpus, contributes to the set of much-needed annotated corpora for AES research. Shengjie Li 0002, Vincent Ng 0001 |
NAACL-HLT | 2 |
| 2024 | MemeIntent: Benchmarking Intent Description Generation for MemesabstractJeongsik Park, Khoi P. N. Nguyen, Terrence Li, Suyesh Shrestha, Megan Kim Vu, Jerry Yining Wang, Vincent Ng. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Jeongsik Park, Khoi P. N. Nguyen, Terrence Li, Suyesh Shrestha, Megan Kim Vu, Jerry Yining Wang, Vincent Ng 0001 |
SIGDIAL | 7 |
| 2024 | PassSum: Leveraging paths of abstract syntax trees and self-supervision for code summarizationabstractAbstract Code summarization is to provide a high‐level comment for a code snippet that typically describes the function and intent of the given code. Recent years have seen the successful application of data‐driven code summarization. To improve the performance of the model, numerous approaches use abstract syntax trees (ASTs) to represent the structural information of the code, which is considered by most researchers to be the main factor that distinguishes code from natural language. Then, such data‐driven methods are trained on large‐scale labeled datasets to obtain a model with strong generalization capabilities that can be applied to new examples. Nevertheless, we argue that state‐of‐the‐art approaches suffer from two key weaknesses: (1) inefficient encoding of ASTs; (2) reliance on a large labeled corpus for model training. As a result, such drawbacks lead to (1) oversized model, slow training, information loss and instability; (2) inability to be applied to programming languages with only a small amount of labeled data. In light of these weaknesses, we propose PassSum, a code summarization approach that addresses the aforementioned weaknesses via (1) a novel input representation which contains an efficient AST encoding method; (2) introducing three pretraining objectives and pretraining our model with a large amount of (easy‐to‐obtain) unlabeled data under the guidance of self‐supervised learning. Experimental results on code summarization for Java, Python, and Ruby methods demonstrate the superiority of PassSum to state‐of‐the‐art methods. Further experiments demonstrate that the input representation we use has both temporal and spatial advantages in addition to performance leadership. In addition, pretraining is also shown to make the model more generalizable with less labeled data, and also to speed up the convergence of the model during training. Changan Niu, Chuanyi Li, Vincent Ng 0001, Jidong Ge, LiGuo Huang, Bin Luo 0003 |
J. Softw. Evol. Process. | 3 |
| 2024 | Comparing the Pretrained Models of Source Code by Re-pretraining Under a Unified SetupabstractRecent years have seen the successful application of large pretrained models of source code (CodePTMs) to code representation learning, which have taken the field of software engineering (SE) from task-specific solutions to task-agnostic generic models. By the remarkable results, CodePTMs are seen as a promising direction in both academia and industry. While a number of CodePTMs have been proposed, they are often not directly comparable because they differ in experimental setups such as pretraining dataset, model size, evaluation tasks, and datasets. In this article, we first review the experimental setup used in previous work and propose a standardized setup to facilitate fair comparisons among CodePTMs to explore the impacts of their pretraining tasks. Then, under the standardized setup, we re-pretrain CodePTMs using the same model architecture, input modalities, and pretraining tasks, as they declared and fine-tune each model on each evaluation SE task for evaluating. Finally, we present the experimental results and make a comprehensive discussion on the relative strength and weakness of different pretraining tasks with respect to each SE task. We hope our view can inspire and advance the future study of more powerful CodePTMs. Changan Niu, Chuanyi Li, Vincent Ng 0001, Bin Luo 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | PTM-APIRec: Leveraging Pre-trained Models of Source Code in API RecommendationabstractRecommending APIs is a practical and essential feature of IDEs. Improving the accuracy of API recommendations is an effective way to improve coding efficiency. With the success of deep learning in software engineering, the state-of-the-art (SOTA) performance of API recommendation is also achieved by deep-learning-based approaches. However, existing SOTAs either only consider the API sequences in the code snippets or rely on complex operations for extracting hand-crafted features, all of which have potential risks in under-encoding the input code snippets and further resulting in sub-optimal recommendation performance. To this end, this article proposes to utilize the code understanding ability of existing general code P re- T raining M odels to fully encode the input code snippet to improve the accuracy of API Rec ommendation, namely, PTM-APIRec . To ensure that the code semantics of the input are fully understood and the API recommended actually exists, we use separate vocabularies for the input code snippet and the APIs to be predicted. The experimental results on the JDK and Android datasets show that PTM-APIRec surpasses existing approaches. Besides, an effective way to improve the performance of PTM-APIRec is to enhance the pre-trained model with more pre-training data (which is easier to obtain than API recommendation datasets). Chuanyi Li, Ze Tang 0002, Wanhong Huang 0003, Jidong Ge, Bin Luo 0003, Vincent Ng 0001 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2023 | Multimodal Propaganda ProcessingabstractPropaganda campaigns have long been used to influence public opinion via disseminating biased and/or misleading information. Despite the increasing prevalence of propaganda content on the Internet, few attempts have been made by AI researchers to analyze such content. We introduce the task of multimodal propaganda processing, where the goal is to automatically analyze propaganda content. We believe that this task presents a long-term challenge to AI researchers and that successful processing of propaganda could bring machine understanding one important step closer to human understanding. We discuss the technical challenges associated with this task and outline the steps that need to be taken to address it. Vincent Ng 0001, Shengjie Li 0002 |
AAAI | 1 |
| 2023 | PairSpanBERT: An Enhanced Language Model for Bridging ResolutionabstractWe present PAIRSPANBERT, a SPANBERTbased pre-trained model specialized for bridging resolution.PAIRSPANBERT is pre-trained with a novel objective that aims to learn the contexts in which two mentions are implicitly linked to each other from a large amount of data automatically generated either heuristically or via distance supervision with a knowledge graph.Despite the noise inherent in the automatically generated data, we achieve the best results reported to date on three evaluation datasets for bridging resolution when replacing SPANBERT with PAIRSPANBERT in a stateof-the-art resolver that jointly performs entity coreference resolution and bridging resolution. Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001 |
ACL (1) | 3 |
| 2023 | An Empirical Comparison of Pre-Trained Models of Source CodeabstractWhile a large number of pre-trained models of source code have been successfully developed and applied to a variety of software engineering (SE) tasks in recent years, our understanding of these pre-trained models is arguably fairly limited. With the goal of advancing our understanding of these models, we perform the first systematic empirical comparison of 19 recently-developed pre-trained models of source code on 13 SE tasks. To gain additional insights into these models, we adopt a recently -developed 4-dimensional categorization of pre-trained models, and subsequently investigate whether there are correlations between different categories of pre-trained models and their performances on different SE tasks. Changan Niu, Chuanyi Li, Vincent Ng 0001, Dongxiao Chen, Jidong Ge, Bin Luo 0003 |
ICSE | 3 |
| 2023 | CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code ModelsabstractDespite the recent advances showing that a model pre-trained on large-scale source code data is able to gain appreciable generalization capability, it still requires a sizeable amount of data on the target task for fine-tuning. And the effectiveness of the model generalization is largely affected by the size and quality of the fine-tuning data, which is detrimental for target tasks with limited or unavailable resources. Therefore, cross-task generalization, with the goal of improving the generalization of the model to unseen tasks that have not been seen before, is of strong research and application value. In this paper, we propose a large-scale benchmark that includes 216 existing code-related tasks. Then, we annotate each task with the corresponding meta information such as task description and instruction, which contains detailed information about the task and a solution guide. This also helps us to easily create a wide variety of “training/evaluation” task splits to evaluate the various cross-task generalization capabilities of the model. Then we perform some preliminary experiments to demonstrate that the cross-task generalization of models can be largely improved by in-context learning methods such as few-shot learning and learning from task instructions, which shows the promising prospects of conducting cross-task learning research on our benchmark. We hope that the collection of the datasets and our benchmark will facilitate future work that is not limited to cross-task generalization. Changan Niu, Chuanyi Li, Vincent Ng 0001, Bin Luo 0003 |
ICSE | 3 |
| 2023 | Machine/Deep Learning for Software Engineering: A Systematic Literature ReviewabstractSince 2009, the deep learning revolution, which was triggered by the introduction of ImageNet, has stimulated the synergy between Software Engineering (SE) and Machine Learning (ML)/Deep Learning (DL). Meanwhile, critical reviews have emerged that suggest that ML/DL should be used cautiously. To improve the applicability and generalizability of ML/DL-related SE studies, we conducted a 12-year Systematic Literature Review (SLR) on 1,428 ML/DL-related SE papers published between 2009 and 2020. Our trend analysis demonstrated the impacts that ML/DL brought to SE. We examined the complexity of applying ML/DL solutions to SE problems and how such complexity led to issues concerning the reproducibility and replicability of ML/DL studies in SE. Specifically, we investigated how ML and DL differ in data preprocessing, model training, and evaluation when applied to SE tasks, and what details need to be provided to ensure that a study can be reproduced or replicated. By categorizing the rationales behind the selection of ML/DL techniques into five themes, we analyzed how model performance, robustness, interpretability, complexity, and data simplicity affected the choices of ML/DL models. LiGuo Huang, Amiao Gao, Jidong Ge, Haitao Feng, Ishna Satyarth, Ming Li 0005, He Zhang 0001, Vincent Ng 0001 |
IEEE Trans. Software Eng. | 10 |
| 2022 | Commonsense Knowledge Reasoning and Generation with Pre-trained Language Models: A SurveyabstractWhile commonsense knowledge acquisition and reasoning has traditionally been a core research topic in the knowledge representation and reasoning community, recent years have seen a surge of interest in the natural language processing community in developing pre-trained models and testing their ability to address a variety of newly designed commonsense knowledge reasoning and generation tasks. This paper presents a survey of these tasks, discusses the strengths and weaknesses of state-of-the-art pre-trained models for commonsense reasoning and generation as revealed by these tasks, and reflects on future research directions. Prajjwal Bhargava, Vincent Ng 0001 |
AAAI | 2 |
| 2022 | Legal Judgment Prediction via Event Extraction with ConstraintsabstractWhile significant progress has been made on the task of Legal Judgment Prediction (LJP) in recent years, the incorrect predictions made by SOTA LJP models can be attributed in part to their failure to (1) locate the key event information that determines the judgment, and (2) exploit the cross-task consistency constraints that exist among the subtasks of LJP.To address these weaknesses, we propose EPM, an Event-based Prediction Model with constraints, which surpasses existing SOTA models in performance on a standard LJP dataset. Yi Feng 0005, Chuanyi Li, Vincent Ng 0001 |
ACL (1) | 3 |
| 2022 | Constrained Multi-Task Learning for Bridging ResolutionabstractWe examine the extent to which supervised bridging resolvers can be improved without employing additional labeled bridging data by proposing a novel constrained multi-task learning framework for bridging resolution, within which we (1) design cross-task consistency constraints to guide the learning process; (2) pretrain the entity coreference model in the multitask framework on the large amount of publicly available coreference data; and (3) integrate prior knowledge encoded in rule-based resolvers.Our approach achieves state-of-theart results on three standard evaluation corpora. Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001 |
ACL (1) | 3 |
| 2022 | End-to-End Neural Bridging ResolutionabstractThe state of bridging resolution research is rather unsatisfactory: not only are state-of-the-art resolvers evaluated in unrealistic settings, but the neural models underlying these resolvers are weaker than those used for entity coreference resolution. In light of these problems, we evaluate bridging resolvers in an end-to-end setting, strengthen them with better encoders, and attempt to gain a better understanding of them via perturbation experiments and a manual analysis of their outputs. Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001 |
COLING | 3 |
| 2022 | DiscoSense: Commonsense Reasoning with Discourse ConnectivesabstractWe present DISCOSENSE, a benchmark for commonsense reasoning via understanding a wide variety of discourse connectives.We generate compelling distractors in DISCOSENSE using Conditional Adversarial Filtering, an extension of Adversarial Filtering that employs conditional generation.We show that state-ofthe-art pre-trained language models struggle to perform well on DISCOSENSE, which makes this dataset ideal for evaluating next-generation commonsense reasoning systems. Prajjwal Bhargava, Vincent Ng 0001 |
EMNLP | 2 |
| 2022 | End-to-End Neural Discourse Deixis Resolution in DialogueabstractWe adapt Lee et al.'s (2018) span-based entity coreference model to the task of end-to-end discourse deixis resolution in dialogue, specifically by proposing extensions to their model that exploit task-specific characteristics.The resulting model, dd-utt, achieves state-ofthe-art results on the four datasets in the CODI-CRAC 2021 shared task. Shengjie Li 0002, Vincent Ng 0001 |
EMNLP | 2 |
| 2022 | SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsabstractRecent years have seen the successful application of large pre-trained models to code representation learning, resulting in substantial improvements on many code-related downstream tasks. But there are issues surrounding their application to SE tasks. First, the majority of the pre-trained models focus on pre-training only the encoder of the Transformer. For generation tasks that are addressed using models with the encoder-decoder architecture, however, there is no reason why the decoder should be left out during pre-training. Second, many existing pre-trained models, including state-of-the-art models such as T5-learning, simply reuse the pretraining tasks designed for natural languages. Moreover, to learn the natural language description of source code needed eventually for code-related tasks such as code summarization, existing pretraining tasks require a bilingual corpus composed of source code and the associated natural language description, which severely limits the amount of data for pre-training. To this end, we propose SPT-Code, a sequence-to-sequence pre-trained model for source code. In order to pre-train SPT-Code in a sequence-to-sequence manner and address the aforementioned weaknesses associated with existing pre-training tasks, we introduce three pre-training tasks that are specifically designed to enable SPT-Code to learn knowledge of source code, the corresponding code structure, as well as a natural language description of the code without relying on any bilingual corpus, and eventually exploit these three sources of information when it is applied to downstream tasks. Experimental results demonstrate that SPT-Code achieves state-of-the-art performance on five code-related downstream tasks after fine-tuning. Changan Niu, Chuanyi Li, Vincent Ng 0001, Jidong Ge, LiGuo Huang, Bin Luo 0003 |
ICSE | 3 |
| 2022 | Legal Judgment Prediction: A Survey of the State of the ArtabstractAutomatic legal judgment prediction (LJP) has recently received increasing attention in the natural language processing community in part because of its practical values as well as the associated research challenges. We present an overview of the major milestones made in LJP research covering multiple jurisdictions and multiple languages, and conclude with promising future research directions. Yi Feng 0005, Chuanyi Li, Vincent Ng 0001 |
IJCAI | 3 |
| 2022 | DeexaggerationabstractWe introduce a new task in hyperbole processing, deexaggeration, which concerns the recovery of the meaning of what is being exaggerated in a hyperbolic sentence in the form of a structured representation. In this paper, we lay the groundwork for the computational study of understanding hyperbole by (1) defining a structured representation to encode what is being exaggerated in a hyperbole in a non-hyperbolic manner, (2) annotating the hyperbolic sentences in two existing datasets, HYPO and HYPO-cn, using this structured representation, (3) conducting an empirical analysis of our annotated corpora, and (4) presenting preliminary results on the deexaggeration task. Chuanyi Li, Vincent Ng 0001 |
IJCAI | 3 |
| 2022 | Deep Learning Meets Software Engineering: A Survey on Pre-Trained Models of Source CodeabstractRecent years have seen the successful application of deep learning to software engineering (SE). In particular, the development and use of pre-trained models of source code has enabled state-of-the-art results to be achieved on a wide variety of SE tasks. This paper provides an overview of this rapidly advancing field of research and reflects on future research directions. Changan Niu, Chuanyi Li, Bin Luo 0003, Vincent Ng 0001 |
IJCAI | 4 |
| 2022 | Predicting Product Review Helpfulness - A Hybrid MethodabstractRecent years have seen a rapidly growing number of online reviews of products. As a result, it is often not possible for customers to go through each review before making purchase decisions. One way to address this problem is to build a system for automatically addressing the helpfulness of reviews and present only those reviews that are determined to be helpful by the system to an end user. The vast majority of existing approaches to the task of review helpfulness prediction are based on hand-crafted features, thus making system performance heavily dependent on the quality of these features. In light of this weakness, we propose a new model of review helpfulness prediction using a combination of Convolutional Neural Network (CNN) and TransE wherein hand-crafted features can also be incorporated to improve the output. Specifically, CNN enables us to learn the semantic information from a review and TransE is used to capture the relationship between different entities mentioned in the review. Experiments on the Amazon product review datasets demonstrate that our approach significantly outperforms the state of the art. Chuanyi Li, Jidong Ge, Vincent Ng 0001, Bin Luo 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Span-Based Event Coreference ResolutionabstractMotivated by the recent successful application of span-based models to entity-based information extraction tasks, we investigate span-based models for event coreference resolution, focusing on determining (1) whether the successes of span-based models of entity coreference can be extended to event coreference; (2) whether exploiting the dependency between event coreference and the related subtask of trigger detection; and (3) whether automatically computed entity coreference information can benefit span-based event coreference resolution. Empirical results on the standard evaluation dataset provide affirmative answers to all three questions. Vincent Ng 0001 |
AAAI | 2 |
| 2021 | Conundrums in Event Coreference Resolution: Making Sense of the State of the ArtabstractDespite recent promising results achieved by span-based approaches to event coreference resolution, there is a lack of understanding of what has been improved.We present an empirical analysis of our state-of-the-art span-based event coreference resolver (Lu and Ng, 2021) with the goal of providing the general NLP audience with a better understanding of the state of the art and coreference researchers with directions for future research. Vincent Ng 0001 |
EMNLP (1) | 2 |
| 2021 | Bridging Resolution: Making Sense of the State of the ArtabstractWhile Yu and Poesio (2020) have recently demonstrated the superiority of their neural multi-task learning (MTL) model to rulebased approaches for bridging anaphora resolution, there is little understanding of (1) how it is better than the rule-based approaches (e.g., are the two approaches making similar or complementary mistakes?)and ( 2) what should be improved.To shed light on these issues, we (1) propose a hybrid rule-based and MTL approach that would enable a better understanding of their comparative strengths and weaknesses; and (2) perform a manual analysis of the errors made by the MTL model. Hideo Kobayashi, Vincent Ng 0001 |
NAACL-HLT | 2 |
| 2021 | Constrained Multi-Task Learning for Event Coreference ResolutionabstractWe propose a neural event coreference model in which event coreference is jointly trained with five tasks: trigger detection, entity coreference, anaphoricity determination, realis detection, and argument extraction.To guide the learning of this complex model, we incorporate cross-task consistency constraints into the learning process as soft constraints via designing penalty functions.In addition, we propose the novel idea of viewing entity coreference and event coreference as a single coreference task, which we believe is a step towards a unified model of coreference resolution.The resulting model achieves state-of-the-art results on the KBP 2017 event coreference dataset. Vincent Ng 0001 |
NAACL-HLT | 2 |
| 2021 | Recommending Statutes: A Portable Method Based on Neural NetworksabstractLegal judgment prediction, which aims at predicting judgment results such as penalty, charges, and statutes for cases, has attracted much attention recently. In this article, we focus on building a recommender system to predict the associated statutes for a case given the facts of the case as input. For this purpose, we propose a two-step neural network-based machine learning framework to assist judges as well as ordinary people to reduce their effort in finding applicable statutes. The proposed model takes advantage of recurrent neural networks with a max-pooling layer to obtain contextual representations of documents, i.e., the facts associated with the cases. Moreover, an attention mechanism is used to automatically focus on the important words contributing to the prediction of statutes. In addition, we apply an encoder--decoder ranking approach to extract correlations between statutes to achieve more accurate recommendation results. We evaluate our model on a real-world dataset. Experimental results show that, compared with existing baseline methods, our method can predict statutes that are more likely to appear in real judgments. Yi Feng 0005, Chuanyi Li, Jidong Ge, Bin Luo 0003, Vincent Ng 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2020 | Unveiling Hidden IntentionsabstractRecent years have seen significant advances in machine perception, which have enabled AI systems to become grounded in the world. While AI systems can now "read" and "see", they still cannot read between the lines and see through the lens, unlike humans. We propose the novel task of hidden message and intention identification: given some perceptual input (i.e., a text, an image), the goal is to produce a short description of the message the input transmits and the hidden intention of its author, if any. Not only will a solution to this task enable machine perception technologies to reach the next level of complexity, but it will be an important step towards addressing a task that has recently received a lot of public attention, political manipulation in social media. Gerardo Ocampo Diaz, Vincent Ng 0001 |
AAAI | 2 |
| 2020 | Bridging Resolution: A Survey of the State of the ArtabstractBridging reference resolution is an anaphora resolution task that is arguably more challenging and less studied than entity coreference resolution.Given that significant progress has been made on coreference resolution in recent years, we believe that bridging resolution will receive increasing attention in the NLP community.Nevertheless, progress on bridging resolution is currently hampered in part by the scarcity of large annotated corpora for model training as well as the lack of standardized evaluation protocols.This paper presents a survey of the current state of research on bridging reference resolution and discusses future research directions. Hideo Kobayashi, Vincent Ng 0001 |
COLING | 2 |
| 2020 | Identifying Exaggerated LanguageabstractWhile exaggeration is one of the most prevalent rhetorical devices, it is arguably one of the least studied in the figurative language processing community.We contribute to the computational study of exaggeration by (1) creating the first Chinese corpus focusing on sentence-level hyperbole detection, with the goal of facilitating a cross-lingual study on this phenomenon, (2) performing a statistical and manual analysis of our corpus, with the goal of gaining insights into the strategies humans employ when creating hyperboles, and (3) addressing the automatic hyperbole detection task with deep learning techniques. Chuanyi Li, Jidong Ge, Bin Luo 0003, Vincent Ng 0001 |
EMNLP (1) | 5 |
| 2020 | Conundrums in Entity Coreference Resolution: Making Sense of the State of the ArtabstractDespite the significant progress on entity coreference resolution observed in recent years, there is a general lack of understanding of what has been improved.We present an empirical analysis of state-of-the-art resolvers with the goal of providing the general NLP audience with a better understanding of the state of the art and coreference researchers with directions for future research. Vincent Ng 0001 |
EMNLP (1) | 2 |
| 2020 | Aspect-Based Sentiment Analysis as Fine-Grained Opinion MiningabstractWe show how the general fine-grained opinion mining concepts of opinion target and opinion expression are related to aspect-based sentiment analysis (ABSA) and discuss their benefits for resource creation over popular ABSA annotation schemes. Specifically, we first discuss why opinions modeled solely in terms of (entity, aspect) pairs inadequately captures the meaning of the sentiment originally expressed by authors and how opinion expressions and opinion targets can be used to avoid the loss of information. We then design a meaning-preserving annotation scheme and apply it to two popular ABSA datasets, the 2016 SemEval ABSA Restaurant and Laptop datasets. Finally, we discuss the importance of opinion expressions and opinion targets for next-generation ABSA systems. We make our datasets publicly available for download. Gerardo Ocampo Diaz, Xuanming Zhang, Vincent Ng 0001 |
LREC | 3 |
| 2020 | Unsupervised Argumentation Mining in Student EssaysabstractState-of-the-art systems for argumentation mining are supervised, thus relying on training data containing manually annotated argument components and the relationships between them. To eliminate the reliance on annotated data, we present a novel approach to unsupervised argument mining. The key idea is to bootstrap from a small set of argument components automatically identified using simple heuristics in combination with reliable contextual cues. Results on a Stab and Gurevych’s corpus of 402 essays show that our unsupervised approach rivals two supervised baselines in performance and achieves 73.5-83.7% of the performance of a state-of-the-art neural approach. Isaac Persing, Vincent Ng 0001 |
LREC | 2 |
| 2019 | Abstractive Summarization: A Survey of the State of the ArtabstractThe focus of automatic text summarization research has exhibited a gradual shift from extractive methods to abstractive methods in recent years, owing in part to advances in neural methods. Originally developed for machine translation, neural methods provide a viable framework for obtaining an abstract representation of the meaning of an input text and generating informative, fluent, and human-like summaries. This paper surveys existing approaches to abstractive summarization, focusing on the recently developed neural approaches. Vincent Ng 0001 |
AAAI | 2 |
| 2019 | Give Me More Feedback II: Annotating Thesis Strength and Related Attributes in Student EssaysabstractWhile the vast majority of existing work on automated essay scoring has focused on holistic scoring, researchers have recently begun work on scoring specific dimensions of essay quality.Nevertheless, progress in dimensionspecific essay scoring research is hindered in part by the lack of annotated corpora.To facilitate advances in this area of research, we design a rubric for scoring an important, yet unexplored dimension of persuasive essay quality, thesis strength, and annotate a corpus of essays with thesis strength scores.We additionally identify the attributes that could impact thesis strength and annotate the essays with the values of these attributes, which, when predicted by computational models, could provide feedback to students on why her essay receives a particular thesis strength score. Zixuan Ke, Hrishikesh Inamdar, Vincent Ng 0001 |
ACL (1) | 4 |
| 2019 | Automated Essay Scoring: A Survey of the State of the ArtabstractDespite being investigated for over 50 years, the task of automated essay scoring is far from being solved. Nevertheless, it continues to draw a lot of attention in the natural language processing community in part because of its commercial and educational values as well as the associated research challenges. This paper presents an overview of the major milestones made in automated essay scoring research since its inception. Zixuan Ke, Vincent Ng 0001 |
IJCAI | 2 |
| 2019 | Predicting Licenses for Changed Source CodeabstractOpen source software licenses regulate the circumstances under which software can be redistributed, reused and modified. Ensuring license compatibility and preventing license restriction conflicts among source code during software changes are the key to protect their commercial use. However, selecting the appropriate licenses for software changes requires lots of experience and manual effort that involve examining, assimilating and comparing various licenses as well as understanding their relationships with software changes. Worse still, there is no state-of-the-art methodology to provide this capability. Motivated by this observation, we propose in this paper Automatic License Prediction (ALP), a novel learning-based method and tool for predicting licenses as software changes. An extensive evaluation of ALP on predicting licenses in 700 open source projects demonstrate its effectiveness: ALP can achieve not only a high overall prediction accuracy (92.5% in micro F1 score) but also high accuracies across all license types. LiGuo Huang, Jidong Ge, Vincent Ng 0001 |
ASE | 4 |
| 2019 | Assessing the quality of the steps to reproduce in bug reportsabstractA major problem with user-written bug reports, indicated by developers and documented by researchers, is the (lack of high) quality of the reported steps to reproduce the bugs. Low-quality steps to reproduce lead to excessive manual effort spent on bug triage and resolution. This paper proposes Euler, an approach that automatically identifies and assesses the quality of the steps to reproduce in a bug report, providing feedback to the reporters, which they can use to improve the bug report. The feedback provided by Euler was assessed by external evaluators and the results indicate that Euler correctly identified 98% of the existing steps to reproduce and 58% of the missing ones, while 73% of its quality annotations are correct. Oscar Chaparro, Carlos Bernal-Cárdenas, Kevin Moran, Andrian Marcus, Massimiliano Di Penta, Denys Poshyvanyk, Vincent Ng 0001 |
ESEC/SIGSOFT FSE | 8 |
| 2018 | Give Me More Feedback: Annotating Argument Persuasiveness and Related Attributes in Student EssaysabstractWhile argument persuasiveness is one of the most important dimensions of argumentative essay quality, it is relatively little studied in automated essay scoring research.Progress on scoring argument persuasiveness is hindered in part by the scarcity of annotated corpora.We present the first corpus of essays that are simultaneously annotated with argument components, argument persuasiveness scores, and attributes of argument components that impact an argument's persuasiveness.This corpus could trigger the development of novel computational models concerning argument persuasiveness that provide useful feedback to students on why their arguments are (un)persuasive in addition to how persuasive they are. Winston Carlile, Nishant Gurrapadi, Zixuan Ke, Vincent Ng 0001 |
ACL (1) | 4 |
| 2018 | Modeling and Prediction of Online Product Review Helpfulness: A SurveyabstractAs the popularity of free-form usergenerated reviews in e-commerce and review websites continues to increase, there is a growing need for automatic mechanisms that sift through the vast number of reviews and identify quality content.Online review helpfulness modeling and prediction is a task which studies the factors that determine review helpfulness and attempts to accurately predict it.This survey paper provides an overview of the most relevant work on product review helpfulness prediction and understanding in the past decade, discusses gained insights, and provides guidelines for future research. Gerardo Ocampo Diaz, Vincent Ng 0001 |
ACL (1) | 2 |
| 2018 | Linking Source Code to Untangled Change IntentsabstractPrevious work [13] suggests that tangled changes (i.e., different change intents aggregated in one single commit message) could complicate tracing to different change tasks when developers manage software changes. Identifying links from changed source code to untangled change intents could help developers solve this problem. Manually identifying such links requires lots of experience and review efforts, however. Unfortunately, there is no automatic method that provides this capability. In this paper, we propose AutoCILink, which automatically identifies code to untangled change intent links with a pattern-based link identification system (AutoCILink-P) and a supervised learning-based link classification system (AutoCILink-ML). Evaluation results demonstrate the effectiveness of both systems: the pattern-based AutoCILink-P and the supervised learning-based AutoCILink-ML achieve average accuracy of 74.6% and 81.2%, respectively. LiGuo Huang, Chuanyi Liu, Vincent Ng 0001 |
ICSME | 4 |
| 2018 | Learning to Give Feedback: Modeling Attributes Affecting Argument Persuasiveness in Student EssaysabstractArgument persuasiveness is one of the most important dimensions of argumentative essay quality, yet it is little studied in automated essay scoring research. Using a recently released corpus of essays that are simultaneously annotated with argument components, argument persuasiveness scores, and attributes of argument components that impact an argument’s persuasiveness, we design and train the first set of neural models that predict the persuasiveness of an argument and its attributes in a student essay, enabling useful feedback to be provided to students on why their arguments are (un)persuasive in addition to how persuasive they are. Zixuan Ke, Winston Carlile, Nishant Gurrapadi, Vincent Ng 0001 |
IJCAI | 4 |
| 2018 | Event Coreference Resolution: A Survey of Two Decades of ResearchabstractRecent years have seen a gradual shift of focus from entity-based tasks to event-based tasks in information extraction research. Being a core event-based task, event coreference resolution is less studied but arguably more challenging than entity coreference resolution. This paper provides an overview of the major milestones made in event coreference research since its inception two decades ago. Vincent Ng 0001 |
IJCAI | 2 |
| 2018 | Effective API recommendation without historical software repositoriesabstractIt is time-consuming and labor-intensive to learn and locate the correct API for programming tasks. Thus, it is beneficial to perform API recommendation automatically. The graph-based statistical model has been shown to recommend top-10 API candidates effectively. It falls short, however, in accurately recommending an actual top-1 API. To address this weakness, we propose RecRank, an approach and tool that applies a novel ranking-based discriminative approach leveraging API usage path features to improve top-1 API recommendation. Empirical evaluation on a large corpus of (1385+8) open source projects shows that RecRank significantly improves top-1 API recommendation accuracy and mean reciprocal rank when compared to state-of-the-art API recommendation approaches. LiGuo Huang, Vincent Ng 0001 |
ASE | 3 |
| 2018 | Modeling Trolling in Social Media Conversations
Luis Gerardo Mojica, Vincent Ng 0001 |
LREC | 2 |
| 2018 | Improving Unsupervised Keyphrase Extraction using Background Knowledge
Vincent Ng 0001 |
LREC | 2 |
| 2018 | Automatically classifying user requests in crowdsourcing requirements engineering
Chuanyi Li, LiGuo Huang, Jidong Ge, Bin Luo 0003, Vincent Ng 0001 |
J. Syst. Softw. | 5 |
| 2017 | Machine Learning for Entity Coreference Resolution: A Retrospective Look at Two Decades of ResearchabstractThough extensively investigated since the 1960s, entity coreference resolution, a core task in natural language understanding, is far from being solved. Nevertheless, significant progress has been made on learning-based coreference research since its inception two decades ago. This paper provides an overview of the major milestones made in learning-based coreference research and discusses a hard entity coreference task, the Winograd Schema Challenge, which has recently received a lot of attention in the AI community. Vincent Ng 0001 |
AAAI | 1 |
| 2017 | Joint Learning for Event Coreference ResolutionabstractWhile joint models have been developed for many NLP tasks, the vast majority of event coreference resolvers, including the top-performing resolvers competing in the recent TAC KBP 2016 Event Nugget Detection and Coreference task, are pipelinebased, where the propagation of errors from the trigger detection component to the event coreference component is a major performance limiting factor.To address this problem, we propose a model for jointly learning event coreference, trigger detection, and event anaphoricity.Our joint model is novel in its choice of tasks and its features for capturing cross-task interactions.To our knowledge, this is the first attempt to train a mention-ranking model and employ event anaphoricity for event coreference.Our model achieves the best results to date on the KBP 2016 English and Chinese datasets. Vincent Ng 0001 |
ACL (1) | 2 |
| 2017 | Learning Antecedent Structures for Event Coreference ResolutionabstractThe vast majority of existing work on learning-based event coreference resolution has employed the so-called mentionpair model, which is a binary classifier that determines whether two event mentions are coreferent. Though conceptually simple, this model is known to suffer from several major weaknesses. Rather than making pairwise local decisions, we view event coreference as a structured prediction task, where we propose a probabilistic model that selects an antecedent for each event mention in a given document in a collective manner. Our model achieves the best results reported to date on the new KBP 2016 English and Chinese event coreference resolution datasets. Vincent Ng 0001 |
ICMLA | 2 |
| 2017 | Why Can't You Convince Me? Modeling Weaknesses in Unpersuasive ArgumentsabstractRecent work on argument persuasiveness has focused on determining how persuasive an argument is. Oftentimes, however, it is equally important to understand why an argument is unpersuasive, as it is difficult for an author to make her argument more persuasive unless she first knows what errors made it unpersuasive. Motivated by this practical concern, we (1) annotate a corpus of debate comments with not only their persuasiveness scores but also the errors they contain, (2) propose an approach to persuasiveness scoring and error identification that outperforms competing baselines, and (3) show that the persuasiveness scores computed by our approach can indeed be explained by the errors it identifies. Isaac Persing, Vincent Ng 0001 |
IJCAI | 2 |
| 2017 | Lightly-Supervised Modeling of Argument PersuasivenessabstractWe propose the first lightly-supervised approach to scoring an argument’s persuasiveness. Key to our approach is the novel hypothesis that lightly-supervised persuasiveness scoring is possible by explicitly modeling the major errors that negatively impact persuasiveness. In an evaluation on a new annotated corpus of online debate arguments, our approach rivals its fully-supervised counterparts in performance by four scoring metrics when using only 10% of the available training instances. Isaac Persing, Vincent Ng 0001 |
IJCNLP(1) | 2 |
| 2017 | Tracing requirements in software designabstractSoftware requirement analysis is an essential step in software development process, which defines what is to be built in a project. Requirements are mostly written in text and will later evolve to fine-grained and actionable artifacts with details about system configurations, technology stacks, etc. Tracing the evolution of requirements enables stakeholders to determine the origin of each requirement and understand how well the software's design reflects to its requirements. Reckoning requirements traceability is not a trivial task, we focus on applying machine learning approach to classify traceability between various associated requirements. In particular, we investigate a 2-learner, ontology-based approach, where we train two classifiers to separately exploit two types of features, lexical features and features derived from a hand-built ontology. In comparison to a supervised baseline system that uses only lexical features, our approach yields a relative error reduction of 25.9%. Most interestingly, results do not deteriorate when the hand-built ontology is replaced with its automatically constructed counterpart. Zeheng Li, LiGuo Huang, Vincent Ng 0001, Ruili Geng |
ICSSP | 4 |
| 2017 | Detecting missing information in bug descriptionsabstractBug reports document unexpected software behaviors experienced by users. To be effective, they should allow bug triagers to easily understand and reproduce the potential reported bugs, by clearly describing the Observed Behavior (OB), the Steps to Reproduce (S2R), and the Expected Behavior (EB). Unfortunately, while considered extremely useful, reporters often miss such pieces of information in bug reports and, to date, there is no effective way to automatically check and enforce their presence. We manually analyzed nearly 3k bug reports to understand to what extent OB, EB, and S2R are reported in bug reports and what discourse patterns reporters use to describe such information. We found that (i) while most reports contain OB (i.e., 93.5%), only 35.2% and 51.4% explicitly describe EB and S2R, respectively; and (ii) reporters recurrently use 154 discourse patterns to describe such content. Based on these findings, we designed and evaluated an automated approach to detect the absence (or presence) of EB and S2R in bug descriptions. With its best setting, our approach is able to detect missing EB (S2R) with 85.9% (69.2%) average precision and 93.2% (83%) average recall. Our approach intends to improve bug descriptions quality by alerting reporters about missing EB and S2R at reporting time. Oscar Chaparro, Fiorella Zampetti, Laura Moreno, Massimiliano Di Penta, Andrian Marcus, Gabriele Bavota, Vincent Ng 0001 |
ESEC/SIGSOFT FSE | 8 |
| 2016 | Joint Inference over a Lightly Supervised Information Extraction Pipeline: Towards Event Coreference Resolution for Resource-Scarce LanguagesabstractWe address two key challenges in end-to-end event coreference resolution research: (1) the error propagation problem, where an event coreference resolver has to assume as input the noisy outputs produced by its upstream components in the standard information extraction (IE) pipeline; and (2) the data annotation bottleneck, where manually annotating data for all the components in the IE pipeline is prohibitively expensive. This is the case in the vast majority of the world's natural languages, where such annotated resources are not readily available. To address these problems, we propose to perform joint inference over a lightly supervised IE pipeline, where all the models are trained using either active learning or unsupervised learning. Using our approach, only 25% of the training sentences in the Chinese portion of the ACE 2005 corpus need to be annotated with entity and event mentions in order for our event coreference resolver to surpass its fully supervised counterpart in performance. Chen Chen 0004, Vincent Ng 0001 |
AAAI | 2 |
| 2016 | Chinese Zero Pronoun Resolution with Deep Neural NetworksabstractWhile unsupervised anaphoric zero pronoun (AZP) resolvers have recently been shown to rival their supervised counterparts in performance, it is relatively difficult to scale them up to reach the next level of performance due to the large amount of feature engineering efforts involved and their ineffectiveness in exploiting lexical features.To address these weaknesses, we propose a supervised approach to AZP resolution based on deep neural networks, taking advantage of their ability to learn useful task-specific representations and effectively exploit lexical features via word embeddings.Our approach achieves stateof-the-art performance when resolving the Chinese AZPs in the OntoNotes corpus. Chen Chen 0004, Vincent Ng 0001 |
ACL (1) | 2 |
| 2016 | Modeling Stance in Student EssaysabstractEssay stance classification, the task of determining how much an essay's author agrees with a given proposition, is an important yet under-investigated subtask in understanding an argumentative essay's overall content.We introduce a new corpus of argumentative student essays annotated with stance information and propose a computational model for automatically predicting essay stance.In an evaluation on 826 essays, our approach significantly outperforms four baselines, one of which relies on features previously developed specifically for stance classification in student essays, yielding relative error reductions of at least 11.3% and 5.3%, in micro and macro F-score, respectively. Isaac Persing, Vincent Ng 0001 |
ACL (1) | 2 |
| 2016 | Joint Inference for Event Coreference ResolutionabstractEvent coreference resolution is a challenging problem since it relies on several components of the information extraction pipeline that typically yield noisy outputs. We hypothesize that exploiting the inter-dependencies between these components can significantly improve the performance of an event coreference resolver, and subsequently propose a novel joint inference based event coreference resolver using Markov Logic Networks (MLNs). However, the rich features that are important for this task are typically very hard to explicitly encode as MLN formulas since they significantly increase the size of the MLN, thereby making joint inference and learning infeasible. To address this problem, we propose a novel solution where we implicitly encode rich features into our model by augmenting the MLN distribution with low dimensional unit clauses. Our approach achieves state-of-the-art results on two standard evaluation corpora. Deepak Venugopal, Vibhav Gogate, Vincent Ng 0001 |
COLING | 4 |
| 2016 | Event Coreference Resolution with Multi-Pass Sieves
Vincent Ng 0001 |
LREC | 2 |
| 2016 | Markov Logic Networks for Text Mining: A Qualitative and Empirical Comparison with Integer Linear Programming
Luis Gerardo Mojica, Vincent Ng 0001 |
LREC | 2 |
| 2016 | End-to-End Argumentation Mining in Student EssaysabstractUnderstanding the argumentative structure of a persuasive essay involves addressing two challenging tasks: identifying the components of the essay's argument and identifying the relations that occur between them.We examine the under-investigated task of end-toend argument mining in persuasive student essays, where we (1) present the first results on end-to-end argument mining in student essays using a pipeline approach; (2) address error propagation inherent in the pipeline approach by performing joint inference over the outputs of the tasks in an Integer Linear Programming (ILP) framework; and (3) propose a novel objective function that enables F-score to be maximized directly by an ILP solver.We evaluate our joint-inference approach with our novel objective function on a publiclyavailable corpus of 90 essays, where it yields an 18.5% relative error reduction in F-score over the pipeline system. Isaac Persing, Vincent Ng 0001 |
HLT-NAACL | 2 |
| 2015 | Chinese Common Noun Phrase Resolution: An Unsupervised Probabilistic Model Rivaling Supervised ResolversabstractPronoun resolution and common noun phrase resolution are the two most challenging subtasks of coreference resolution. While a lot of work has focused on pronoun resolution, common noun phrase resolution has almost always been tackled in the context of the larger coreference resolution task. In fact, to our knowledge, there has been no attempt to address Chinese common noun phrase resolution as a standalone task. In this paper, we propose a generative model for unsupervised Chinese common noun phrase resolution that not only allows easy incorporation of linguistic constraints on coreference but also performs joint resolution and anaphoricity determination. When evaluated on the Chinese portion of the OntoNotes 5.0 corpus, our model rivals its supervised counterpart in performance. Chen Chen 0004, Vincent Ng 0001 |
AAAI | 2 |
| 2015 | Modeling Argument Strength in Student EssaysabstractIsaac Persing, Vincent Ng. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Isaac Persing, Vincent Ng 0001 |
ACL (1) | 2 |
| 2015 | Recovering Traceability Links in Requirements DocumentsabstractSoftware system development is guided by the evolution of requirements.In this paper, we address the task of requirements traceability, which is concerned with providing bi-directional traceability between various requirements, enabling users to find the origin of each requirement and track every change made to it.We propose a knowledge-rich approach to the task, where we extend a supervised baseline system with (1) additional training instances derived from human-provided annotator rationales; and (2) additional features derived from a hand-built ontology.Experiments demonstrate that our approach yields a relative error reduction of 11.1-19.7%. Zeheng Li, LiGuo Huang, Vincent Ng 0001 |
CoNLL | 4 |
| 2015 | Sieve-Based Spatial Relation Extraction with Expanding Parse TreesabstractA key challenge introduced by the recent SpaceEval shared task on spatial relation extraction is the identification of MOVELINKs, a type of spatial relation in which up to eight spatial elements can participate.To handle the complexity of extracting MOVELINKs, we combine two ideas that have been successfully applied to information extraction tasks, namely tree kernels and multi-pass sieves, proposing the use of an expanding parse tree as a novel structured feature for training MOVELINK classifiers.Our approach yields state-of-the-art results on two key tasks in SpaceEval. Jennifer D'Souza 0001, Vincent Ng 0001 |
EMNLP | 2 |
| 2015 | Fine-Grained Opinion Extraction with Markov Logic NetworksabstractMarkov Logic Networks, a joint inference framework that combines logical and probabilistic representations, enable effective modeling of the dependencies that exist between different instances of a data sample. While its ability to capture relational dependencies makes it an ideal framework for predicting the structures inherent in many natural language processing (NLP) tasks, it is arguably underused in NLP, especially in comparison to other joint inference frameworks such as integer linear programming. In this paper, we present the first Markov logic model for the NLP task of fine-grained opinion extraction that exploits a factuality lexicon. When evaluated on a standard evaluation corpus, our approach surpasses a state-of-the-art approach in performance. Luis Gerardo Mojica, Vincent Ng 0001 |
ICMLA | 2 |
| 2015 | Chinese Event Coreference Resolution: An Unsupervised Probabilistic Model Rivaling Supervised ResolversabstractRecent work has successfully leveraged the semantic information extracted from lexical knowledge bases such as WordNet and FrameNet to improve English event coreference resolvers.The lack of comparable resources in other languages, however, has made the design of high-performance non-English event coreference resolvers, particularly those employing unsupervised models, very difficult.We propose a generative model for the under-studied task of Chinese event coreference resolution that rivals its supervised counterparts in performance when evaluated on the ACE 2005 corpus. Chen Chen 0004, Vincent Ng 0001 |
HLT-NAACL | 2 |
| 2015 | AutoODC: Automated generation of orthogonal defect classifications
LiGuo Huang, Vincent Ng 0001, Isaac Persing, Zeheng Li, Ruili Geng, Jeff Tian |
Autom. Softw. Eng. | 2 |
| 2015 | SMPLearner: learning to predict software maintainability
LiGuo Huang, Vincent Ng 0001, Jidong Ge |
Autom. Softw. Eng. | 3 |
| 2014 | Chinese Overt Pronoun Resolution: A Bilingual ApproachabstractMuch research has been done on the problem of English pronoun resolution, but there has been relatively little work on the corresponding problem of Chinese pronoun resolution. While pronoun resolution in both languages remains a challenging task, Chinese pronoun resolution is further complicated by (1) the lack of publicly available Chinese word lists or dictionaries that can be used to look up essential mention attributes such as gender and number; and (2) the relative dearth of Chinese coreference-annotated data. Existing approaches to Chinese pronoun resolution are monolingual, training and testing a pronoun resolver on Chinese data. In contrast, we propose a bilingual approach to Chinese pronoun resolution, aiming to improve the resolution of Chinese pronouns by leveraging the publicly available English dictionaries and coreference annotations. Experiments on the OntoNotes 5.0 corpus demonstrate that our bilingual approach to Chinese pronoun resolution significantly surpasses the performance of state-of-the-art monolingual approaches. Chen Chen 0004, Vincent Ng 0001 |
AAAI | 2 |
| 2014 | Chinese Zero Pronoun Resolution: An Unsupervised Approach Combining Ranking and Integer Linear ProgrammingabstractState-of-the-art approaches to Chinese zero pronoun resolution are supervised, requiring training documents with manually resolved zero pronouns. To eliminate the reliance on annotated data, we propose an unsupervised approach to this task. Underlying our approach is the novel idea of employing a model trained on manually resolved overt pronouns to resolve zero pronouns. Experimental results on the OntoNotes 5.0 corpus are encouraging: our unsupervised model surpasses its supervised counterparts in performance. Chen Chen 0004, Vincent Ng 0001 |
AAAI | 2 |
| 2014 | Automatic Keyphrase Extraction: A Survey of the State of the ArtabstractWhile automatic keyphrase extraction has been examined extensively, state-of-the-art performance on this task is still much lower than that on many core natural lan-guage processing tasks. We present a sur-vey of the state of the art in automatic keyphrase extraction, examining the major sources of errors made by existing systems and discussing the challenges ahead. 1 Kazi Saidul Hasan, Vincent Ng 0001 |
ACL (1) | 2 |
| 2014 | Modeling Prompt Adherence in Student EssaysabstractRecently, researchers have begun exploring methods of scoring student essays with respect to particular dimensions of quality such as coherence, technical errors, and prompt adherence.The work on modeling prompt adherence, however, has been focused mainly on whether individual sentences adhere to the prompt.We present a new annotated corpus of essaylevel prompt adherence scores and propose a feature-rich approach to scoring essays along the prompt adherence dimension.Our approach significantly outperforms a knowledge-lean baseline prompt adherence scoring system yielding improvements of up to 16.6%. Isaac Persing, Vincent Ng 0001 |
ACL (1) | 2 |
| 2014 | Ensemble-Based Medical Relation Classification
Jennifer D'Souza 0001, Vincent Ng 0001 |
COLING | 2 |
| 2014 | Chinese Zero Pronoun Resolution: An Unsupervised Probabilistic Model Rivaling Supervised ResolversabstractState-of-the-art Chinese zero pronoun res-olution systems are supervised, thus re-lying on training data containing manu-ally resolved zero pronouns. To elimi-nate the reliance on annotated data, we present a generative model for unsuper-vised Chinese zero pronoun resolution. At the core of our model is a novel hy-pothesis: a probabilistic pronoun resolver trained on overt pronouns in an unsuper-vised manner can be used to resolve zero pronouns. Experiments demonstrate that our unsupervised model rivals its state-of-the-art supervised counterparts in perfor-mance when resolving the Chinese zero pronouns in the OntoNotes corpus. 1 Chen Chen 0004, Vincent Ng 0001 |
EMNLP | 2 |
| 2014 | Why are You Taking this Stance? Identifying and Classifying Reasons in Ideological DebatesabstractRecent years have seen a surge of interest in stance classification in online debates.Oftentimes, however, it is important to determine not only the stance expressed by an author in her debate posts, but also the reasons behind her supporting or opposing the issue under debate.We therefore examine the new task of reason classification in this paper.Given the close interplay between stance classification and reason classification, we design computational models for examining how automatically computed stance information can be profitably exploited for reason classification.Experiments on our reason-annotated corpus of ideological debate posts from four domains demonstrate that sophisticated models of stances and reasons can indeed yield more accurate reason and stance classification results than their simpler counterparts. Kazi Saidul Hasan, Vincent Ng 0001 |
EMNLP | 2 |
| 2014 | Vote Prediction on Comments in Social PollsabstractA poll consists of a question and a set of predefined answers from which voters can select. We present the new problem of vote prediction on comments, which involves determining which of these answers a voter selected given a comment she wrote after voting. To address this task, we ex-ploit not only the information extracted from the comments but also extra-textual information such as user demographic in-formation and inter-comment constraints. In an evaluation involving nearly one mil-lion comments collected from the popu-lar SodaHead social polling website, we show that a vote prediction system that ex-ploits only textual information can be im-proved significantly when extended with extra-textual information. 1 Isaac Persing, Vincent Ng 0001 |
EMNLP | 2 |
| 2014 | Relieving the Computational Bottleneck: Joint Inference for Event Extraction with High-Dimensional FeaturesabstractSeveral state-of-the-art event extraction systems employ models based on Support Vector Machines (SVMs) in a pipeline architecture, which fails to exploit the joint dependencies that typically exist among events and arguments.While there have been attempts to overcome this limitation using Markov Logic Networks (MLNs), it remains challenging to perform joint inference in MLNs when the model encodes many high-dimensional sophisticated features such as those essential for event extraction.In this paper, we propose a new model for event extraction that combines the power of MLNs and SVMs, dwarfing their limitations.The key idea is to reliably learn and process high-dimensional features using SVMs; encode the output of SVMs as low-dimensional, soft formulas in MLNs; and use the superior joint inferencing power of MLNs to enforce joint consistency constraints over the soft formulas.We evaluate our approach for the task of extracting biomedical events on the BioNLP 2013BioNLP , 2011BioNLP and 2009 Genia shared task datasets.Our approach yields the best F1 score to date on the BioNLP'13 (53.61) and BioNLP'11 (58.07) datasets and the second-best F1 score to date on the BioNLP'09 dataset (58.16). Deepak Venugopal, Chen Chen 0004, Vibhav Gogate, Vincent Ng 0001 |
EMNLP | 4 |
| 2014 | SinoCoreferencer: An End-to-End Chinese Event Coreference Resolver
Chen Chen 0004, Vincent Ng 0001 |
LREC | 2 |
| 2014 | Annotating Inter-Sentence Temporal Relations in Clinical Notes
Jennifer D'Souza 0001, Vincent Ng 0001 |
LREC | 2 |
| 2013 | Modeling Thesis Clarity in Student Essays
Isaac Persing, Vincent Ng 0001 |
ACL (1) | 2 |
| 2013 | Frame Semantics for Stance Classification
Kazi Saidul Hasan, Vincent Ng 0001 |
CoNLL | 2 |
| 2013 | Chinese Zero Pronoun Resolution: Some Recent AdvancesabstractWe extend Zhao and Ng's (2007) Chinese anaphoric zero pronoun resolver by (1) using a richer set of features and (2) exploiting the coreference links between zero pronouns during resolution.Results on OntoNotes show that our approach significantly outperforms two state-of-the-art anaphoric zero pronoun resolvers.To our knowledge, this is the first work to report results obtained by an end-toend Chinese zero pronoun resolver. Chen Chen 0004, Vincent Ng 0001 |
EMNLP | 2 |
| 2013 | Fast Strong Planning for FOND Problems with Multi-root Directed Acyclic GraphsabstractWe present a planner for addressing a difficult, yet under-investigated class of planning problems: Fully Observable Non-Deterministic planning problems with strong solutions. Our strong planner employs a new data structure, MRDAG (multi-root directed acyclic graph), to define how the solution space should be expanded. We further equip a MRDAG with heuristics to ensure planning towards the relevant search direction. We performed extensive experiments to evaluate MRDAG and the heuristics. Results show that our strong algorithm achieves impressive performance on a variety of benchmark problems: on average it runs more than three orders of magnitude faster than the state-of-the-art planners, MBP and Gamer, and demonstrates significantly better scalability. Jicheng Fu, Andres Calderon Jaramillo, Vincent Ng 0001, Farokh B. Bastani, I-Ling Yen |
ICTAI | 3 |
| 2013 | Chinese Event Coreference Resolution: Understanding the State of the Art
Chen Chen 0004, Vincent Ng 0001 |
IJCNLP | 2 |
| 2013 | Linguistically Aware Coreference Evaluation Metrics
Chen Chen 0004, Vincent Ng 0001 |
IJCNLP | 2 |
| 2013 | Stance Classification of Ideological Debates: Data, Models, Features, and Constraints
Kazi Saidul Hasan, Vincent Ng 0001 |
IJCNLP | 2 |
| 2013 | Classifying Temporal Relations with Rich Linguistic Knowledge
Jennifer D'Souza 0001, Vincent Ng 0001 |
HLT-NAACL | 2 |
| 2013 | Classifying temporal relations in clinical data: A hybrid, knowledge-rich approach
Jennifer D'Souza 0001, Vincent Ng 0001 |
J. Biomed. Informatics | 2 |
| 2012 | Clustering Documents Along Multiple DimensionsabstractTraditional clustering algorithms are designed to search for a single clustering solution despite the fact that multiple alternative solutions might exist for a particular dataset. For example, a set of news articles might be clustered by topic or by the author's gender or age. Similarly, book reviews might be clustered by sentiment or comprehensiveness. In this paper, we address the problem of identifying alternative clustering solutions by developing a Probabilistic Multi-Clustering (PMC) model that discovers multiple, maximally different clusterings of a data sample. Empirical results on six datasets representative of real-world applications show that our PMC model exhibits superior performance to comparable multi-clustering algorithms. Sajib Dasgupta, Richard M. Golden, Vincent Ng 0001 |
AAAI | 3 |
| 2012 | Joint Modeling for Chinese Event Extraction with Rich Linguistic Features
Chen Chen 0004, Vincent Ng 0001 |
COLING | 2 |
| 2012 | Learning the Fine-Grained Information Status of Discourse Entities
Altaf Rahman, Vincent Ng 0001 |
EACL | 2 |
| 2012 | Resolving Complex Cases of Definite Pronouns: The Winograd Schema Challenge
Altaf Rahman, Vincent Ng 0001 |
EMNLP-CoNLL | 2 |
| 2012 | Handling Planning Failures with Virtual ActionsabstractArtificial intelligence (AI) planners have been widely used in many fields, such as intelligent agents, autonomous robots, web service compositions, etc. However, existing AI planners share a common problem: When given a problem to solve, they either return a solution if one exists or report that no solution is found. However, simply reporting failure leaves no clues for people to trace the causes of the planning failure. In this paper, we present a novel approach that can propose virtual actions in the event of planning failure. Virtual actions enable traditional planners to succeed and hence return an incomplete plan instead of merely an error message. More importantly, the specifications of the virtual actions suggest what the missing parts may contain, thus providing important clues to users as to the nature of the failure. Experimental results show that our approach constantly returns useful and comprehensible information for humans, thus making AI planning more practical when solving real-world problems. Jicheng Fu, Sijie Tian, Vincent Ng 0001, Farokh B. Bastani, I-Ling Yen |
ICTAI | 3 |
| 2012 | Translation-Based Projection for Multilingual Coreference Resolution
Altaf Rahman, Vincent Ng 0001 |
HLT-NAACL | 2 |
| 2011 | Coreference Resolution with World Knowledge
Altaf Rahman, Vincent Ng 0001 |
ACL | 2 |
| 2011 | Learning the Information Status of Noun Phrases in Spoken Dialogues
Altaf Rahman, Vincent Ng 0001 |
EMNLP | 2 |
| 2011 | Learning Cause Identifiers from Annotator RationalesabstractIn the aviation safety research domain, cause identification refers to the task of identifying the possible causes responsible for the incident described in an aviation safety incident report. This task presents a number of challenges, including the scarcity of labeled data and the difficulties in finding the relevant portions of the text. We investigate the use of annotator rationales to overcome these challenges, proposing several new ways of utilizing rationales and showing that through judicious use of the rationales, it is possible to achieve significant improvement over a unigram SVM baseline. Muhammad Arshad Ul Abedin, Vincent Ng 0001, Latifur Khan |
IJCAI | 2 |
| 2011 | Simple and Fast Strong Cyclic Planning for Fully-Observable Nondeterministic Planning ProblemsabstractWe address a difficult, yet under-investigated class of planning problems: fully-observable nondeterministic (FOND) planning problems with strong cyclic solutions. The difficulty of these strong cyclic FOND planning problems stems from the large size of the state space. Hence, to achieve efficient planning, a planner has to cope with the explosion in the size of the state space by planning along the directions that allow the goal to be reached quickly. A major challenge is: how would one know which states and search directions are relevant before the search for a solution has even begun? We first describe an NDP-motivated strong cyclic algorithm that, without addressing the above challenge, can already outperform state-of-the-art FOND planners, and then extend this NDP-motivated planner with a novel heuristic that addresses the challenge. 1 Jicheng Fu, Vincent Ng 0001, Farokh B. Bastani, I-Ling Yen |
IJCAI | 2 |
| 2011 | Ensemble-Based Coreference ResolutionabstractWe investigate new methods for creating and applying ensembles for coreference resolution. While existing ensembles for coreference resolution are typically created using different learning algorithms, clustering algorithms or training sets, we harness recent advances in coreference modeling and propose to create our ensemble from a variety of supervised coreference models. However, the presence of pairwise and non-pairwise coreference models in our ensemble presents a challenge to its application: it is not immediately clear how to combine the coreference decisions made by these models. We investigate different methods for applying a model-heterogeneous ensemble for coreference resolution. Empirical results on the ACE data sets demonstrate the promise of ensemble approaches: all ensemble-based systems significantly outperform the best member of the ensemble. Altaf Rahman, Vincent Ng 0001 |
IJCAI | 2 |
| 2011 | Syntactic Parsing for Ranking-Based Coreference Resolution
Altaf Rahman, Vincent Ng 0001 |
IJCNLP | 2 |
| 2011 | AutoODC: Automated generation of Orthogonal Defect ClassificationsabstractOrthogonal Defect Classification (ODC), the most influential framework for software defect classification and analysis, provides valuable in-process feedback to system development and maintenance. Conducting ODC classification on existing organizational defect reports is human intensive and requires experts' knowledge of both ODC and system domains. This paper presents AutoODC, an approach and tool for automating ODC classification by casting it as a supervised text classification problem. Rather than merely apply the standard machine learning framework to this task, we seek to acquire a better ODC classification system by integrating experts' ODC experience and domain knowledge into the learning process via proposing a novel Relevance Annotation Framework. We evaluated AutoODC on an industrial defect report from the social network domain. AutoODC is a promising approach: not only does it leverage minimal human effort beyond the human annotations typically required by standard machine learning approaches, but it achieves an overall accuracy of 80.2% when using manual classifications as a basis of comparison. LiGuo Huang, Vincent Ng 0001, Isaac Persing, Ruili Geng, Jeff Tian |
ASE | 2 |
| 2011 | Narrowing the Modeling Gap: A Cluster-Ranking Approach to Coreference ResolutionabstractTraditional learning-based coreference resolvers operate by training the mention-pair model for determining whether two mentions are coreferent or not. Though conceptually simple and easy to understand, the mention-pair model is linguistically rather unappealing and lags far behind the heuristic-based coreference models proposed in the pre-statistical NLP era in terms of sophistication. Two independent lines of recent research have attempted to improve the mention-pair model, one by acquiring the mention-ranking model to rank preceding mentions for a given anaphor, and the other by training the entity-mention model to determine whether a preceding cluster is coreferent with a given mention. We propose a cluster-ranking approach to coreference resolution, which combines the strengths of the mention-ranking model and the entity-mention model, and is therefore theoretically more appealing than both of these models. In addition, we seek to improve cluster rankers via two extensions: (1) lexicalization and (2) incorporating knowledge of anaphoricity by jointly modeling anaphoricity determination and coreference resolution. Experimental results on the ACE data sets demonstrate the superior performance of cluster rankers to competing approaches as well as the effectiveness of our two extensions. Altaf Rahman, Vincent Ng 0001 |
J. Artif. Intell. Res. | 2 |
| 2010 | Supervised Noun Phrase Coreference Research: The First Fifteen Years
Vincent Ng 0001 |
ACL | 1 |
| 2010 | Inducing Fine-Grained Semantic Classes via Hierarchical and Collective Classification
Altaf Rahman, Vincent Ng 0001 |
COLING | 2 |
| 2010 | Modeling Organization in Student Essays
Isaac Persing, Alan Davis, Vincent Ng 0001 |
EMNLP | 3 |
| 2010 | Mining Clustering Dimensions
Sajib Dasgupta, Vincent Ng 0001 |
ICML | 2 |
| 2010 | Towards subjectifying text clusteringabstractAlthough it is common practice to produce only a single clustering of a dataset, in many cases text documents can be clustered along different dimensions. Unfortunately, not only do traditional text clustering algorithms fail to produce multiple clusterings of a dataset, the only clustering they produce may not be the one that the user desires. In this paper, we propose a simple active clustering algorithm that is capable of producing multiple clusterings of the same data according to user interest. In comparison to previous work on feedback-oriented clustering, the amount of user feedback required by our algorithm is minimal. In fact, the feedback turns out to be as simple as a cursory look at a list of words. Experimental results are very promising: our system is able to generate clusterings along the user-specified dimensions with reasonable accuracies on several challenging text classification tasks, thus providing suggestive evidence that our approach is viable. Sajib Dasgupta, Vincent Ng 0001 |
SIGIR | 2 |
| 2010 | Cause Identification from Aviation Safety Incident Reports via Weakly Supervised Semantic Lexicon ConstructionabstractThe Aviation Safety Reporting System collects voluntarily submitted reports on aviation safety incidents to facilitate research work aiming to reduce such incidents. To effectively reduce these incidents, it is vital to accurately identify why these incidents occurred. More precisely, given a set of possible causes, or shaping factors, this task of cause identification involves identifying all and only those shaping factors that are responsible for the incidents described in a report. We investigate two approaches to cause identification. Both approaches exploit information provided by a semantic lexicon, which is automatically constructed via Thelen and Riloff's Basilisk framework augmented with our linguistic and algorithmic modifications. The first approach labels a report using a simple heuristic, which looks for the words and phrases acquired during the semantic lexicon learning process in the report. The second approach recasts cause identification as a text classification problem, employing supervised and transductive text classification algorithms to learn models from incident reports labeled with shaping factors and using the models to label unseen reports. Our experiments show that both the heuristic-based approach and the learning-based approach (when given sufficient training data) outperform the baseline system significantly. Muhammad Arshad Ul Abedin, Vincent Ng 0001, Latifur Khan |
J. Artif. Intell. Res. | 2 |
| 2010 | Which Clustering Do You Want? Inducing Your Ideal Clustering with Minimal FeedbackabstractWhile traditional research on text clustering has largely focused on grouping documents by topic, it is conceivable that a user may want to cluster documents along other dimensions, such as the author's mood, gender, age, or sentiment. Without knowing the user's intention, a clustering algorithm will only group documents along the most prominent dimension, which may not be the one the user desires. To address the problem of clustering documents along the user-desired dimension, previous work has focused on learning a similarity metric from data manually annotated with the user's intention or having a human construct a feature space in an interactive manner during the clustering process. With the goal of reducing reliance on human knowledge for fine-tuning the similarity function or selecting the relevant features required by these approaches, we propose a novel active clustering algorithm, which allows a user to easily select the dimension along which she wants to cluster the documents by inspecting only a small number of words. We demonstrate the viability of our algorithm on a variety of commonly-used sentiment datasets. Sajib Dasgupta, Vincent Ng 0001 |
J. Artif. Intell. Res. | 2 |
| 2009 | Mine the Easy, Classify the Hard: A Semi-Supervised Approach to Automatic Sentiment Classification
Sajib Dasgupta, Vincent Ng 0001 |
ACL/IJCNLP | 2 |
| 2009 | Semi-Supervised Cause Identification from Aviation Safety Reports
Isaac Persing, Vincent Ng 0001 |
ACL/IJCNLP | 2 |
| 2009 | Weakly Supervised Part-of-Speech Tagging for Morphologically-Rich, Resource-Scarce Languages
Kazi Saidul Hasan, Vincent Ng 0001 |
EACL | 2 |
| 2009 | Learning-Based Named Entity Recognition for Morphologically-Rich, Resource-Scarce Languages
Kazi Saidul Hasan, Altaf Rahman, Vincent Ng 0001 |
EACL | 3 |
| 2009 | Topic-wise, Sentiment-wise, or Otherwise? Identifying the Hidden Dimension for Unsupervised Text Classification
Sajib Dasgupta, Vincent Ng 0001 |
EMNLP | 2 |
| 2009 | Supervised Models for Coreference Resolution
Altaf Rahman, Vincent Ng 0001 |
EMNLP | 2 |
| 2009 | Graph-Cut-Based Anaphoricity Determination for Coreference Resolution
Vincent Ng 0001 |
HLT-NAACL | 1 |
| 2008 | Unsupervised Models for Coreference Resolution
Vincent Ng 0001 |
EMNLP | 1 |
| 2008 | FIP: A Fast Planning-Graph-Based Iterative PlannerabstractWe present a fast iterative planner (FIP) that aims to handle planning problems involving nondeterministic actions. In contrast to existing iterative planners, FIP is built upon Graphplan's intrinsic features, thus, enabling Graphplan variants, including SGP and FF, to be enhanced with the capability of iterative planning. In addition, FIP is able to produce program-like plans with conditional and loop constructs, and achieves efficient planning via novel algorithms for manipulating planning graphs. Experimental results on several nondeterministic planning problems show that FIP is more efficient than the well-known planner, MBP, especially as the size of the problems increases. Jicheng Fu, Farokh B. Bastani, Vincent Ng 0001, I-Ling Yen, Yansheng Zhang |
ICTAI (1) | 3 |
| 2008 | Semisupervised Learning for Computational Linguistics Steven Abney (University of Michigan) Boca Raton, FL: Chapman & Hall / CRC (Computer science and data analysis series, edited by David Madigan et al.), 2007, xi+308 pp; hardbound, ISBN 978-1-58488-559-7abstractdo discuss self-training, they do so only in the context of Yarowsky's word sense disambiguation algorithm. Vincent Ng 0001 |
Comput. Linguistics | 1 |
| 2007 | Semantic Class Induction and Coreference Resolution
Vincent Ng 0001 |
ACL | 1 |
| 2007 | Unsupervised Part-of-Speech Acquisition for Resource-Scarce Languages
Sajib Dasgupta, Vincent Ng 0001 |
EMNLP-CoNLL | 2 |
| 2007 | Shallow Semantics for Coreference Resolution
Vincent Ng 0001 |
IJCAI | 1 |
| 2007 | High-Performance, Language-Independent Morphological Segmentation
Sajib Dasgupta, Vincent Ng 0001 |
HLT-NAACL | 2 |
| 2006 | Examining the Role of Linguistic Knowledge Sources in the Automatic Identification and Classification of Reviews
Vincent Ng 0001, Sajib Dasgupta, S. M. Niaz Arifin |
ACL | 1 |
| 2005 | Supervised Ranking for Pronoun Resolution: Some Recent Improvements
Vincent Ng 0001 |
AAAI | 1 |
| 2005 | Machine Learning for Coreference Resolution: From Local Classification to Global RankingabstractIn this paper, we view coreference resolution as a problem of ranking candidate partitions generated by different coreference systems. We propose a set of partition-based features to learn a ranking model for distinguishing good and bad partitions. Our approach compares favorably to two state-of-the-art coreference systems when evaluated on three standard coreference data sets. Vincent Ng 0001 |
ACL | 1 |
| 2004 | Learning Noun Phrase Anaphoricity to Improve Conference Resolution: Issues in Representation and OptimizationabstractKnowledge of the anaphoricity of a noun phrase might be profitably exploited by a coreference system to bypass the resolution of non-anaphoric noun phrases. Perhaps surprisingly, recent attempts to incorporate automatically acquired anaphoricity information into coreference systems, however, have led to the degradation in resolution performance. This paper examines several key issues in computing and using anaphoricity information to improve learning-based coreference systems. In particular, we present a new corpus-based approach to anaphoricity determination. Experiments on three standard coreference data sets demonstrate the effectiveness of our approach. Vincent Ng 0001 |
ACL | 1 |
| 2003 | Bootstrapping Coreference Classifiers with Multiple Machine Learning Algorithms
Vincent Ng 0001, Claire Cardie |
EMNLP | 1 |
| 2003 | Weakly Supervised Natural Language Learning Without Redundant Views
Vincent Ng 0001, Claire Cardie |
HLT-NAACL | 1 |
| 2002 | Improving Machine Learning Approaches to Coreference ResolutionabstractWe present a noun phrase coreference system that extends the work of Soon et al. (2001) and, to our knowledge, produces the best results to date on the MUC-6 and MUC-7 coreference resolution data sets -F-measures of 70.4 and 63.4,respectively.Improvements arise from two sources: extra-linguistic changes to the learning framework and a large-scale expansion of the feature set to include more sophisticated linguistic knowledge. Vincent Ng 0001, Claire Gardent |
ACL | 1 |
| 2002 | Identifying Anaphoric and Non-Anaphoric Noun Phrases to Improve Coreference Resolution
Vincent Ng 0001, Claire Cardie |
COLING | 1 |
| 2002 | Combining Sample Selection and Error-Driven Pruning for Machine Learning of Coreference RulesabstractMost machine learning solutions to noun phrase coreference resolution recast the problem as a classification task. We examine three potential problems with this reformulation, namely, skewed class distributions, the inclusion of "hard" training instances, and the loss of transitivity inherent in the original coreference relation. We show how these problems can be handled via intelligent sample selection and error-driven pruning of classification rule-sets. The resulting system achieves an F-measure of 69.5 and 63.4 on the MUC-6 and MUC-7 coreference resolution data sets, respectively, surpassing the performance of the best MUC-6 and MUC-7 coreference systems. In particular, the system outperforms the best-performing learning-based coreference system to date. Vincent Ng 0001, Claire Cardie |
EMNLP | 1 |