VLDB 2026 Research / reviewers in the wild / expert
Wenjuan Han
dblp:188/9071
· DBLP profile ↗
38ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STOLA: Self-Adaptive Touch-Language Framework for Tactile Commonsense Reasoning in Open-Ended ScenariosabstractThis paper explores the challenges of integrating tactile sensing into intelligent systems for multimodal reasoning, particularly in enabling commonsense reasoning about the open-ended physical world. We identify two key challenges: modality discrepancy, where existing touch-language models often treat touch as a mere sub-modality of language without further addressing the semantic differences, and open-ended tactile data scarcity, where current datasets lack the diversity, open-endedness, and complexity needed for reasoning. To overcome these challenges, we introduce SToLa, a Self-Adaptive Touch-Language framework. SToLa utilizes Mixture of Experts (MoE) to dynamically process, unify, and manage tactile and language modalities, capturing their unique characteristics. Crucially, we also present a comprehensive tactile commonsense reasoning dataset and benchmark featuring free-form questions and responses, 8 physical properties, 4 interactive characteristics, and diverse commonsense knowledge. Experiments show SToLa exhibits competitive performance compared to existing models on the PHYSICLEAR benchmark and self-constructed datasets, proving the effectiveness of the Mixture of Experts architecture in multimodal management and the performance advantages for open-scenario tactile commonsense reasoning tasks. Jin An Xu, Jialing Chen, Bin Fang 0003, Wenjuan Han |
AAAI | 5 |
| 2026 | EMAformer: Enhancing Transformer Through Embedding Armor for Time Series ForecastingabstractMultivariate time series forecasting is crucial across a wide range of domains. While presenting notable progress for the Transformer architecture, iTransformer still lags behind the latest MLP-based models. We attribute this performance gap to unstable inter-channel relationships. To bridge this gap, we propose EMAformer, a simple yet effective model that enhances the Transformer with an auxiliary embedding suite, akin to armor that reinforces its ability. By introducing three key inductive biases, i.e., global stability, phase sensitivity, and cross-axis specificity, EMAformer unlocks the further potential of the Transformer architecture, achieving state-of-the-art performance on 12 real-world benchmarks and reducing forecasting errors by an average of 2.73% in MSE and 5.15% in MAE. This significantly advances the practical applicability of Transformer-based approaches for multivariate time series forecasting. Xinyi Du, Xuanchi Guo, Wenjuan Han |
AAAI | 5 |
| 2026 | VisiFold: Long-Term Traffic Forecasting via Temporal Folding Graph and Node VisibilityabstractTraffic forecasting is a cornerstone of intelligent transportation systems. While existing research has made significant progress in short-term prediction, long-term forecasting remains a largely uncharted and challenging frontier. Extending the prediction horizon intensifies two critical issues: escalating computational resource consumption and increasingly complex spatial-temporal dependencies. Current approaches, which rely on spatial-temporal graphs and process temporal and spatial dimensions separately, suffer from snapshot-stacking inflation and cross-step fragmentation. To overcome these limitations, we propose \textit{VisiFold}. Our framework introduces a novel temporal folding graph that consolidates a sequence of temporal snapshots into a single graph. Furthermore, we present a node visibility mechanism that incorporates node-level masking and subgraph sampling to overcome the computational bottleneck imposed by large node counts. Extensive experiments show that VisiFold not only drastically reduces resource consumption but also outperforms existing baselines in long-term forecasting tasks. Remarkably, even with a high mask ratio of 80\%, VisiFold maintains its performance advantage. By effectively breaking the resource constraints in both temporal and spatial dimensions, our work paves the way for more realistic long-term traffic forecasting. The code is available at~ https://github.com/PlanckChang/VisiFold. Xinyi Du, Xuanchi Guo, Wenjuan Han |
ICDE | 5 |
| 2026 | GLINT: Global-local fusion with attention intervention for MLLM hallucination mitigation caused by insufficient visual resolution
Jin An Xu, Songming Zhang 0001, Chengkai Wang, Wenjuan Han |
Inf. Process. Manag. | 5 |
| 2026 | IntentQA: Intent Question Answering in Videos by Cognitive Context ReasoningabstractVideo understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions-often termed the "dark matter" of social intelligence. To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR(eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the lack of transparency in traditional models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets. Jiapeng Li 0003, Ping Wei 0001, Wenjuan Han, Song-Chun Zhu, Lifeng Fan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language ModelsabstractAlthough Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandatory tool use, which does not always align with real-world situations, where the necessity for tools is uncertain, and incorrect or unnecessary use of tools can damage the general abilities of LLMs. Therefore, we propose to explore whether LLMs can discern their ability boundaries and use tools flexibly. We then introduce the Whether-or-not tool usage Evaluation benchmark (WTU-Eval) to assess LLMs with eleven datasets, where six of them are tool-usage datasets, and five are general datasets. LLMs are prompted to use tools according to their needs. The results of eight LLMs on WTU-Eval reveal that LLMs frequently struggle to determine tool use in general datasets, and LLMs’ performance in tool-usage datasets improves when their ability is similar to ChatGPT. Jian Liu 0032, Kangyun Ning, Yisong Su, Wenjuan Han, Jin An Xu, Yuanzhe Zhang |
ICASSP | 4 |
| 2025 | Bridging the reality gap: A benchmark for physical reasoning in general world models with various physical phenomena beyond mechanics
Jin An Xu, Huiqi Hu, Xiuwen Xu, Zijian Jin, Fandong Meng, Jie Zhou 0016, Wenjuan Han |
Expert Syst. Appl. | 10 |
| 2024 | PersonalityScanner: Exploring the Validity of Personality Assessment Based on Multimodal Signals in Virtual Reality
Huiqi Hu, Xianhao Yu, Jin An Xu, Yujia Peng, Wenjuan Han |
CogSci | 9 |
| 2024 | CollabKG: A Learnable Human-Machine-Cooperative Information Extraction Toolkit for (Event) Knowledge Graph ConstructionabstractIn order to construct or extend entity-centric and event-centric knowledge graphs (KG and EKG), the information extraction (IE) annotation toolkit is essential. However, existing IE toolkits have several non-trivial problems, such as not supporting multi-tasks, and not supporting automatic updates. In this work, we present CollabKG, a learnable human-machine-cooperative IE toolkit for KG and EKG construction. Specifically, for the multi-task issue, CollabKG unifies different IE subtasks, including named entity recognition (NER), entity-relation triple extraction (RE), and event extraction (EE), and supports both KG and EKG. Then, combining advanced prompting-based IE technology, the human-machine-cooperation mechanism with Large Language Models (LLMs) as the assistant machine is presented which can provide a lower cost as well as a higher performance. Lastly, owing to the two-way interaction between the human and machine, CollabKG with learning ability allows self-renewal. Besides, CollabKG has several appealing features (e.g., customization, training-free, and label propagation) that make the system powerful and high-productivity. We holistically compare our toolkit with other existing tools on these features. Human evaluation quantitatively illustrates that CollabKG significantly improves annotation quality, efficiency, and stability simultaneously. Yufeng Chen 0005, Xingyu Cui, Jin An Xu, Wenjuan Han |
LREC/COLING | 6 |
| 2024 | CLOVA: A Closed-LOop Visual Assistant with Tool Usage and UpdateabstractUtilizing large language models (LLMs) to compose off-the-shelf visual tools represents a promising avenue of research for developing robust visual assistants capable of addressing diverse visual tasks. However, these methods often overlook the potential for continual learning, typically by freezing the utilized tools, thus limiting their adaptation to environments requiring new knowledge. To tackle this challenge, we propose CLOVA, a Closed-LOop Visual Assistant, which operates within a framework encompassing inference, reflection, and learning phases. During the inference phase, LLMs generate programs and execute corresponding tools to complete assigned tasks. In the reflection phase, a multimodal global-local reflection scheme analyzes human feedback to determine which tools require updating. Lastly, the learning phase employs three flexible approaches to automatically gather training data and introduces a novel prompt tuning scheme to update the tools, allowing CLOVA to efficiently acquire new knowledge. Experimental findings demonstrate that CLOVA surpasses existing tool-usage methods by 5% in visual question answering and multiple-image reasoning, by 10% in knowledge tagging, and by 20% in image editing. These results under-score the significance of the continual learning capability in general visual assistants. Zhi Gao 0002, Yuntao Du 0001, Xiaojian Ma 0001, Wenjuan Han, Song-Chun Zhu, Qing Li 0003 |
CVPR | 5 |
| 2024 | Empowering Vision-Language Models for Reasoning Ability through Large Language ModelsabstractVision-language models (VLM) have shown excellent performance in vision-language tasks. However, they sometimes lack sufficient reasoning ability. In contrast, large language models (LLMs) have emerged with powerful reasoning capabilities. Therefore, we propose a framework called TReE, which transfers the reasoning ability of the LLM to the VLM in learning-free settings. TReE is a three-stage framework: observation, thinking, and re-thinking. The observation stage requires the VLM to obtain overall visual information about the image. Then, the thinking stage combines the visual information and task description as the prompt for the LLM, allowing it to present the thinking process (namely, rationale). Lastly, the re-thinking stage learns useful information from the rationale and then predicts the final result using the VLM. We are the first to explore enhancing the VLM’s reasoning ability without any training, finetuning, or access to the LLM’s parameters, which we refer to as a plug-in mode, leading to the model-agnostic feature. Experiments show that TReE performed well on general visual questionanswering (VQA) tasks and outperformed KOSMOS-1 on the challenging Raven IQ test dataset by 6%. Furthermore, with additional lightweight finetuning using a smaller amount of parameters, TReE achieved a high accuracy of 81.7% on GQA and 67.3% on VQAv2. Jin An Xu, Wenjuan Han |
ICASSP | 4 |
| 2024 | MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningabstractSince the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity.
However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks.
In this paper, we address the limitation above by 1) introducing vision-language Model with **M**ulti-**M**odal **I**n-**C**ontext **L**earning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts.
Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context.
Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC. Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 0001, Kaikai An, Liang Chen 0024, Zixuan Liu 0001, Sheng Wang 0012, Wenjuan Han, Baobao Chang |
ICLR | 9 |
| 2024 | Model-Agnostic Knowledge Distillation Between Heterogeneous Models
Jiaxin Shen, Yanyao Liu, Yong Jiang 0005, Yufeng Chen 0005, Wenjuan Han |
NLPCC (1) | 5 |
| 2024 | LegalAsst: Human-centered and AI-empowered machine to enhance court productivity and legal assistance
Wenjuan Han, Jiaxin Shen, Yanyao Liu, Jin An Xu, Fangxu Hu, Xueli Yu, Huaqing Wang, Zhijing Liu, Yajie Yang, Tianshui Shi, Mengyao Ge |
Inf. Sci. | 1 |
| 2023 | Modeling Instance Interactions for Joint Information Extraction with Neural High-Order Conditional Random FieldabstractPrior works on joint Information Extraction (IE) typically model instance (e.g., event triggers, entities, roles, relations) interactions by representation enhancement, type dependencies scoring, or global decoding.We find that the previous models generally consider binary type dependency scoring of a pair of instances, and leverage local search such as beam search to approximate global solutions.To better integrate cross-instance interactions, in this work, we introduce a joint IE framework (CRFIE) that formulates joint IE as a high-order Conditional Random Field.Specifically, we design binary factors and ternary factors to directly model interactions between not only a pair of instances but also triplets.Then, these factors are utilized to jointly predict labels of all instances.To address the intractability problem of exact high-order inference, we incorporate a high-order neural decoder that is unfolded from a mean-field variational inference method, which achieves consistent learning and inference.The experimental results show that our approach achieves consistent improvements on three IE tasks compared with our baseline and prior work. Zixia Jia, Zhaohui Yan 0001, Wenjuan Han, Zilong Zheng, Kewei Tu |
ACL (1) | 3 |
| 2023 | Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationabstractSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, Jinan Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Songming Zhang 0001, Yunlong Liang, Shuaibo Wang, Yufeng Chen 0005, Wenjuan Han, Jian Liu 0032, Jin An Xu |
ACL (1) | 5 |
| 2023 | CORD: A Three-Stage Coarse-to-Fine Framework for Relation Detection in Knowledge Base Question AnsweringabstractAs a fundamental subtask of Knowledge Base Question Answering (KBQA), Relation Detection (KBQA-RD) plays a crucial role to detect the KB relations between entities or variables in natural language questions. It remains, however, a challenging task, particularly for significant large-scale relations and in the presence of easily confused relations. Recent state-of-the-art methods not only struggle with such scenarios, but often take into account only one facet and fail to incorporate the subtle discrepancy among the relations. In this paper, we propose a simple and efficient three-stage framework to exploit the coarse-to-fine paradigm. Specifically, we employ a natural clustering over all KB relations and perform a coarse-to-fine relation recognition process based on the relation clustering. In this way, our framework (i.e., CORD) refines the detection of relations, so as to scale well with large-scale relations. Experiments on both single-relation (i.e., SimpleQuestions (SQ)) and multi-relation (i.e., WebQSP (WQ)) benchmarks show that CORD not only achieves the outstanding relation detection performance in KBQA-RD subtask; but more importantly, further improves the accuracy of KBQA systems. Yanzeng Li, Sen Hu 0005, Wenjuan Han, Lei Zou 0001 |
CIKM | 3 |
| 2023 | A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase GenerationabstractExisting syntactically-controlled paraphrase generation (SPG) models perform promisingly with human-annotated or well-chosen syntactic templates.However, the difficulty of obtaining such templates actually hinders the practical application of SPG models.For one thing, the prohibitive cost makes it unfeasible to manually design decent templates for every source sentence.For another, the templates automatically retrieved by current heuristic methods are usually unreliable for SPG models to generate qualified paraphrases.To escape this dilemma, we propose a novel Quality-based Syntactic Template Retriever (QSTR) to retrieve templates based on the quality of the to-be-generated paraphrases.Furthermore, for situations requiring multiple paraphrases for each source sentence, we design a Diverse Templates Search (DTS) algorithm, which can enhance the diversity between paraphrases without sacrificing quality.Experiments demonstrate that QSTR can significantly surpass existing retrieval methods in generating high-quality paraphrases and even perform comparably with human-annotated templates in terms of reference-free metrics.Additionally, human evaluation and the performance on downstream tasks using our generated paraphrases for data augmentation showcase the potential of our QSTR and DTS algorithm in practical scenarios. Songming Zhang 0001, Yunlong Liang, Yufeng Chen 0005, Jian Liu 0032, Wenjuan Han, Jin An Xu |
EMNLP | 6 |
| 2023 | IntentQA: Context-aware Video Intent ReasoningabstractIn this paper, we propose a novel task IntentQA, a special VideoQA task focusing on video intent reasoning, which has become increasingly important for AI with its advantages in equipping AI agents with the capability of reasoning beyond mere recognition in daily tasks. We also contribute a large-scale VideoQA dataset for this task. We propose a Context-aware Video Intent Reasoning model (CaVIR) consisting of i) Video Query Language (VQL) for better cross-modal representation of the situational context, ii) Contrastive Learning module for utilizing the contrastive context, and iii) Commonsense Reasoning module for incorporating the commonsense context. Comprehensive experiments on this challenging task demonstrate the effectiveness of each model component, the superiority of our full model over other baselines, and the generalizability of our model to a new VideoQA task. The dataset and codes are open-sourced at: https://github.com/JoseponLee/IntentQA.git. Jiapeng Li 0003, Ping Wei 0001, Wenjuan Han, Lifeng Fan |
ICCV | 3 |
| 2023 | On the Complexity of Bayesian GeneralizationabstractWe examine concept generalization at a large scale in the natural visual spectrum. Established computational modes (*i.e.*, rule-based or similarity-based) are primarily studied isolated, focusing on confined and abstract problem spaces. In this work, we study these two modes when the *problem space* scales up and when the *complexity* of concepts becomes diverse. At the **representational level**, we investigate how the complexity varies when a visual concept is mapped to the representation space. Prior literature has shown that two types of complexities (Griffiths & Tenenbaum, 2003) build an inverted-U relation (Donderi, 2006; Sun & Firestone, 2021). Leveraging *Representativeness of Attribute* (RoA), we computationally confirm: Models use attributes with high RoA to describe visual concepts, and the description length falls in an inverted-U relation with the increment in visual complexity. At the **computational level**, we examine how the complexity of representation affects the shift between the rule- and similarity-based generalization. We hypothesize that category-conditioned visual modeling estimates the co-occurrence frequency between visual and categorical attributes, thus potentially serving as the prior for the natural visual world. Experimental results show that representations with relatively high subjective complexity outperform those with relatively low subjective complexity in rule-based generalization, while the trend is the opposite in similarity-based generalization. Yu-Zhe Shi, Manjie Xu, John E. Hopcroft, Kun He 0001, Josh Tenenbaum, Song-Chun Zhu, Ying Nian Wu, Wenjuan Han, Yixin Zhu 0001 |
ICML | 8 |
| 2023 | Evaluating and Inducing Personality in Pre-trained Language ModelsabstractStandardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the study of personality delves into how individuals differ in thinking, feeling, and behaving. Toward building and understanding human-like social machines, we are motivated to ask: Can we assess machine behaviors by leveraging human psychometric tests in a **principled** and **quantitative** manner? If so, can we induce a specific personality in LLMs? To answer these questions, we introduce the Machine Personality Inventory (MPI) tool for studying machine behaviors; MPI follows standardized
personality tests, built upon the Big Five Personality Factors (Big Five) theory and personality assessment inventories. By systematically evaluating LLMs with MPI, we provide the first piece of evidence demonstrating the efficacy of MPI in studying LLMs behaviors. We further devise a Personality Prompting (P$^2$) method to induce LLMs with specific personalities in a **controllable** way, capable of producing diverse and verifiable behaviors. We hope this work sheds light on future studies by adopting personality as the essential indicator for various downstream tasks, and could further motivate research into equally intriguing human-like machine behaviors. Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 4 |
| 2022 | On the Robustness of Question Rewriting Systems to Questions of Varying HardnessabstractIn conversational question answering (CQA), the task of question rewriting (QR) in context aims to rewrite a context-dependent question into an equivalent self-contained question that gives the same answer.In this paper, we are interested in the robustness of a QR system to questions varying in rewriting hardness or difficulty.Since there is a lack of questions classified based on their rewriting hardness, we first propose a heuristic method to automatically classify questions into subsets of varying hardness, by measuring the discrepancy between a question and its rewrite.To find out what makes questions hard or easy for rewriting, we then conduct a human evaluation to annotate the rewriting hardness of questions.Finally, to enhance the robustness of QR systems to questions of varying hardness, we propose a novel learning framework for QR that first trains a QR model independently on each subset of questions of a certain level of hardness, then combines these QR models as one joint model for inference.Experimental results on two datasets show that our framework improves the overall performance compared to the baselines 1 . Hai Ye, Hwee Tou Ng, Wenjuan Han |
ACL (1) | 3 |
| 2022 | Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsabstractUnderstanding realistic visual scene images together with language descriptions is a fundamental task towards generic visual understanding. Previous works have shown compelling comprehensive results by building hierarchical structures for visual scenes (e.g., scene graphs) and natural languages (e.g., dependency trees), individually. However, how to construct a joint vision-language (VL) structure has barely been investigated. More challenging but worthwhile, we introduce a new task that targets on inducing such a joint VL structure in an unsupervised manner. Our goal is to bridge the visual scene graphs and linguistic dependency trees seamlessly. Due to the lack of VL structural data, we start by building a new dataset VLParse. Rather than using labor-intensive labeling from scratch, we propose an automatic alignment procedure to produce coarse structures followed by human refinement to produce high-quality ones. Moreover, we benchmark our dataset by proposing a contrastive learning (CL)-based framework VLGAE, short for Vision-Language Graph Autoencoder. Our model obtains superior performance on two derived tasks, i.e., language grammar induction and VL phrase grounding. Ablations show the effectiveness of both visual cues and dependency relationships on fine-grained VL structure construction. Chao Lou, Wenjuan Han, Yuhuan Lin, Zilong Zheng |
CVPR | 2 |
| 2022 | Unsupervised Vision-Language Grammar Induction with Shared Structure Modeling
Wenjuan Han, Zilong Zheng, Tinne Tuytelaars |
ICLR | 2 |
| 2021 | Adapting Unsupervised Syntactic Parsing Methodology for Discourse Dependency ParsingabstractLiwen Zhang, Ge Wang, Wenjuan Han, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ge Wang 0005, Wenjuan Han, Kewei Tu |
ACL/IJCNLP (1) | 3 |
| 2021 | Diversity-Driven Combination for Grammatical Error CorrectionabstractGrammatical error correction (GEC) is the task of detecting and correcting errors in a written text. The idea of combining multiple system outputs has been successfully used in GEC. To achieve successful system combination, multiple component systems need to produce corrected sentences that are both diverse and of comparable quality. However, most existing state-of-the-art GEC approaches are based on similar sequence-to-sequence neural networks, so the gains are limited from combining the outputs of component systems similar to one another. In this paper, we present Diversity-Driven Combination (DDC) for GEC, a system combination strategy that encourages diversity among component systems. We evaluate our system combination strategy on the CoNLL-2014 shared task and the BEA-2019 shared task. On both benchmarks, DDC achieves significant performance gain with a small number of training examples and outperforms the component systems by a large margin. Our source code is available at https://github.com/nusnlp/gec-ddc. Wenjuan Han, Hwee Tou Ng |
ICTAI | 1 |
| 2020 | Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationabstractOpen-domain dialogue generation has gained increasing attention in Natural Language Processing.Its evaluation requires a holistic means.Human ratings are deemed as the gold standard.As human evaluation is inefficient and costly, an automated substitute is highly desirable.In this paper, we propose holistic evaluation metrics that capture different aspects of open-domain dialogues.Our metrics consist of (1) GPT-2 based context coherence between sentences in a dialogue, (2) GPT-2 based fluency in phrasing, (3) n-gram based diversity in responses to augmented queries, and (4) textual-entailment-inference based logical self-consistency.The empirical validity of our metrics is demonstrated by strong correlations with human judgments.We open source the code and relevant materials.1 Bo Pang 0004, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Kewei Tu |
ACL | 3 |
| 2020 | A Survey of Unsupervised Dependency ParsingabstractSyntactic dependency parsing is an important task in natural language processing.Unsupervised dependency parsing aims to learn a dependency parser from sentences that have no annotation of their correct parse trees.Despite its difficulty, unsupervised parsing is an interesting research direction because of its capability of utilizing almost unlimited unannotated text data.It also serves as the basis for other research in low-resource parsing.In this paper, we survey existing approaches to unsupervised dependency parsing, identify two major classes of approaches, and discuss recent trends.We hope that our survey can provide insights for researchers and facilitate future research on this topic. Wenjuan Han, Yong Jiang 0005, Hwee Tou Ng, Kewei Tu |
COLING | 1 |
| 2020 | Second-Order Unsupervised Neural Dependency ParsingabstractMost of the unsupervised dependency parsers are based on first-order probabilistic generative models that only consider local parent-child information.Inspired by second-order supervised dependency parsing, we proposed a second-order extension of unsupervised neural dependency models that incorporate grandparent-child or sibling information.We also propose novel design of the neural parameterization and optimization methods of the dependency models.In secondorder models, the number of grammar rules grows cubically with the increase of vocabulary size, making it difficult to train lexicalized models that may contain thousands of words.To circumvent this problem while still benefiting from both second-order parsing and lexicalization, we use the agreement-based learning framework to jointly train a second-order unlexicalized model and a first-order lexicalized model.Experiments on multiple datasets show the effectiveness of our second-order models compared with recent state-of-the-art methods.Our joint model achieves a 10% improvement over the previous state-of-the-art parser on the full WSJ test set. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
COLING | 3 |
| 2020 | ToHRE: A Top-Down Classification Strategy with Hierarchical Bag Representation for Distantly Supervised Relation ExtractionabstractDistantly Supervised Relation Extraction (DSRE) has proven to be effective to find relational facts from texts, but it still suffers from two main problems: the wrong labeling problem and the long-tail problem.Most of the existing approaches address these two problems through flat classification, which lacks hierarchical information of relations.To leverage the informative relation hierarchies, we formulate DSRE as a hierarchical classification task and propose a novel hierarchical classification framework, which extracts the relation in a top-down manner.Specifically, in our proposed framework, 1) we use a hierarchically-refined representation method to achieve hierarchy-specific representation; 2) a top-down classification strategy is introduced instead of training a set of local classifiers.The experiments on NYT dataset demonstrate that our approach significantly outperforms other state-of-the-art approaches, especially for the long-tail problem. Erxin Yu, Wenjuan Han, Yuan Tian 0016, Yi Chang 0001 |
COLING | 2 |
| 2020 | Adversarial Attack and Defense of Structured Prediction ModelsabstractBuilding an effective adversarial attacker and elaborating on countermeasures for adversarial attacks for natural language processing (NLP) have attracted a lot of research in recent years.However, most of the existing approaches focus on classification problems.In this paper, we investigate attacks and defenses for structured prediction tasks in NLP.Besides the difficulty of perturbing discrete words and the sentence fluency problem faced by attackers in any NLP tasks, there is a specific challenge to attackers of structured prediction models: the structured output of structured prediction models is sensitive to small perturbations in the input.To address these problems, we propose a novel and unified framework that learns to attack a structured prediction model using a sequence-to-sequence model with feedbacks from multiple reference models of the same structured prediction task.Based on the proposed attack, we further reinforce the victim model with adversarial training, making its prediction more robust and accurate.We evaluate the proposed framework in dependency parsing and part-of-speech tagging.Automatic and human evaluations show that our proposed framework succeeds in both attacking state-of-the-art structured prediction models and boosting them with adversarial training. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP (1) | 1 |
| 2019 | Enhancing Unsupervised Generative Dependency Parser with Contextual InformationabstractMost of the unsupervised dependency parsers are based on probabilistic generative models that learn the joint distribution of the given sentence and its parse.Probabilistic generative models usually explicit decompose the desired dependency tree into factorized grammar rules, which lack the global features of the entire sentence.In this paper, we propose a novel probabilistic model called discriminative neural dependency model with valence (D-NDMV) that generates a sentence and its parse from a continuous latent representation, which encodes global contextual information of the generated sentence.We propose two approaches to model the latent representation: the first deterministically summarizes the representation from the sentence and the second probabilistically models the representation conditioned on the sentence.Our approach can be regarded as a new type of autoencoder model to unsupervised dependency parsing that combines the benefits of both generative and discriminative techniques.In particular, our approach breaks the context-free independence assumption in previous generative approaches and therefore becomes more expressive.Our extensive experimental results on seventeen datasets from various sources show that our approach achieves competitive accuracy compared with both generative and discriminative state-of-the-art unsupervised dependency parsers. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
ACL (1) | 1 |
| 2019 | Multilingual Grammar Induction with Continuous Language IdentificationabstractWenjuan Han, Ge Wang, Yong Jiang, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wenjuan Han, Ge Wang 0005, Yong Jiang 0005, Kewei Tu |
EMNLP/IJCNLP (1) | 1 |
| 2019 | A Regularization-based Framework for Bilingual Grammar InductionabstractYong Jiang, Wenjuan Han, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Lexicalized Neural Unsupervised Dependency Parsing
Wenjuan Han, Yong Jiang 0005, Kewei Tu |
Neurocomputing | 1 |
| 2017 | Dependency Grammar Induction with Neural Lexicalization and Big Training DataabstractWe study the impact of big models (in terms of the degree of lexicalization) and big data (in terms of the training corpus size) on dependency grammar induction.We experimented with L-DMV, a lexicalized version of Dependency Model with Valence (Klein and Manning, 2004) and L-NDMV, our lexicalized extension of the Neural Dependency Model with Valence (Jiang et al., 2016).We find that L-DMV only benefits from very small degrees of lexicalization and moderate sizes of training corpora.L-NDMV can benefit from big training data and lexicalization of greater degrees, especially when enhanced with good model initialization, and it achieves a result that is competitive with the current state-of-the-art. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP | 1 |
| 2017 | Combining Generative and Discriminative Approaches to Unsupervised Dependency Parsing via Dual DecompositionabstractUnsupervised dependency parsing aims to learn a dependency parser from unannotated sentences.Existing work focuses on either learning generative models using the expectation-maximization algorithm and its variants, or learning discriminative models using the discriminative clustering algorithm.In this paper, we propose a new learning strategy that learns a generative model and a discriminative model jointly based on the dual decomposition method.Our method is simple and general, yet effective to capture the advantages of both models and improve their learning results.We tested our method on the UD treebank and achieved a state-ofthe-art performance on thirty languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 2 |
| 2016 | Unsupervised Neural Dependency ParsingabstractUnsupervised dependency parsing aims to learn a dependency grammar from text annotated with only POS tags.Various features and inductive biases are often used to incorporate prior knowledge into learning.One useful type of prior information is that there exist correlations between the parameters of grammar rules involving different POS tags.Previous work employed manually designed features or special prior distributions to encode such information.In this paper, we propose a novel approach to unsupervised dependency parsing that uses a neural model to predict grammar rule probabilities based on distributed representation of POS tags.The distributed representation is automatically learned from data and captures the correlations between POS tags.Our experiments show that our approach outperforms previous approaches utilizing POS correlations and is competitive with recent state-of-the-art approaches on nine different languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 2 |