VLDB 2026 Research / reviewers in the wild / expert
Hongming Zhang 0009
dblp:48/859-9
· DBLP profile ↗
59ranked-venue papers
9as first author
46since 2021 · last 2026
0000-0001-8172-6484ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 8 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsabstractRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong |
ACL (1) | 10 |
| 2025 | OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and OptimizationabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 5 |
| 2025 | Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsabstractLarge language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms. Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001 |
ACL (1) | 5 |
| 2025 | WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World ModelabstractAgent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 . Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001 |
EMNLP | 2 |
| 2025 | Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and ExtrapolationabstractMamba's theoretical infinite-context potential is limited in practice when sequences far exceed training lengths.This work explores unlocking Mamba's long-context memory ability by a simple-yet-effective method, Recall with Reasoning (RwR), by distilling chain-ofthought (CoT) summarization from a teacher model.Specifically, RwR prepends these summarization as CoT prompts during fine-tuning, teaching Mamba to actively recall and reason over long contexts.Experiments on LONG-MEMEVAL and HELMET show that RwR outperforms existing long-term memory methods on the Mamba model.Furthermore, under similar pre-training conditions, RwR improves the long-context performance of Mamba relative to comparable Transformer/hybrid baselines while preserving short-context capabilities, all without changing the architecture. Jun-Yu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001 |
EMNLP | 4 |
| 2025 | Retrieval-augmented GUI Agents with Generative GuidelinesabstractGUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks.However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require longtailed knowledge covering rare, unseen scenarios.We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time.RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF).Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent.Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios. Ran Xu 0002, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Joyce C. Ho, Carl Yang 0001, Dong Yu 0001 |
EMNLP | 4 |
| 2025 | DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?abstractLarge Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents. Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001 |
ICLR | 7 |
| 2025 | RepoGraph: Enhancing AI Software Engineering with Repository-level Code GraphabstractLarge Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks.
Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph. Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001 |
ICLR | 8 |
| 2025 | UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionabstractLarge language models are increasingly capable of handling long-context inputs, but the memory overhead of KV cache remains a major bottleneck for general-purpose deployment. While many compression strategies have been explored, sequence-level compression is particularly challenging due to its tendency to lose important details. We present UniGist, a gist token-based long context compression framework that removes the need for chunk-wise training, enabling the model to learn how to compress and utilize long-range context during training. To fully exploit the sparsity, we introduce a gist shift trick that transforms the attention layout into a right-aligned block structure and develop a block-table-free sparse attention kernel based on it. UniGist further supports one-pass training and flexible chunk sizes during inference, allowing efficient and adaptive context processing. Experiments across multiple long-context tasks show that UniGist significantly improves compression quality, with especially strong performance in recalling details and long-range dependency modeling. Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Tianqing Fang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Zhicheng Dou |
NeurIPS | 6 |
| 2024 | WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 6 |
| 2024 | CLOMO: Counterfactual Logical Modification with Large Language ModelsabstractYinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao, Zhicheng Yang, Dong Yu, Changshui Zhang, Xiaodan Liang, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yinya Huang, Ruixin Hong, Hongming Zhang 0009, Wei Shao 0009, Dong Yu 0001, Changshui Zhang, Xiaodan Liang, Linqi Song |
ACL (1) | 3 |
| 2024 | AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility EstimationabstractZhaowei Wang, Wei Fan, Qing Zong, Hongming Zhang, Sehyun Choi, Tianqing Fang, Xin Liu, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaowei Wang 0003, Wei Fan 0001, Qing Zong, Hongming Zhang 0009, Sehyun Choi, Tianqing Fang, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See |
ACL (1) | 4 |
| 2024 | Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language ModelsabstractRetrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources.However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response generation.Moreover, standard RALMs frequently neglect their intrinsic knowledge due to the interference from retrieved information.In instances where the retrieved information is irrelevant, RALMs should ideally utilize their intrinsic knowledge or, in the absence of both intrinsic and retrieved knowledge, opt to respond with "unknown" to avoid hallucination.In this paper, we introduces CHAIN-OF-NOTE (CON), a novel approach to improve robustness of RALMs in facing noisy, irrelevant documents and in handling unknown scenarios.The core idea of CON is to generate sequential reading notes for each retrieved document, enabling a thorough evaluation of their relevance to the given question and integrating this information to formulate the final answer.Our experimental results show that GPT-4, when equipped with CON, outperforms the CHAIN-OF-THOUGHT approach.Besides, we utilized GPT-4 to create 10K CON data, subsequently trained on LLaMa-2 7B model.Our experiments across four open-domain QA benchmarks show that fine-tuned RALMs equipped with CON significantly outperform standard fine-tuned RALMs. Wenhao Yu 0002, Hongming Zhang 0009, Xiaoman Pan, Peixin Cao, Kaixin Ma, Hongwei Wang 0001, Dong Yu 0001 |
EMNLP | 2 |
| 2024 | MixGR: Enhancing Retriever Generalization for Scientific Domain through Complementary GranularityabstractRecent studies show the growing significance of document retrieval in the generation of LLMs, i.e., RAG, within the scientific domain by bridging their knowledge gap.However, dense retrievers often struggle with domainspecific retrieval and complex query-document relationships, particularly when query segments correspond to various parts of a document.To alleviate such prevalent challenges, this paper introduces MixGR, which improves dense retrievers' awareness of query-document matching across various levels of granularity in queries and documents using a zero-shot approach.MixGR fuses various metrics based on these granularities to a united score that reflects a comprehensive query-document similarity.Our experiments demonstrate that MixGR outperforms previous document retrieval by 24.7%, 9.8%, and 6.9% on nDCG@5 with unsupervised, supervised, and LLM-based retrievers, respectively, averaged on queries containing multiple subqueries from five scientific retrieval datasets.Moreover, the efficacy of two downstream scientific question-answering tasks highlights the advantage of MixGR to boost the application of LLMs in the scientific domain.The code and experimental datasets are available.1 Fengyu Cai, Hongming Zhang 0009, Iryna Gurevych, Heinz Koeppl |
EMNLP | 5 |
| 2024 | Dense X Retrieval: What Retrieval Granularity Should We Use?abstractDense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks.When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g.document, passage, or sentence.We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks.Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval.Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format.We conduct an empirical comparison of different retrieval granularity.Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks.Moreover, constructing prompts with fine-grained retrieved units for retrievalaugmented language models improves the performance of downstream QA tasks given a specific computation budget. Hongwei Wang 0010, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Dong Yu 0001 |
EMNLP | 7 |
| 2024 | Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic RepresentationsabstractSihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hongming Zhang 0009, Ben Zhou, Wenhao Yu 0002, Dian Yu 0001, Baolin Peng, Hongwei Wang 0010, Dan Roth 0001, Dong Yu 0001 |
NAACL-HLT | 2 |
| 2024 | A Closer Look at the Self-Verification Abilities of Large Language Models in Logical ReasoningabstractRuixin Hong, Hongming Zhang, Xinyu Pang, Dong Yu, Changshui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ruixin Hong, Hongming Zhang 0009, Xinyu Pang, Dong Yu 0001, Changshui Zhang |
NAACL-HLT | 2 |
| 2023 | COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference PerspectiveabstractZhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhaowei Wang 0003, Quyet V. Do, Hongming Zhang 0009, Jiayao Zhang 0001, Weiqi Wang 0001, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See |
ACL (1) | 3 |
| 2023 | Faithful Question Answering with Monte-Carlo PlanningabstractAlthough large language models demonstrate remarkable question-answering performances, revealing the intermediate reasoning steps that the models faithfully follow remains challenging.In this paper, we propose FAME (FAithful question answering with MontE-carlo planning) to answer questions based on faithful reasoning steps.The reasoning steps are organized as a structured entailment tree, which shows how premises are used to produce intermediate conclusions that can prove the correctness of the answer.We formulate the task as a discrete decision-making problem and solve it through the interaction of a reasoning environment and a controller.The environment is modular and contains several basic task-oriented modules, while the controller proposes actions to assemble the modules.Since the search space could be large, we introduce a Monte-Carlo planning algorithm to do a look-ahead search and select actions that will eventually lead to highquality steps.FAME achieves advanced performance on the standard benchmark.It can produce valid and faithful reasoning steps compared with large language models with a much smaller model size. Ruixin Hong, Hongming Zhang 0009, Dong Yu 0001, Changshui Zhang |
ACL (1) | 2 |
| 2023 | Extracting or Guessing? Improving Faithfulness of Event Temporal Relation ExtractionabstractIn this paper, we seek to improve the faithfulness of TEMPREL extraction models from two perspectives.The first perspective is to extract genuinely based on contextual description.To achieve this, we propose to conduct counterfactual analysis to attenuate the effects of two significant types of training biases: the event trigger bias and the frequent label bias.We also add tense information into event representations to explicitly place an emphasis on the contextual description.The second perspective is to provide proper uncertainty estimation and abstain from extraction when no relation is described in the text.By parameterization of Dirichlet Prior over the model-predicted categorical distribution, we improve the model estimates of the correctness likelihood and make TEMPREL predictions more selective.We also employ temperature scaling to recalibrate the model confidence measure after bias mitigation.Through experimental analysis on MATRES, MATRES-DS, and TDDiscourse, we demonstrate that our model extracts TEMPREL and timelines more faithfully compared to SOTA methods, especially under distribution shifts. Haoyu Wang 0005, Hongming Zhang 0009, Yuqian Deng, Jacob R. Gardner, Dan Roth 0001, Muhao Chen 0001 |
EACL | 2 |
| 2023 | How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation MethodabstractWenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
EACL | 3 |
| 2023 | Are All Steps Equally Important? Benchmarking Essentiality Detection in Event ProcessesabstractNatural language expresses events with varying granularities, where coarse-grained events (goals) can be broken down into finer-grained event sequences (steps).A critical yet overlooked aspect of understanding event processes is recognizing that not all step events hold equal importance toward the completion of a goal.In this paper, we address this gap by examining the extent to which current models comprehend the essentiality of step events in relation to a goal event.Cognitive studies suggest that such capability enables machines to emulate human commonsense reasoning about preconditions and necessary efforts of everyday tasks.We contribute a high-quality corpus of (goal, step) pairs gathered from the community guideline website WikiHow, with steps manually annotated for their essentiality concerning the goal by experts.The high inter-annotator agreement demonstrates that humans possess a consistent understanding of event essentiality.However, after evaluating multiple statistical and largescale pre-trained language models, we find that existing approaches considerably underperform compared to humans.This observation highlights the need for further exploration into this critical and challenging task 1 . Haoyu Wang 0005, Hongming Zhang 0009, Yueguan Wang, Yuqian Deng, Muhao Chen 0001, Dan Roth 0001 |
EMNLP | 2 |
| 2023 | StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingabstractCheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, Zheng Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Cheng Jiayang, Tsz Ho Chan, Tianqing Fang, Weiqi Wang 0001, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang 0009, Yangqiu Song, Yue Zhang 0004, Zheng Zhang 0001 |
EMNLP | 9 |
| 2023 | Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsabstractTraditional sentence embedding models encode sentences into vector representations to capture useful properties such as the semantic similarity between sentences.However, in addition to similarity, sentence semantics can also be interpreted via compositional operations such as sentence fusion or difference.It is unclear whether the compositional semantics of sentences can be directly reflected as compositional operations in the embedding space.To more effectively bridge the continuous embedding and discrete text spaces, we explore the plausibility of incorporating various compositional properties into the sentence embedding space that allows us to interpret embedding transformations as compositional sentence operations.We propose INTERSENT, an end-toend framework for learning interpretable sentence embeddings that supports compositional sentence operations in the embedding space.Our method optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddings.Experimental results demonstrate that our method significantly improves the interpretability of sentence embeddings on four textual generation tasks over existing approaches while maintaining strong performance on traditional semantic similarity tasks. 1 . James Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 0009, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 4 |
| 2023 | Video State-Changing Object SegmentationabstractDaily objects commonly experience state changes. For example, slicing a cucumber changes its state from whole to sliced. Learning about object state changes in Video Object Segmentation (VOS) is crucial for understanding and interacting with the visual world. Conventional VOS benchmarks do not consider this challenging yet crucial problem. This paper makes a pioneering effort to introduce a weakly-supervised benchmark on Video State-Changing Object Segmentation (VSCOS). We construct our VSCOS benchmark by selecting state-changing videos from existing datasets. In advocate of an annotation-efficient approach towards state-changing object segmentation, we only annotate the first and last frames of training videos, which is different from conventional VOS. Notably, an open-vocabulary setting is included to evaluate the generalization to novel types of objects or state changes. We empirically illustrate that state-of-the-art VOS models struggle with state-changing objects and lose track after the state changes. We analyze the main difficulties of our VSCOS task and identify three technical improvements, namely, fine-tuning strategies, representation learning, and integrating motion information. Applying these improvements results in a strong baseline for segmenting state-changing objects consistently. Our benchmark and baseline methods are publicly available at https://github.com/venom12138/VSCOS. Jiangwei Yu, Hongming Zhang 0009, Yu-Xiong Wang |
ICCV | 4 |
| 2023 | Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
Xiaoman Pan, Wenlin Yao, Hongming Zhang 0009, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
ICLR | 3 |
| 2023 | Thrust: Adaptively Propels Large Language Models with External KnowledgeabstractAlthough large-scale pre-trained language models (PTLMs) are shown to encode rich knowledge in their model parameters, the inherent knowledge in PTLMs can be opaque or static, making external knowledge necessary. However, the existing information retrieval techniques could be costly and may even introduce noisy and sometimes misleading knowledge. To address these challenges, we propose the instance-level adaptive propulsion of external knowledge (IAPEK), where we only conduct the retrieval when necessary. To achieve this goal, we propose to model whether a PTLM contains enough knowledge to solve an instance with a novel metric, Thrust, which leverages the representation distribution of a small amount of seen instances. Extensive experiments demonstrate that Thrust is a good measurement of models' instance-level knowledgeability. Moreover, we can achieve higher cost-efficiency with the Thrust score as the retrieval indicator than the naive usage of external knowledge on 88% of the evaluated tasks with 26% average performance improvement. Such findings shed light on the real-world practice of knowledge-enhanced LMs with a limited budget for knowledge seeking due to computation latency or costs. Hongming Zhang 0009, Xiaoman Pan, Wenlin Yao, Dong Yu 0001, Jianshu Chen |
NeurIPS | 2 |
| 2023 | OpenFact: Factuality Enhanced Open Knowledge ExtractionabstractAbstract We focus on the factuality property during the extraction of an OpenIE corpus named OpenFact, which contains more than 12 million high-quality knowledge triplets. We break down the factuality property into two important aspects—expressiveness and groundedness—and we propose a comprehensive framework to handle both aspects. To enhance expressiveness, we formulate each knowledge piece in OpenFact based on a semantic frame. We also design templates, extra constraints, and adopt human efforts so that most OpenFact triplets contain enough details. For groundedness, we require the main arguments of each triplet to contain linked Wikidata1 entities. A human evaluation suggests that the OpenFact triplets are much more accurate and contain denser information compared to OPIEC-Linked (Gashteovski et al., 2019), one recent high-quality OpenIE corpus grounded to Wikidata. Further experiments on knowledge base completion and knowledge base question answering show the effectiveness of OpenFact over OPIEC-Linked as supplementary knowledge to Wikidata as the major KG. Linfeng Song, Ante Wang, Xiaoman Pan, Hongming Zhang 0009, Dian Yu 0001, Lifeng Jin, Haitao Mi, Jinsong Su, Yue Zhang 0004, Dong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | Rare and Zero-shot Word Sense Disambiguation using Z-ReweightingabstractWord sense disambiguation (WSD) is a crucial problem in the natural language processing (NLP) community.Current methods achieve decent performance by utilizing supervised learning and large pre-trained language models.However, the imbalanced training dataset leads to poor performance on rare senses and zero-shot senses.There are more training instances and senses for words with top frequency ranks than those with low frequency ranks in the training dataset.We investigate the statistical relation between word frequency rank and word sense number distribution.Based on the relation, we propose a Z-reweighting method on the word level to adjust the training on the imbalanced dataset.The experiments show that the Z-reweighting strategy achieves performance gain on the standard English all words WSD benchmark.Moreover, the strategy can help models generalize better on rare and zero-shot senses. Hongming Zhang 0009, Yangqiu Song, Tong Zhang 0001 |
ACL (1) | 2 |
| 2022 | Multilingual Word Sense Disambiguation with Unified Sense RepresentationabstractAs a key natural language processing (NLP) task, word sense disambiguation (WSD) evaluates how well NLP models can understand the fine-grained semantics of words under specific contexts. Benefited from the large-scale annotation, current WSD systems have achieved impressive performances in English by combining supervised learning with lexical knowledge. However, such success is hard to be replicated in other languages, where we only have very limited annotations. In this paper, based on that the multilingual lexicon BabelNet describing the same set of concepts across languages, we propose to build knowledge and supervised based Multilingual Word Sense Disambiguation (MWSD) systems. We build unified sense representations for multiple languages and address the annotation scarcity problem for MWSD by transferring annotations from rich sourced languages. With the unified sense representations, annotations from multiple languages can be jointly trained to benefit the MWSD tasks. Evaluations of SemEval-13 and SemEval-15 datasets demonstrate the effectiveness of our methodology. Hongming Zhang 0009, Yangqiu Song, Tong Zhang 0001 |
COLING | 2 |
| 2022 | MetaLogic: Logical Reasoning Explanations with Fine-Grained StructureabstractIn this paper, we propose a comprehensive benchmark to investigate models' logical reasoning capabilities in complex real-life scenarios.Current explanation datasets often employ synthetic data with simple reasoning structures.Therefore, it cannot express more complex reasoning processes, such as the rebuttal to a reasoning step and the degree of certainty of the evidence.To this end, we propose a comprehensive logical reasoning explanation form.Based on the multi-hop chain of reasoning, the explanation form includes three main components:(1) The condition of rebuttal that the reasoning node can be challenged; (2) Logical formulae that uncover the internal texture of reasoning nodes; (3) Reasoning strength indicated by degrees of certainty.The fine-grained structure conforms to the real logical reasoning scenario, better fitting the human cognitive process but, simultaneously, is more challenging for the current models.We evaluate the current best models' performance on this new explanation form.The experimental results show that generating reasoning graphs remains a challenging task for current models, even with the help of giant pre-trained language models. Yinya Huang, Hongming Zhang 0009, Ruixin Hong, Xiaodan Liang, Changshui Zhang, Dong Yu 0001 |
EMNLP | 2 |
| 2022 | Salience Allocation as Guidance for Abstractive SummarizationabstractFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 3 |
| 2022 | SubeventWriter: Iterative Sub-event Sequence Generation with Coherence ControllerabstractIn this paper, we propose a new task of subevent generation for an unseen process to evaluate the understanding of the coherence of subevent actions and objects.To solve the problem, we design SubeventWriter, a sub-event sequence generation framework with a coherence controller.Given an unseen process, the framework can iteratively construct the subevent sequence by generating one sub-event at each iteration.We also design a very effective coherence controller to decode more coherent sub-events.As our extensive experiments and analysis indicate, SubeventWriter 1 can generate more reliable and meaningful sub-event sequences for unseen processes. Zhaowei Wang 0003, Hongming Zhang 0009, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See |
EMNLP | 2 |
| 2022 | Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationabstractLarge-scale pretrained language models have made significant advances in solving downstream language understanding tasks.However, they generally suffer from reporting bias, the phenomenon describing the lack of explicit commonsense knowledge in written text, e.g., "an orange is orange".To overcome this limitation, we develop a novel approach, Z-LaVI, to endow language models with visual imagination capabilities.Specifically, we leverage two complementary types of "imaginations": (i) recalling existing images through retrieval and (ii) synthesizing nonexistent images via text-toimage generation.Jointly exploiting the language inputs and the imagination, a pretrained vision-language model (e.g., CLIP) eventually composes a zero-shot solution to the original language tasks.Notably, fueling language models with imagination can effectively leverage visual knowledge to solve plain language tasks.In consequence, Z-LaVI consistently improves the zero-shot performance of existing language models across a diverse set of language tasks. 1 * Work was done during the internship at Tencent AI Lab.(a) Word Sense Disambiguation (b) Science Question Answering (c) Topic Classification sense1: bank(institute) sense2: bank(geography) Input: The species prefers the {bank} of pond. Wenlin Yao, Hongming Zhang 0009, Xiaoyang Wang 0001, Dong Yu 0001, Jianshu Chen |
EMNLP | 3 |
| 2022 | ROCK: Causal Inference Principles for Reasoning about Commonsense CausalityabstractCommonsense causality reasoning (CCR) aims at identifying plausible causes and effects in natural language descriptions that are deemed reasonable by an average person. Although being of great academic and practical interest, this problem is still shadowed by the lack of a well-posed theoretical framework; existing work usually relies on deep language models wholeheartedly, and is potentially susceptible to confounding co-occurrences. Motivated by classical causal principles, we articulate the central question of CCR and draw parallels between human subjects in observational studies and natural languages to adopt CCR to the potential-outcomes framework, which is the first such attempt for commonsense tasks. We propose a novel framework, ROCK, to Reason O(A)bout Commonsense K(C)ausality, which utilizes temporal signals as incidental supervision, and balances confounding effects using temporal propensities that are analogous to propensity scores. The ROCK implementation is modular and zero-shot, and demonstrates good CCR capabilities. Jiayao Zhang 0001, Hongming Zhang 0009, Weijie J. Su, Dan Roth 0001 |
ICML | 2 |
| 2022 | ARCLIN: Automated API Mention Resolution for Unformatted TextsabstractOnline technical forums (e.g., StackOverflow) are popular platforms for developers to discuss technical problems such as how to use a specific Application Programming Interface (API), how to solve the programming tasks, or how to fix bugs in their code. These discussions can often provide auxiliary knowledge of how to use the software that is not covered by the official documents. The automatic extraction of such knowledge may support a set of downstream tasks like API searching or indexing. However, unlike official documentation written by experts, discussions in open forums are made by regular developers who write in short and informal texts, including spelling errors or abbreviations. There are three major challenges for the accurate APIs recognition and linking mentioned APIs from unstructured natural language documents to an entry in the API repository: (1) distinguishing API mentions from common words; (2) identifying API mentions without a fully qualified name; and (3) disambiguating API mentions with similar method names but in a different library. In this paper, to tackle these challenges, we propose an ARCLIN tool, which can effectively distinguish and link APIs without using human annotations. Specifically, we first design an API recognizer to automatically extract API mentions from natural language sentences by a Conditional Random Field (CRF) on the top of a Bi-directional Long Short-Term Memory (Bi-LSTM) module, then we apply a context-aware scoring mechanism to compute the mention-entry similarity for each entry in an API repository. Compared to previous approaches with heuristic rules, our proposed tool without manual inspection outperforms by 8% in a high-quality dataset Py-mention, which contains 558 mentions and 2,830 sentences from five popular Python libraries. To our best knowledge, ARCLIN is the first approach to achieve full automation of API mention resolution from unformatted text without manually collected labels. Yintong Huo, Yuxin Su 0001, Hongming Zhang 0009, Michael R. Lyu |
ICSE | 3 |
| 2022 | PCR4ALL: A Comprehensive Evaluation Benchmark for Pronoun Coreference Resolution in EnglishabstractPronoun Coreference Resolution (PCR) is the task of resolving pronominal expressions to all mentions they refer to. The correct resolution of pronouns typically involves the complex inference over both linguistic knowledge and general world knowledge. Recently, with the help of pre-trained language representation models, the community has made significant progress on various PCR tasks. However, as most existing works focus on developing PCR models for specific datasets and measuring the accuracy or F1 alone, it is still unclear whether current PCR systems are reliable in real applications. Motivated by this, we propose PCR4ALL, a new benchmark and a toolbox that evaluates and analyzes the performance of PCR systems from different perspectives (i.e., knowledge source, domain, data size, frequency, relevance, and polarity). Experiments demonstrate notable performance differences when the models are examined from different angles. We hope that PCR4ALL can motivate the community to pay more attention to solving the overall PCR problem and understand the performance comprehensively. All data and codes are available at: https://github.com/HKUST-KnowComp/PCR4ALL. Hongming Zhang 0009, Yangqiu Song |
LREC | 2 |
| 2022 | ASER: Towards large-scale commonsense knowledge acquisition via higher-order selectional preference over eventualities
Hongming Zhang 0009, Xin Liu 0039, Haojie Pan, Haowen Ke, Jiefu Ou, Tianqing Fang, Yangqiu Song |
Artif. Intell. | 1 |
| 2022 | VD-PCR: Improving visual dialog with pronoun coreference resolution
Xintong Yu 0002, Hongming Zhang 0009, Ruixin Hong, Yangqiu Song, Changshui Zhang |
Pattern Recognit. | 2 |
| 2022 | Learning Event Extraction From a Few Guideline ExamplesabstractExisting fully supervised event extraction models achieve advanced performance with large-scale labeled data. However, when new event types emerge and annotations are scarce, it is hard for the supervised models to master the new types with limited annotations. In contrast, humans can learn to understand new event types with only a few examples in the event extraction guideline. In this paper, we work on a challenging yet more realistic setting, the few-example event extraction. It requires models to learn event extraction with only a few sentences in guidelines as training data, so that we do not need to collect large-scale annotations each time when new event types emerge. As models tend to overfit when trained with only a few examples, we propose knowledge-guided data augmentation to generate valid and diverse sentences from the guideline examples. To help models better leverage the augmented data, we add a consistency regularization to guarantee consistent representations between the augmented sentences and the original ones. Experiments on the standard benchmark ACE-2005 indicate that our method can extract event triggers and arguments effectively with only a few guideline examples. Ruixin Hong, Hongming Zhang 0009, Xintong Yu 0002, Changshui Zhang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Joint Coreference Resolution and Character Linking for Multiparty ConversationabstractCharacter linking, the task of linking mentioned people in conversations to the real world, is crucial for understanding the conversations.For the efficiency of communication, humans often choose to use pronouns (e.g., "she") or normal phrases (e.g., "that girl") rather than named entities (e.g., "Rachel") in the spoken language, which makes linking those mentions to real people a much more challenging than a regular entity linking task.To address this challenge, we propose to incorporate the richer context from the coreference relations among different mentions to help the linking.On the other hand, considering that finding coreference clusters itself is not a trivial task and could benefit from the global character information, we propose to jointly solve these two tasks.Specifically, we propose C 2 , the joint learning model of Coreference resolution and Character linking.The experimental results demonstrate that C 2 can significantly outperform previous works on both tasks.Further analyses are conducted to analyze the contribution of all modules in the proposed model and the effect of all hyper-parameters. Jiaxin Bai, Hongming Zhang 0009, Yangqiu Song, Kun Xu 0005 |
EACL | 2 |
| 2021 | Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaabstractThe Winograd Schema (WS) has been proposed as a test for measuring commonsense capabilities of models.Recently, pre-trained language model-based approaches have boosted performance on some WS benchmarks but the source of improvement is still not clear.This paper suggests that the apparent progress on WS may not necessarily reflect progress in commonsense reasoning.To support this claim, we first show that the current evaluation method of WS is sub-optimal and propose a modification that uses twin sentences for evaluation.We also propose two new baselines that indicate the existence of artifacts in WS benchmarks.We then develop a method for evaluating WS-like sentences in a zero-shot setting to account for the commonsense reasoning abilities acquired during the pretraining and observe that popular language models perform randomly in this setting when using our more strict evaluation.We conclude that the observed progress is mostly due to the use of supervision in training WS models, which is not likely to successfully support all the required commonsense reasoning skills and knowledge.1 Yanai Elazar, Hongming Zhang 0009, Yoav Goldberg, Dan Roth 0001 |
EMNLP (1) | 2 |
| 2021 | Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation DatasetabstractReasoning over commonsense knowledge bases (CSKBs) whose elements are in the form of free-text is an important yet hard task in NLP.While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions from external resources.In this task, CSKBs are grounded to a large-scale eventuality (activity, state, and event) graph to discriminate whether novel triples from the eventuality graph are plausible or not.However, existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation).In this paper, we benchmark the CSKB population task with a new large-scale dataset by first aligning four popular CSKBs, and then presenting a highquality human-annotated evaluation set to probe neural models' commonsense reasoning ability.We also propose a novel inductive commonsense reasoning model that reasons over graphs.Experimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task.Models achieving high accuracy during training perform poorly on the evaluation set, with a large gap between human performance. Tianqing Fang, Weiqi Wang 0001, Sehyun Choi, Shibo Hao, Hongming Zhang 0009, Yangqiu Song |
EMNLP (1) | 5 |
| 2021 | Learning Constraints and Descriptive Segmentation for Subevent DetectionabstractEvent mentions in text correspond to realworld events of varying degrees of granularity.The task of subevent detection aims to resolve this granularity issue, recognizing the membership of multi-granular events in event complexes.Since knowing the span of descriptive contexts of event complexes helps infer the membership of events, we propose the task of event-based text segmentation (EVENTSEG) as an auxiliary task to improve the learning for subevent detection.To bridge the two tasks together, we propose an approach to learning and enforcing constraints that capture dependencies between subevent detection and EVENTSEG prediction, as well as guiding the model to make globally consistent inference.Specifically, we adopt Rectifier Networks for constraint learning and then convert the learned constraints to a regularization term in the loss function of the neural model.Experimental results show that the proposed method outperforms baseline methods by 2.3% and 2.5% on benchmark datasets for subevent detection, HiEve and IC, respectively, while achieving a decent performance on EVENTSEG prediction 1 . Haoyu Wang 0005, Hongming Zhang 0009, Muhao Chen 0001, Dan Roth 0001 |
EMNLP (1) | 2 |
| 2021 | Exophoric Pronoun Resolution in Dialogues with Topic RegularizationabstractResolving pronouns to their referents has long been studied as a fundamental natural language understanding problem.Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions in text while ignoring the exophoric scenario.Exophoric pronouns are common in daily communications, where speakers may directly use pronouns to refer to some objects present in the environment without introducing the objects first.Although such objects are not mentioned in the dialogue text, they can often be disambiguated by the general topics of the dialogue.Motivated by this, we propose to jointly leverage the local context and global topics of dialogues to solve the out-of-text PCR problem.Extensive experiments demonstrate the effectiveness of adding topic regularization for resolving exophoric pronouns. Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Changshui Zhang, Kun Xu 0005, Dong Yu 0001 |
EMNLP (1) | 2 |
| 2021 | DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeabstractCommonsense knowledge is crucial for artificial intelligence systems to understand natural language. Previous commonsense knowledge acquisition approaches typically rely on human annotations (for example, ATOMIC) or text generation models (for example, COMET.) Human annotation could provide high-quality commonsense knowledge, yet its high cost often results in relatively small scale and low coverage. On the other hand, generation models have the potential to automatically generate more knowledge. Nonetheless, machine learning models often fit the training data well and thus struggle to generate high-quality novel knowledge. To address the limitations of previous approaches, in this paper, we propose an alternative commonsense knowledge acquisition framework DISCOS (from DIScourse to COmmonSense), which automatically populates expensive complex commonsense knowledge to more affordable linguistic knowledge resources. Experiments demonstrate that we can successfully convert discourse knowledge about eventualities from ASER, a large-scale discourse knowledge graph, into if-then commonsense knowledge defined in ATOMIC without any additional annotation effort. Further study suggests that DISCOS significantly outperforms previous supervised approaches in terms of novelty and diversity with comparable quality. In total, we can acquire 3.4M ATOMIC-like inferential commonsense knowledge by populating ATOMIC on the core part of ASER. Codes and data are available at https://github.com/HKUST-KnowComp/DISCOS-commonsense. Tianqing Fang, Hongming Zhang 0009, Weiqi Wang 0001, Yangqiu Song |
WWW | 2 |
| 2020 | WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema ChallengeabstractIn this paper, we present the first comprehensive categorization of essential commonsense knowledge for answering the Winograd Schema Challenge (WSC).For each of the questions, we invite annotators to first provide reasons for making correct decisions and then categorize them into six major knowledge categories.By doing so, we better understand the limitation of existing methods (i.e., what kind of knowledge cannot be effectively represented or inferred with existing methods) and shed some light on the commonsense knowledge that we need to acquire in the future for better commonsense reasoning.Moreover, to investigate whether current WSC models can understand the commonsense or they simply solve the WSC questions based on the statistical bias of the dataset, we leverage the collected reasons to develop a new task called WinoWhy, which requires models to distinguish plausible reasons from very similar but wrong reasons for all WSC questions.Experimental results prove that even though pre-trained language representation models have achieved promising progress on the original WSC dataset, they are still struggling at WinoWhy.Further experiments show that even though supervised models can achieve better performance, the performance of these models can be sensitive to the dataset distribution. Hongming Zhang 0009, Yangqiu Song |
ACL | 1 |
| 2020 | What Are You Trying to Do? Semantic Typing of Event ProcessesabstractThis paper studies a new cognitively motivated semantic typing task, multi-axis event process typing, that, given an event process, attempts to infer free-form type labels describing (i) the type of action made by the process and (ii) the type of object the process seeks to affect.This task is inspired by computational and cognitive studies of event understanding, which suggest that understanding processes of events is often directed by recognizing the goals, plans or intentions of the protagonist(s).We develop a large dataset containing over 60k event processes, featuring ultra fine-grained typing on both the action and object type axes with very large (10 3 ∼ 10 4 ) label vocabularies.We then propose a hybrid learning framework, P2GT, which addresses the challenging typing problem with indirect supervision from glosses 1 and a joint learning-to-rank framework.As our experiments indicate, P2GT supports identifying the intent of processes, as well as the fine semantic type of the affected object.It also demonstrates the capability of handling fewshot cases, and strong generalizability on outof-domain processes.2 * This work was done when the author was visiting the University of Pennsylvania.1 A gloss provides a sense definition for a lexeme. 2 The contributed learning resources, software and a system demonstration are available at http://cogcomp.org/page/publication_view/915. Muhao Chen 0001, Hongming Zhang 0009, Haoyu Wang 0005, Dan Roth 0001 |
CoNLL | 2 |
| 2020 | Joint Constrained Learning for Event-Event Relation ExtractionabstractUnderstanding natural language involves recognizing how multiple event mentions structurally and temporally interact with each other.In this process, one can induce event complexes that organize multi-granular events with temporal order and membership relations interweaving among them.Due to the lack of jointly labeled data for these relational phenomena and the restriction on the structures they articulate, we propose a joint constrained learning framework for modeling event-event relations.Specifically, the framework enforces logical constraints within and across multiple temporal and subevent relations by converting these constraints into differentiable learning objectives.We show that our joint constrained learning approach effectively compensates for the lack of jointly labeled data, and outperforms SOTA methods on benchmarks for both temporal relation extraction and event hierarchy construction, replacing a commonly used but more expensive global inference process.We also present a promising case study showing the effectiveness of our approach in inducing event complexes on an external corpus. 1 Haoyu Wang 0005, Muhao Chen 0001, Hongming Zhang 0009, Dan Roth 0001 |
EMNLP (1) | 3 |
| 2020 | When Hearst Is not Enough: Improving Hypernymy Detection from Corpus with Distributional ModelsabstractWe address hypernymy detection, i.e., whether an is-a relationship exists between words (x, y), with the help of large textual corpora.Most conventional approaches to this task have been categorized to be either pattern-based or distributional.Recent studies suggest that pattern-based ones are superior, if large-scale Hearst pairs are extracted and fed, with the sparsity of unseen (x, y) pairs relieved.However, they become invalid in some specific sparsity cases, where x or y is not involved in any pattern.For the first time, this paper quantifies the non-negligible existence of those specific cases.We also demonstrate that distributional methods are ideal to make up for patternbased ones in such cases.We devise a complementary framework, under which a patternbased and a distributional model collaborate seamlessly in cases which they each prefer.On several benchmark datasets, our framework achieves competitive improvements and the case study shows its better interpretability. Changlong Yu, Jialong Han, Peifeng Wang, Yangqiu Song, Hongming Zhang 0009, Wilfred Ng, Shuming Shi 0001 |
EMNLP (1) | 5 |
| 2020 | Analogous Process Structure Induction for Sub-event Sequence PredictionabstractComputational and cognitive studies of event understanding suggest that identifying, comprehending, and predicting events depend on having structured representations of a sequence of events and on conceptualizing (abstracting) its components into (soft) event categories.Thus, knowledge about a known process such as "buying a car" can be used in the context of a new but analogous process such as "buying a house".Nevertheless, most event understanding work in NLP is still at the ground level and does not consider abstraction.In this paper, we propose an Analogous Process Structure Induction (APSI) framework, which leverages analogies among processes and conceptualization of sub-event instances to predict the whole sub-event sequence of previously unseen open-domain processes.As our experiments and analysis indicate, APSI 1 supports the generation of meaningful sub-event sequences for unseen processes and can help predict missing events. Hongming Zhang 0009, Muhao Chen 0001, Haoyu Wang 0005, Yangqiu Song, Dan Roth 0001 |
EMNLP (1) | 1 |
| 2020 | TransOMCS: From Linguistic Graphs to Commonsense KnowledgeabstractCommonsense knowledge acquisition is a key problem for artificial intelligence. Conventional methods of acquiring commonsense knowledge generally require laborious and costly human annotations, which are not feasible on a large scale. In this paper, we explore a practical way of mining commonsense knowledge from linguistic graphs, with the goal of transferring cheap knowledge obtained with linguistic patterns into expensive commonsense knowledge. The result is a conversion of ASER [Zhang et al., 2020], a large-scale selectional preference knowledge resource, into TransOMCS, of the same representation as ConceptNet [Liu and Singh, 2004] but two orders of magnitude larger. Experimental results demonstrate the transferability of linguistic knowledge to commonsense knowledge and the effectiveness of the proposed approach in terms of quantity, novelty, and quality. TransOMCS is publicly available at: https://github.com/HKUST-KnowComp/TransOMCS. Hongming Zhang 0009, Daniel Khashabi, Yangqiu Song, Dan Roth 0001 |
IJCAI | 1 |
| 2020 | ASER: A Large-scale Eventuality Knowledge GraphabstractUnderstanding human’s language requires complex world knowledge. However, existing large-scale knowledge graphs mainly focus on knowledge about entities while ignoring knowledge about activities, states, or events, which are used to describe how entities or things act in the real world. To fill this gap, we develop ASER (activities, states, events, and their relations), a large-scale eventuality knowledge graph extracted from more than 11-billion-token unstructured textual data. ASER contains 15 relation types belonging to five categories, 194-million unique eventualities, and 64-million unique edges among them. Both intrinsic and extrinsic evaluations demonstrate the quality and effectiveness of ASER. Hongming Zhang 0009, Xin Liu 0039, Haojie Pan, Yangqiu Song, Cane Wing-ki Leung |
WWW | 1 |
| 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference AcquisitionabstractSelectional Preference (SP) is a commonly observed language phenomenon and proved to be useful in many natural language processing tasks.To provide a better evaluation method for SP models, we introduce SP-10K, a largescale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.Three representative SP acquisition methods based on pseudo-disambiguation are evaluated with SP-10K.To demonstrate the importance of our dataset, we investigate the relationship between SP-10K and the commonsense knowledge in ConceptNet5 and show the potential of using SP to represent the commonsense knowledge.We also use the Winograd Schema Challenge to prove that the proposed new SP relations are essential for the hard pronoun coreference resolution problem. Hongming Zhang 0009, Hantian Ding, Yangqiu Song |
ACL (1) | 1 |
| 2019 | Knowledge-aware Pronoun Coreference ResolutionabstractResolving pronoun coreference requires knowledge support, especially for particular domains (e.g., medicine).In this paper, we explore how to leverage different types of knowledge to better resolve pronoun coreference with a neural model.To ensure the generalization ability of our model, we directly incorporate knowledge in the format of triplets, which is the most common format of modern knowledge graphs, instead of encoding it with features or rules as that in conventional approaches.Moreover, since not all knowledge is helpful in certain contexts, to selectively use them, we propose a knowledge attention module, which learns to select and use informative knowledge based on contexts, to enhance our model.Experimental results on two datasets from different domains prove the validity and effectiveness of our model, where it outperforms state-of-the-art baselines by a large margin.Moreover, since our model learns to use external knowledge rather than only fitting the training data, it also demonstrates superior performance to baselines in the cross-domain setting. Hongming Zhang 0009, Yan Song 0003, Yangqiu Song, Dong Yu 0001 |
ACL (1) | 1 |
| 2019 | Multilingual and Multi-Aspect Hate Speech AnalysisabstractNedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, Dit-Yan Yeung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang 0009, Yangqiu Song, Dit-Yan Yeung |
EMNLP/IJCNLP (1) | 3 |
| 2019 | What You See is What You Get: Visual Pronoun Coreference Resolution in DialoguesabstractXintong Yu, Hongming Zhang, Yangqiu Song, Yan Song, Changshui Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Yan Song 0003, Changshui Zhang |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Multiplex Word Embeddings for Selectional Preference AcquisitionabstractHongming Zhang, Jiaxin Bai, Yan Song, Kun Xu, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hongming Zhang 0009, Jiaxin Bai, Yan Song 0003, Kun Xu 0005, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Scalable Multiplex Network EmbeddingabstractNetwork embedding has been proven to be helpful for many real-world problems. In this paper, we present a scalable multiplex network embedding model to represent information of multi-type relations into a unified embedding space. To combine information of different types of relations while maintaining their distinctive properties, for each node, we propose one high-dimensional common embedding and a lower-dimensional additional embedding for each type of relation. Then multiple relations can be learned jointly based on a unified network embedding model. We conduct experiments on two tasks: link prediction and node classification using six different multiplex networks. On both tasks, our model achieved better or comparable performance compared to current state-of-the-art models with less memory use. Hongming Zhang 0009, Liwei Qiu, Lingling Yi, Yangqiu Song |
IJCAI | 1 |