Rui Zhang 0037

dblp:60/2536-37 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
33since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
abstract
Priyanka Kargupta, Nan Zhang, Yunyi Zhang, Rui Zhang, Prasenjit Mitra, Jiawei Han. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Priyanka Kargupta, Yunyi Zhang 0001, Rui Zhang 0037, Prasenjit Mitra 0001, Jiawei Han 0001
ACL (1)4
2025 Accelerating Causal Network Discovery of Alzheimer's Disease Biomarkers via Scientific Literature-Based Retrieval Augmented Generation
Xiaofan Zhou, Liangjie Huang, Pinyang Chen, Wenpeng Yin 0001, Rui Zhang 0037, Wenrui Hao, Lu Cheng 0001
IEEE Big Data5
2025 HRScene: How Far are VLMs from Effective High-Resolution Image Understanding?
abstract
High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 $\times$ 1,024 to 35,503 $\times$ 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic to radiology images, street views, long-range pictures, and telescope images. It includes HRIs of real-world objects, scanned documents, and composite multi-image. The two diagnostic evaluation datasets are synthesized by combining the target image with the gold answer and distracting images in different orders, assessing how well models utilize regions in HRI. We conduct extensive experiments involving 28 VLMs, including Gemini 2.0 Flash and GPT-4o. Experiments on HRScene show that current VLMs achieve an average accuracy of around 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on synthetic datasets reveal that VLMs struggle to effectively utilize HRI regions, showing significant Regional Divergence and lost-in-middle, shedding light on future research.
Yusen Zhang 0001, Wenliang Zheng, Aashrith Madasu, Peng Shi 0010, Ryo Kamoi, Zhuoyang Zou, Shu Zhao 0006, Sarkar Snigdha Sarathi Das, Xiaoxin Lu, Ranran Haoran Zhang, Avitej Iyer, Renze Lou, Wenpeng Yin 0001, Rui Zhang 0037
ICCV17
2025 GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
abstract
The effectiveness of large language models (LLMs) is closely tied to the design of prompts, making prompt optimization essential for enhancing their performance across a wide range of tasks. Although recent advancements have focused on automating prompt engineering, many existing approaches rely exclusively on textual feedback, refining prompts based solely on inference errors identified by large, computationally expensive LLMs. Unfortunately, smaller models struggle to generate high-quality feedback, resulting in complete dependence on large LLM judgment. Moreover, these methods fail to leverage more direct and finer-grained information, such as gradients, due to operating purely in text space. To this end, we introduce, we introduce *GReaTer*, a novel prompt optimization technique that directly incorporates *gradient information over task-specific reasoning*. By utilizing task loss gradients, *GReaTer* enables self-optimization of prompts for smaller, lightweight language models (LM) without the need for costly closed-source LLMs, while maintaining reasonable prompt structures. This allows high-performance prompt optimization without dependence on massive LLMs, closing the gap between smaller models and the sophisticated reasoning often needed for prompt refinement. Extensive evaluations across diverse tasks demonstrate that \ours consistently outperforms previous methods, even those reliant on powerful LLMs. Additionally, *GReaTer*-optimized prompts frequently exhibit better transferability and, in some cases, boost task performance to levels comparable to or surpassing those achieved by larger language models, highlighting the effectiveness of *"gradient over reasoning"*-based prompt optimization. Code of *GReaTer* is available at: https://github.com/psunlpgroup/GreaTer
Sarkar Snigdha Sarathi Das, Ryo Kamoi, Bo Pang 0004, Yusen Zhang 0001, Caiming Xiong, Rui Zhang 0037
ICLR6
2025 SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
abstract
Indexing is an important step towards strong performance in retrieval-augmented generation (RAG) systems. However, existing methods organize data based on either semantic similarity (similarity) or related information (relatedness), but do not cover both perspectives comprehensively. Our analysis reveals that modeling only one perspective results in insufficient knowledge synthesis, leading to suboptimal performance on complex tasks requiring multihop reasoning. In this paper, we propose SiReRAG, a novel RAG indexing approach that explicitly considers both similar and related information. On the similarity side, we follow existing work and explore some variances to construct a similarity tree based on recursive summarization. On the relatedness side, SiReRAG extracts propositions and entities from texts, groups propositions via shared entities, and generates recursive summaries to construct a relatedness tree. We index and flatten both similarity and relatedness trees into a unified retrieval pool. Our experiments demonstrate that SiReRAG consistently outperforms state-of-the-art indexing methods on three multihop datasets (MuSiQue, 2WikiMultiHopQA, and HotpotQA), with an average 1.9% improvement in F1 scores. As a reasonably efficient solution, SiReRAG enhances existing reranking methods significantly, with up to 7.8% improvement in average F1 scores. Our code is available at https://github.com/SalesforceAIResearch/SiReRAG.
Prafulla Kumar Choubey, Alexander R. Fabbri, Gabriel Bernadett-Shapiro, Rui Zhang 0037, Prasenjit Mitra 0001, Caiming Xiong, Chien-Sheng Wu
ICLR5
2025 Coverage-based Fairness in Multi-document Summarization
abstract
Haoyuan Li, Yusen Zhang, Rui Zhang, Snigdha Chaturvedi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yusen Zhang 0001, Rui Zhang 0037, Snigdha Chaturvedi
NAACL (Long Papers)3
2024 DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial Documents
abstract
Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yilun Zhao 0001, Yitao Long, Hongjun Liu 0001, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu 0003, Xiangru Tang, Rui Zhang 0037, Arman Cohan
ACL (1)9
2024 KnowledgeFMath: A Knowledge-Intensive Math Reasoning Dataset in Finance Domains
abstract
We introduce FinanceMATH, a novel benchmark designed to evaluate LLMs' capabilities in solving knowledge-intensive math reasoning problems.Compared to prior works, this study features three core advancements.First, FinanceMATH includes 1,200 problems with a hybrid of textual and tabular content.These problems require college-level knowledge in the finance domain for effective resolution.Second, we provide expert-annotated, detailed solution references in Python program format, ensuring a high-quality benchmark for LLM assessment.We also construct a finance-domain knowledge bank and investigate various knowledge integration strategies.Finally, we evaluate a wide spectrum of 51 LLMs with both Chainof-Thought and Program-of-Thought prompting methods.Our experimental results reveal that the current best-performing system (i.e., GPT-4o) achieves only 60.9% accuracy using CoT prompting, leaving substantial room for improvement.Moreover, while augmenting LLMs with external knowledge can improve model performance (e.g., 47.5% → 54.5% for Gemini-1.5-Pro),their accuracy remains significantly lower than the estimated human expert performance of 92%.We believe that Fi-nanceMATH can advance future research in the area of domain-specific knowledge retrieval and integration, particularly within the context of solving reasoning-intensive tasks. * Equal ContributionQuestion: In 2018, Company A had a passive equity ownership interest of 15% in Company B. By the close of 2018, Company A decided to increase its ownership in Company B to 50%, effective as of 1st January 2019, through a cash purchase.There have been no financial transactions between Company A and Company B. Based on the data in the following table with
Yilun Zhao 0001, Hongjun Liu 0001, Yitao Long, Rui Zhang 0037, Chen Zhao 0013, Arman Cohan
ACL (1)4
2024 LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
abstract
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001
EMNLP38
2024 FOLIO: Natural Language Reasoning with First-Order Logic
abstract
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev
EMNLP24
2024 MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
abstract
Traditionally, Machine Translation (MT) Evaluation has been treated as a regression problem -- producing an absolute translation-quality score. This approach has two limitations: i) the scores lack interpretability, and human annotators struggle with giving consistent scores; ii) most scoring methods are based on (reference, translation) pairs, limiting their applicability in real-world scenarios where references are absent. In practice, we often care about whether a new MT system is better or worse than some competitors. In addition, reference-free MT evaluation is increasingly practical and necessary. Unfortunately, these two practical considerations have yet to be jointly explored. In this work, we formulate the reference-free MT evaluation into a pairwise ranking problem. Given the source sentence and a pair of translations, our system predicts which translation is better. In addition to proposing this new formulation, we further show that this new paradigm can demonstrate superior correlation with human judgments by merely using indirect supervision from natural language inference and weak supervision from our synthetic data. In the context of reference-free evaluation, MT-Ranker, trained without any human annotations, achieves state-of-the-art results on the WMT Shared Metrics Task benchmarks DARR20, MQM20, and MQM21. On a more challenging benchmark, ACES, which contains fine-grained evaluation criteria such as addition, omission, and mistranslation errors, MT-Ranker marks state-of-the-art against reference-free as well as reference-based baselines.
Ibraheem Muhammad Moosa, Rui Zhang 0037, Wenpeng Yin 0001
ICLR2
2024 Veiled Pathways: Investigating Covert and Side Channels Within GPU Uncore
abstract
With the emergence of GPUs as first-class compute engines, more concentrated focus has been put into covert and side channel discovery in these architectures. However, most of the covert and side channels uncovered on GPUs to date are rooted in “GPU cores”, which include computational cores, cache and core interconnects, but they do not consider “GPU uncore”, which include non-computational engines, GPU DRAM, host-G PU links and inter-GPulinks. In this paper, we delve into the less-explored domains of GPU uncore, unveiling four novel leakage sources for covert and side channel exploitation: (1) GPU DRAM frequency scaling; (2) NVENC utilization; (3) NVDEC utilization; (4) NVJPEG utilization. What makes these covert and side channels interesting is that they all take effect under the GPU MPS mode - which fractionalizes GPU cores and GPU memory on both desktop-scale and server-scale GPUs. Furthermore, our study reevaluates PCI-e bandwidth allocation on GPUs. Notably, we have engineered covert and side channel capable of bypassing GPU MIG isolation - a mechanism implemented by NVIDIA to physically segregate hardware resources on server-scale GPUs. Our research showcases concrete examples of these covert and side channels, highlighting their potency in breaching system security, all achieved without necessitating root privileges. This underscores the practical implications and urgency of addressing these vulnerabilities in GPU architectures.
Yuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang, Gang Tan, Rui Zhang 0037, Mahmut T. Kandemir
MICRO6
2024 Fair Abstractive Summarization of Diverse Perspectives
abstract
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yusen Zhang 0001, Yixin Liu 0003, Alexander R. Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao 0001, Dragomir R. Radev, Kathy McKeown, Rui Zhang 0037
NAACL-HLT12
2024 Chain of Agents: Large Language Models Collaborating on Long-Context Tasks
abstract
Addressing the challenge of effectively processing long contexts has become a critical issue for Large Language Models (LLMs). Two common strategies have emerged: 1) reducing the input length, such as retrieving relevant chunks by Retrieval-Augmented Generation (RAG), and 2) expanding the context window limit of LLMs. However, both strategies have drawbacks: input reduction has no guarantee of covering the part with needed information, while window extension struggles with focusing on the pertinent information for solving the task. To mitigate these limitations, we propose Chain-of-Agents (CoA), a novel framework that harnesses multi-agent collaboration through natural language to enable information aggregation and context reasoning across various LLMs over long-context tasks. CoA consists of multiple worker agents who sequentially communicate to handle different segmented portions of the text, followed by a manager agent who synthesizes these contributions into a coherent final output. CoA processes the entire input by interleaving reading and reasoning, and it mitigates long context focus issues by assigning each agent a short context. We perform a comprehensive evaluation of CoA on a wide range of long-context tasks in question answering, summarization, and code completion, demonstrating significant improvements by up to 10% over strong baselines of RAG, Full-Context, and multi-agent LLMs.
Yusen Zhang 0001, Ruoxi Sun 0002, Yanfei Chen, Tomas Pfister, Rui Zhang 0037, Sercan Ö. Arik
NeurIPS5
2024 Hermes: Unlocking Security Analysis of Cellular Network Protocols by Synthesizing Finite State Machines from Natural Language Specifications
Abdullah Al Ishtiaq, Sarkar Snigdha Sarathi Das, Syed Md. Mukit Rashid, Ali Ranjbar, Kai Tu, Tianwei Wu, Zhezheng Song, Mujtahid Akon, Rui Zhang 0037, Syed Rafiul Hussain
USENIX Security Symposium10
2024 When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
abstract
Abstract Self-correction is an approach to improving responses from large language models (LLMs) by refining the responses using LLMs during inference. Prior work has proposed various self-correction frameworks using different sources of feedback, including self-evaluation and external feedback. However, there is still no consensus on the question of when LLMs can correct their own mistakes, as recent studies also report negative results. In this work, we critically survey broad papers and discuss the conditions required for successful self-correction. We first find that prior studies often do not define their research questions in detail and involve impractical frameworks or unfair evaluations that over-evaluate self-correction. To tackle these issues, we categorize research questions in self-correction research and provide a checklist for designing appropriate experiments. Our critical survey based on the newly categorized research questions shows that (1) no prior work demonstrates successful self-correction with feedback from prompted LLMs, except for studies in tasks that are exceptionally suited for self-correction, (2) self-correction works well in tasks that can use reliable external feedback, and (3) large-scale fine-tuning enables self-correction.
Ryo Kamoi, Yusen Zhang 0001, Jiawei Han 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics5
2023 XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning Representations
abstract
Cross-Lingual Semantic Parsing (CLSP) aims to translate queries in multiple natural languages (NLs) into meaning representations (MRs) such as SQL, lambda calculus, and logic forms.However, existing CLSP models are separately proposed and evaluated on datasets of limited tasks and applications, impeding a comprehensive and unified evaluation of CLSP on a diverse range of NLs and MRs.To this end, we present XSEMPLR, a unified benchmark for cross-lingual semantic parsing featured with 22 natural languages and 8 meaning representations by examining and selecting 9 existing datasets to cover 5 tasks and 164 domains.We use XSEMPLR to conduct a comprehensive benchmark study on a wide range of multilingual language models including encoder-based models (mBERT, XLM-R), encoder-decoder models (mBART, mT5), and decoder-based models (Codex, BLOOM).We design 6 experiment settings covering various lingual combinations (monolingual, multilingual, cross-lingual) and numbers of learning samples (full dataset, few-shot, and zero-shot).Our experiments show that encoder-decoder models (mT5) achieve the highest performance compared with other popular models, and multilingual training can further improve the average performance.Notably, multilingual large language models (e.g., BLOOM) are still inadequate to perform CLSP tasks.We also find that the performance gap between monolingual training and cross-lingual transfer learning is still significant for multilingual models, though it can be mitigated by cross-lingual fewshot training.Our dataset and code are available at https://github.com/psunlpgroup/ XSemPLR.
Yusen Zhang 0001, Jun Wang 0122, Zhiguo Wang 0006, Rui Zhang 0037
ACL (1)4
2023 ConEntail: An Entailment-based Framework for Universal Zero and Few Shot Classification with Supervised Contrastive Pretraining
abstract
A universal classification model aims to generalize to diverse classification tasks in both zero and few shot settings.A promising way toward universal classification is to cast heterogeneous data formats into a dataset-agnostic "meta-task" (e.g., textual entailment, question answering) then pretrain a model on the combined meta dataset.The existing work is either pretrained on specific subsets of classification tasks, or pretrained on both classification and generation data but the model could not fulfill its potential in universality and reliability.These also leave a massive amount of annotated data under-exploited.To fill these gaps, we propose CONENTAIL, a new framework for universal zero and few shot classification with supervised contrastive pretraining.Our unified meta-task for classification is based on nested entailment.It can be interpreted as "Does sentence a entails [sentence b entails label c]".This formulation enables us to make better use of 57 annotated classification datasets for supervised contrastive pretraining and universal evaluation.In this way, CONENTAIL helps the model (1) absorb knowledge from different datasets, and (2) gain consistent performance gain with more pretraining data.In experiments, we compare our model with discriminative and generative models pretrained on the same dataset.The results confirm that our framework effectively exploits existing annotated data and outperforms baselines in both zero (9.4% average improvement) and few shot settings (3.5% average improvement).Our code is available at https: //github.com/psunlpgroup/ConEntail.
Ranran Haoran Zhang, Aysa Xuemo Fan, Rui Zhang 0037
EACL3
2023 Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse Finetuning
abstract
Unified Sequence Labeling articulates different sequence labeling tasks such as Named Entity Recognition, Relation Extraction, Semantic Role Labeling, etc. in a generalized sequence-to-sequence format.Unfortunately, this requires formatting different tasks into specialized augmented formats which are unfamiliar to the base pretrained language model (PLMs).This necessitates model fine-tuning and significantly bounds its usefulness in datalimited settings where fine-tuning large models cannot properly generalize to the target format.To address this challenge and leverage PLM knowledge effectively, we propose FISH-DIP, a sample-aware dynamic sparse finetuning strategy.It selectively finetunes a fraction of parameters informed by highly regressing examples during the fine-tuning process.By leveraging the dynamism of sparsity, our approach mitigates the impact of well-learned samples and prioritizes underperforming instances for improvement in generalization.Across five tasks of sequence labeling, we demonstrate that FISH-DIP can smoothly optimize the model in low-resource settings, offering up to 40% performance improvements over full fine-tuning depending on target evaluation settings.Also, compared to in-context learning and other parameter-efficient fine-tuning (PEFT) approaches, FISH-DIP performs comparably or better, notably in extreme lowresource settings.
Sarkar Snigdha Sarathi Das, Ranran Haoran Zhang, Peng Shi 0010, Wenpeng Yin 0001, Rui Zhang 0037
EMNLP5
2023 FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization
abstract
Summaries of medical text shall be faithful by being consistent and factual with source inputs, which is an important but understudied topic for safety and efficiency in healthcare.In this paper, we investigate and improve faithfulness in summarization on a broad range of medical summarization tasks.Our investigation reveals that current summarization models often produce unfaithful outputs for medical input text.We then introduce FAMESUMM, a framework to improve faithfulness by fine-tuning pre-trained language models based on medical knowledge.FAMESUMM performs contrastive learning on designed sets of faithful and unfaithful summaries, and it incorporates medical terms and their contexts to encourage faithful generation of medical terms.We conduct comprehensive experiments on three datasets in two languages: health question and radiology report summarization datasets in English, and a patient-doctor dialogue dataset in Chinese.Results demonstrate that FAMESUMM is flexible and effective by delivering consistent improvements over mainstream language models such as BART, T5, mT5, and PEGASUS, yielding state-of-the-art performances on metrics for faithfulness and general quality.Human evaluation by doctors also shows that FAMESUMM generates more faithful outputs.
Yusen Zhang 0001, Wu Guo, Prasenjit Mitra 0001, Rui Zhang 0037
EMNLP5
2023 Selective Annotation Makes Language Models Better Few-Shot Learners
Hongjin Su, Jungo Kasai, Chen Henry Wu, Jiayi Xin, Rui Zhang 0037, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009
ICLR7
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics9
2022 SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents
abstract
Yusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Awadallah, Dragomir Radev, Rui Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yusen Zhang 0001, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu 0001, Budhaditya Deb, Ahmed Awadallah 0001, Dragomir R. Radev, Rui Zhang 0037
ACL (1)9
2022 CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning
abstract
Named Entity Recognition (NER) in Few-Shot setting is imperative for entity tagging in low resource domains.Existing approaches only learn class-specific semantic features and intermediate representations from source domains.This affects generalizability to unseen target domains, resulting in suboptimal performances.To this end, we present CONTAINER, a novel contrastive learning technique that optimizes the inter-token distribution distance for Few-Shot NER.Instead of optimizing class-specific attributes, CONTAINER optimizes a generalized objective of differentiating between token categories based on their Gaussian-distributed embeddings.This effectively alleviates overfitting issues originating from training domains.Our experiments in several traditional test domains (OntoNotes, CoNLL'03, WNUT '17, GUM) and a new large scale Few-Shot NER dataset (Few-NERD) demonstrate that, on average, CONTAINER outperforms previous methods by 3%-13% absolute F1 points while showing consistent performance trends, even in challenging scenarios where previous approaches could not achieve appreciable performance.The source code of CONTAINER will be available at: https://github.com/ psunlpgroup/CONTaiNER.
Sarkar Snigdha Sarathi Das, Arzoo Katiyar, Rebecca J. Passonneau, Rui Zhang 0037
ACL (1)4
2022 DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization
abstract
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Awadallah, Dragomir Radev. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 0001, Rui Zhang 0037, Tao Yu 0009, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev
ACL (1)5
2022 MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data
abstract
Numerical reasoning over hybrid data containing both textual and tabular content (e.g., financial reports) has recently attracted much attention in the NLP community.However, existing question answering (QA) benchmarks over hybrid data only include a single flat table in each document and thus lack examples of multistep numerical reasoning across multiple hierarchical tables.To facilitate data analytical progress, we construct a new large-scale benchmark, MULTIHIERTT, with QA pairs over Multi Hierarchical Tabular and Textual data.MULTIHIERTT is built from a wealth of financial reports and has the following unique characteristics: 1) each document contain multiple tables and longer unstructured texts; 2) most of tables contained are hierarchical; 3) the reasoning process required for each question is more complex and challenging than existing benchmarks; and 4) fine-grained annotations of reasoning processes and supporting facts are provided to reveal complex numerical reasoning.We further introduce a novel QA model termed MT2Net, which first applies facts retrieving to extract relevant supporting facts from both tables and text and then uses a reasoning module to perform symbolic reasoning over retrieved facts.We conduct comprehensive experiments on various baselines.The experimental results show that MULTIHIERTT presents a strong challenge for existing baselines whose results lag far behind the performance of human experts.The dataset and code are publicly available at https://github. com/psunlpgroup/MultiHiertt.How much is the sum of stock purchase rights in 2018 lower than those in 2017?How many years were the sales and client service expenses higher than software development expenses?How much of US corporate debt securities is there in total (in 2009) without consider gross unrealized gain and gross unrelized loss?How many financing activities continues to increase every year from 2017 to 2021?
Yilun Zhao 0001, Chenying Li, Rui Zhang 0037
ACL (1)4
2022 UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
abstract
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009
EMNLP20
2022 ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning Examples
abstract
Reasoning over tabular data requires both table structure understanding and a broad set of table reasoning skills.Current models with tablespecific architectures and pre-training methods perform well on understanding table structures, but they still struggle with tasks that require various table reasoning skills.In this work, we develop REASTAP to show that high-level table reasoning skills can be injected into models during pre-training without a complex tablespecific architecture design.We define 7 table reasoning skills, such as numerical operation, temporal comparison, and conjunction.Each reasoning skill is associated with one example generator, which synthesizes questions over semi-structured tables according to the sampled templates.We model the table pre-training task as a sequence generation task and pretrain REASTAP to generate precise answers to the synthetic examples.REASTAP is evaluated on four benchmarks covering three downstream tasks including: 1) WIKISQL-WEAK and WIKITQ for Table Question Answering; 2) TABFACT for Table Fact Verification; and 3) LOGICNLG for Faithful Table-to-Text Generation.Experimental results demonstrate that REASTAP achieves new state-of-the-art performance on all benchmarks and delivers a significant improvement on low-resource setting.
Yilun Zhao 0001, Linyong Nan, Zhenting Qi, Rui Zhang 0037, Dragomir R. Radev
EMNLP4
2022 FeTaQA: Free-form Table Question Answering
abstract
Abstract Existing table question answering datasets contain abundant factual questions that primarily evaluate a QA system’s comprehension of query and tabular data. However, restricted by their short-form answers, these datasets fail to include question–answer interactions that represent more advanced and naturally occurring information needs: questions that ask for reasoning and integration of information pieces retrieved from a structured knowledge source. To complement the existing datasets and to reveal the challenging nature of the table-based question answering task, we introduce FeTaQA, a new dataset with 10K Wikipedia-based {table, question, free-form answer, supporting table cells} pairs. FeTaQA is collected from noteworthy descriptions of Wikipedia tables that contain information people tend to seek; generation of these descriptions requires advanced processing that humans perform on a daily basis: Understand the question and table, retrieve, integrate, infer, and conduct text planning and surface realization to generate an answer. We provide two benchmark methods for the proposed task: a pipeline method based on semantic parsing-based QA systems and an end-to-end method based on large pretrained text generation models, and show that FeTaQA poses a challenge for both methods.
Linyong Nan, Chiachun Hsieh, Ziming Mao, Xi Victoria Lin, Neha Verma 0001, Rui Zhang 0037, Wojciech Kryscinski, Hailey Schoelkopf, Riley Kong, Xiangru Tang, Mutethia Mutuma, Ben Rosand, Isabel Trindade, Renusree Bandaru, Jacob Cunningham, Caiming Xiong, Dragomir R. Radev
Trans. Assoc. Comput. Linguistics6
2021 Cross-language Sentence Selection via Data Augmentation and Rationale Training
abstract
Yanda Chen, Chris Kedzie, Suraj Nair, Petra Galuscakova, Rui Zhang, Douglas Oard, Kathleen McKeown. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanda Chen, Chris Kedzie, Suraj Nair 0001, Petra Galuscáková, Rui Zhang 0037, Douglas W. Oard, Kathy McKeown
ACL/IJCNLP (1)5
2021 SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Tao Yu 0009, Rui Zhang 0037, Oleksandr Polozov, Christopher Meek, Ahmed Awadallah 0001
ICLR2
2021 DART: Open-Domain Structured Data Record to Text Generation
abstract
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Linyong Nan, Dragomir R. Radev, Rui Zhang 0037, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma 0001, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta 0015, Tao Yu 0009, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani
NAACL-HLT3
2021 Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training
abstract
Recurrent Neural Networks (RNNs), more specifically their Long Short-Term Memory (LSTM) variants, have been widely used as a deep learning tool for tackling sequence-based learning tasks in text and speech. Training of such LSTM applications is computationally intensive due to the recurrent nature of hidden state computation that repeats for each time step. While sparsity in Deep Neural Nets has been widely seen as an opportunity for reducing computation time in both training and inference phases, the usage of non-ReLU activation in LSTM RNNs renders the opportunities for such dynamic sparsity associated with neuron activation and gradient values to be limited or non-existent. In this work, we identify dropout induced sparsity for LSTMs as a suitable mode of computation reduction. Dropout is a widely used regularization mechanism, which randomly drops computed neuron values during each iteration of training. We propose to structure dropout patterns, by dropping out the same set of physical neurons within a batch, resulting in column (row) level hidden state sparsity, which are well amenable to computation reduction at run-time in general-purpose SIMD hardware as well as systolic arrays. We provide a detailed analysis of how the dropout-induced sparsity propagates through the different stages of network training and how it can be leveraged in each stage. More importantly, our proposed approach works as a direct replacement for existing dropout-based application settings. We conduct our experiments for three representative NLP tasks: language modelling on the PTB dataset, OpenNMT based machine translation using the IWSLT De-En and En-Vi datasets, and named entity recognition sequence labelling using the CoNLL-2003 shared task. We demonstrate that our proposed approach can be used to translate dropout-based computation reduction into reduced training time, with improvement ranging from 1.23$\times$ to 1.64$\times$, without sacrificing the target metric.
Anup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang 0037, Mahmut T. Kandemir, Chita R. Das
NeurIPS4
2020 ESPRIT: Explaining Solutions to Physical Reasoning Tasks
abstract
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Nazneen Fatema Rajani, Rui Zhang 0037, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir R. Radev
ACL2
2019 ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks
abstract
Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article’s impacts on research community. This paper provides novel solutions to these two challenges. We 1) develop and release the first large-scale manually-annotated corpus for scientific papers (on computational linguistics) by enabling faster annotation, and 2) propose summarization methods that integrate the authors’ original highlights (abstract) and the article’s actual impacts on the community (citations), to create comprehensive, hybrid summaries. We conduct experiments to demonstrate the efficacy of our corpus in training data-driven models for scientific paper summarization and the advantage of our hybrid summaries over abstracts and traditional citation-based summaries. Our large annotated corpus and hybrid methods provide a new framework for scientific paper summarization research.
Michihiro Yasunaga, Jungo Kasai, Rui Zhang 0037, Alexander R. Fabbri, Irene Li, Dan Friedman, Dragomir R. Radev
AAAI3
2019 SParC: Cross-Domain Semantic Parsing in Context
abstract
Tao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev
ACL (1)2
2019 This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation
abstract
Given the overwhelming number of emails, an effective subject line becomes essential to better inform the recipient of the email's content.In this paper, we propose and study the task of email subject line generation: automatically generating an email subject line from the email body.We create the first dataset for this task and find that email subject line generation favor extremely abstractive summary which differentiates it from news headline generation or news single document summarization.We then develop a novel deep learning method and compare it to several baselines as well as recent state-of-the-art text summarization systems.We also investigate the efficacy of several automatic metrics based on correlations with human judgments and propose a new automatic evaluation metric.Our system outperforms competitive baselines given both automatic and human evaluations.To our knowledge, this is the first work to tackle the problem of effective email subject line generation.
Rui Zhang 0037, Joel R. Tetreault
ACL (1)1
2019 Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations
abstract
Rui Zhang, Caitlin Westerfield, Sungrok Shim, Garrett Bingham, Alexander Fabbri, William Hu, Neha Verma, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Rui Zhang 0037, Caitlin Westerfield, Sungrok Shim, Garrett Bingham, Alexander R. Fabbri, William Hu, Neha Verma 0001, Dragomir R. Radev
ACL (1)1
2019 CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
abstract
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev
EMNLP/IJCNLP (1)2
2019 Editing-Based SQL Query Generation for Cross-Domain Context-Dependent Questions
abstract
Rui Zhang, Tao Yu, Heyang Er, Sungrok Shim, Eric Xue, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Rui Zhang 0037, Tao Yu 0009, Heyang Er, Sungrok Shim, Eric Xue 0001, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir R. Radev
EMNLP/IJCNLP (1)1
2019 Harmonic Unpaired Image-to-image Translation
Rui Zhang 0037, Tomas Pfister, Li-Jia Li 0001
ICLR (Poster)1
2018 Addressee and Response Selection in Multi-Party Conversations With Speaker Interaction RNNs
abstract
In this paper, we study the problem of addressee and response selection in multi-party conversations. Understanding multi-party conversations is challenging because of complex speaker interactions: multiple speakers exchange messages with each other, playing different roles (sender, addressee, observer), and these roles vary across turns. To tackle this challenge, we propose the Speaker Interaction Recurrent Neural Network (SI-RNN). Whereas the previous state-of-the-art system updated speaker embeddings only for the sender, SI-RNN uses a novel dialog encoder to update speaker embeddings in a role-sensitive way. Additionally, unlike the previous work that selected the addressee and response separately, SI-RNN selects them jointly by viewing the task as a sequence prediction problem. Experimental results show that SI-RNN significantly improves the accuracy of addressee and response selection, particularly in complex conversations with many speakers and responses to distant messages many turns in the past.
Rui Zhang 0037, Honglak Lee, Lazaros Polymenakos, Dragomir R. Radev
AAAI1
2018 Improving Text-to-SQL Evaluation Methodology
abstract
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, Dragomir Radev. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang 0039, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang 0037, Dragomir R. Radev
ACL (1)6
2018 SyntaxSQLNet: Syntax Tree Networks for Complex and Cross-Domain Text-to-SQL Task
abstract
Most existing studies in text-to-SQL tasks do not require generating complex SQL queries with multiple clauses or sub-queries, and generalizing to new, unseen databases.In this paper we propose SyntaxSQLNet, a syntax tree network to address the complex and crossdomain text-to-SQL generation task.Syn-taxSQLNet employs a SQL specific syntax tree-based decoder with SQL generation path history and table-aware column attention encoders.We evaluate SyntaxSQLNet on a new large-scale text-to-SQL corpus containing databases with multiple tables and complex SQL queries containing multiple SQL clauses and nested queries.We use a database split setting where databases in the test set are unseen during training.Experimental results show that SyntaxSQLNet can handle a significantly greater number of complex SQL examples than prior work, outperforming the previous state-of-the-art model by 9.5% in exact matching accuracy.To our knowledge, we are the first to study this complex text-to-SQL task.Our task and models with the latest updates are available at https://yale-lily. github.io/seq2sql/spider.
Tao Yu 0009, Michihiro Yasunaga, Rui Zhang 0037, Zifan Li, Dragomir R. Radev
EMNLP4
2018 Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
abstract
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, Dragomir Radev. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Dragomir R. Radev
EMNLP2
2017 Graph-based Neural Multi-Document Summarization
abstract
We propose a neural multi-document summarization (MDS) system that incorporates sentence relation graphs.We employ a Graph Convolutional Network (GCN) on the relation graphs, with sentence embeddings obtained from Recurrent Neural Networks as input node features.Through multiple layer-wise propagation, the GCN generates high-level hidden sentence features for salience estimation.We then use a greedy heuristic to extract salient sentences while avoiding redundancy.In our experiments on DUC 2004, we consider three types of sentence relation graphs and demonstrate the advantage of combining sentence relations in graphs with the representation power of deep neural networks.Our model improves upon traditional graph-based extractive approaches and the vanilla GRU sequence model with no graph, and it achieves competitive results against other state-of-the-art multidocument summarization systems.
Michihiro Yasunaga, Rui Zhang 0037, Kshitijh Meelu, Ayush Pareek, Krishnan Srinivasan, Dragomir R. Radev
CoNLL2
2016 Effects of Creativity and Cluster Tightness on Short Text Clustering Performance
abstract
Properties of corpora, such as the diversity of vocabulary and how tightly related texts cluster together, impact the best way to cluster short texts.We examine several such properties in a variety of corpora and track their effects on various combinations of similarity metrics and clustering algorithms.We show that semantic similarity metrics outperform traditional n-gram and dependency similarity metrics for kmeans clustering of a linguistically creative dataset, but do not help with less creative texts.Yet the choice of similarity metric interacts with the choice of clustering method.We find that graphbased clustering methods perform well on tightly clustered data but poorly on loosely clustered data.Semantic similarity metrics generate loosely clustered output even when applied to a tightly clustered dataset.Thus, the best performing clustering systems could not use semantic metrics.
Catherine Finegan-Dollak, Reed Coke, Rui Zhang 0037, Xiangyi Ye, Dragomir R. Radev
ACL (1)3
2016 Dependency Sensitive Convolutional Neural Networks for Modeling Sentences and Documents
abstract
The goal of sentence and document modeling is to accurately represent the meaning of sentences and documents for various Natural Language Processing tasks.In this work, we present Dependency Sensitive Convolutional Neural Networks (DSCNN) as a generalpurpose classification system for both sentences and documents.DSCNN hierarchically builds textual representations by processing pretrained word embeddings via Long Short-Term Memory networks and subsequently extracting features with convolution operators.Compared with existing recursive neural models with tree structures, DSCNN does not rely on parsers and expensive phrase labeling, and thus is not restricted to sentencelevel tasks.Moreover, unlike other CNNbased models that analyze sentences locally by sliding windows, our system captures both the dependency information within each sentence and relationships across sentences in the same document.Experiment results demonstrate that our approach is achieving state-ofthe-art performance on several tasks, including sentiment analysis, question type classification, and subjectivity classification.
Rui Zhang 0037, Honglak Lee, Dragomir R. Radev
HLT-NAACL1