VLDB 2026 Research / reviewers in the wild / expert
Shafiq R. Joty
dblp:62/2078 · also Shafiq Joty, Shafiq Rayhan Joty
· DBLP profile ↗
176ranked-venue papers
18as first author
92since 2021 · last 2026
0000-0002-9222-2641ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 157 · 17 first-author · 84 since 2021Databases, data management, data science and information retrieval · 20 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathabstractLarge language model (LLM)-based reasoning systems have recently achieved gold medallevel performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also sufficiently supported.To train LLM-based reasoners in such challenging, open-ended settings, strong verifiers capable of catching step-level mistakes are necessary prerequisites.We introduce Hard2Verify, a human-annotated, step-level verification benchmark produced with over 500 hours of human labor.Hard2Verify is designed to rigorously assess step-level verifiers at the frontier: Verifiers must provide step-level annotations or identify the first error in responses generated by frontier LLMs for very recent, challenging, and open-ended math questions.We evaluate 29 generative critics and process reward models, demonstrating that, beyond a few standouts, open-source verifiers lag closed source models.We subsequently analyze what drives poor performance in step-level verification, the impacts of scaling verifier compute, as well as fundamental questions such as self-verification and verification-generation dynamics. Shrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming, Caiming Xiong, Shafiq R. Joty |
ACL (1) | 6 |
| 2026 | J4R: Learning to Judge with Equivalent Initial State Group Relative Policy OptimizationabstractTo keep pace with the increasing velocity of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic evaluation, where LLMs themselves are tasked with assessing and critiquing other model outputs.LLM-as-judge models are a class of generative evaluators that excel in evaluating relatively simple domains, like chat quality, but struggle in reasoning intensive domains where model responses contain more substantive and challenging content.To remedy existing judge shortcomings, we explore training judges with reinforcement learning (RL).We make three key contributions: (1) We propose the Equivalent Initial State Group Relative Policy Optimization (EIS-GRPO) algorithm, which allows us to train our judge to be robust to positional biases that arise in more complex evaluation settings.(2) We introduce ReasoningJudgeBench, a benchmark that evaluates judges in diverse reasoning settings not covered by prior work.(3) We train Judge for Reasoning (J4R), a 7B judge trained with EIS-GRPO that outperforms GPT-4o and the next best small judge by 6.7% and 9%, matching or exceeding the performance of larger GRPOtrained judges on both JudgeBench and Rea-soningJudgeBench. Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong, Shafiq R. Joty |
ACL (1) | 5 |
| 2025 | Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging TasksabstractLarge language models excel at problemsolving but often struggle with complex reasoning and factual accuracy.While chainof-thought and retrieval-augmented generation help break down problems and retrieve knowledge, they still falter on challenging tasks like competitive programming due to frequent reasoning errors and irrelevant retrieval.To address this, we introduce Critic-guided planning with Retrieval-augmentation, CR-Planner, a novel framework that leverages fine-tuned critic models to guide both reasoning and retrieval processes through planning.CR-Planner iteratively selects and executes sub-goals, guided by critic models.A sub-goal critic identifies promising sub-goals from reasoning, query generation, and retrieval, while an execution critic evaluates outputs of sub-goal executions.We employ Monte Carlo Tree Search to collect data for critic training, allowing systematic exploration of action sequences and effective navigation toward the final answer.We evaluate CR-Planner on challenging domain-knowledgeintensive and reasoning-heavy tasks, including competitive programming, theorem-driven math reasoning, and complex domain retrieval problems.It significantly outperforms baselines, demonstrating effectiveness in both reasoning and retrieval.Our code is available at https://github.com/xingxuanli/CR-Planner. Xingxuan Li, Weiwen Xu, Fangkai Jiao, Shafiq R. Joty, Lidong Bing |
ACL (1) | 5 |
| 2025 | What Makes a Good Natural Language Prompt?abstractDo Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq Joty, Min-Yen Kan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq R. Joty, Min-Yen Kan |
ACL (1) | 6 |
| 2025 | Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningabstractChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen, Yuchen Hu, Bosheng Ding, Ruirui Chen, Shafiq Joty. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen 0075, Bosheng Ding, Ruirui Chen 0002, Shafiq R. Joty |
ACL (1) | 8 |
| 2025 | Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsabstractThe large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-deployment monitoring. While judge models—LLMs finetuned to specialize in assessing and critiquing model outputs—have been touted as general purpose evaluators, they are typically evaluated only on non-contextual scenarios, such as instruction following. The omission of contextual settings—those where external information is used as context to generate an output—is surprising given the increasing prevalence of retrieval-augmented generation (RAG) and summarization use cases. Contextual assessment is uniquely challenging, as evaluation often depends on practitioner priorities, leading to conditional evaluation criteria (e.g., comparing responses based on factuality and then considering completeness if they are equally factual). To address the gap, we propose ContextualJudgeBench, a judge benchmark with 2,000 challenging response pairs across eight splits inspired by real-world contextual evaluation scenarios. We build our benchmark with a multi-pronged data construction pipeline that leverages both existing human annotations and model-based perturbations. Our comprehensive study across 11 judge models and 7 general purpose models, reveals that the contextual information and assessment criteria present a significant challenge to even state-of-the-art models. For example, o1, the best-performing model, barely reaches 55% consistent accuracy. Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, Shafiq R. Joty |
ACL (1) | 5 |
| 2025 | DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and AggregationabstractThe acceleration of Large Language Models (LLMs) research has opened up new possibilities for evaluating generated text. Though LLMs serve as scalable and economical evaluators, how reliable these evaluators is still under-explored. Prior research efforts in the meta-evaluation of LLMs as judges limit the prompting of an LLM to a single use to obtain a final evaluation decision. They then compute the agreement between LLMs’ outputs and human labels. This lacks interpretability in understanding the evaluation capability of LLMs. In light of this challenge, we propose DnA-Eval, which breaks down the evaluation process into decomposition and aggregation stages based on pedagogical practices. Our experiments show that it not only provides a more interpretable window for how well LLMs evaluate, but also leads to improvements up to 39.6% for different LLMs on a variety of meta-evaluation benchmarks. Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq R. Joty, Nancy F. Chen, Min-Yen Kan |
COLING | 4 |
| 2025 | CEMTM: Contextual Embedding-based Multimodal Topic ModelingabstractWe introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images.CEMTM builds on fine-tuned large vision language models (LVLMs) to obtain contextualized embeddings, and employs a distributional attention mechanism to weight token-level contributions to topic inference.A reconstruction objective aligns topic-based representations with the document embedding, encouraging semantic consistency across modalities.Unlike existing approaches, CEMTM can process multiple images per document without repeated encoding and maintains interpretability through explicit word-topic and documenttopic distributions.Extensive experiments on six multimodal benchmarks show that CEMTM consistently outperforms unimodal and multimodal baselines, achieving a remarkable average LLM score of 2.61 (1-3 scale).Further analysis shows its effectiveness in downstream few-shot retrieval and its ability to capture visually grounded semantics in complex domains such as scientific articles 1 . Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq R. Joty, Giuseppe Carenini |
EMNLP | 4 |
| 2025 | Demystifying Domain-adaptive Post-training for Financial LLMsabstractDomain-adaptive post-training of large language models (LLMs) has emerged as a promising approach for specialized domains such as medicine and finance.However, significant challenges remain in identifying optimal adaptation criteria and training strategies across varying data and model configurations.To address these challenges, we introduce FINDAP, a systematic and fine-grained investigation into domain-adaptive post-training of LLMs for the finance domain.Our approach consists of four key components: FinCap, which defines the core capabilities required for the target domain; FinRec, an effective training recipe that jointly optimizes continual pre-training and instruction-following, along with a novel preference data distillation method leveraging process signals from a generative reward model; FinTrain, a curated set of training datasets supporting FinRec; and FinEval, a comprehensive evaluation suite aligned with FinCap.The resulting model, Llama-Fin, achieves state-ofthe-art performance across a wide range of financial tasks.Our analysis also highlights how each post-training stage contributes to distinct capabilities, uncovering specific challenges and effective solutions, providing valuable insights for domain adaptation of LLMs. Zixuan Ke, Yifei Ming, Xuan-Phi Nguyen, Caiming Xiong, Shafiq R. Joty |
EMNLP | 5 |
| 2025 | From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-TextabstractRidwan Mahbub, Mohammed Saidul Islam, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar, Mizanur Rahman, Shafiq Joty, Enamul Hoque. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ridwan Mahbub, Mohammed Saidul Islam, Mir Tafseer Nayeem, Md. Tahmid Rahman Laskar, Shafiq R. Joty, Enamul Hoque Prince |
EMNLP | 6 |
| 2025 | Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextabstractAutomated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency.While large language models (LLMs) have shown promise in generating visualizations from natural language, the absence of comprehensive benchmarks limits the rigorous evaluation of their capabilities.We introduce Text2Vis, a benchmark designed to assess textto-visualization models, covering 20+ chart types and diverse data science queries, including trend analysis, correlation, outlier detection, and predictive analytics.It comprises 1,985 samples, each with a data table, natural language query, short answer, visualization code, and annotated charts.The queries involve complex reasoning, conversational turns, and dynamic data retrieval.We benchmark 11 open-source and closed-source models, revealing significant performance gaps, highlighting key challenges, and offering insights for future advancements.To close this gap, we propose the first cross-modal actor-critic agentic framework that jointly refines the textual answer and visualization code, increasing GPT-4o's pass rate from 26% to 42% over the direct approach and improving chart quality.We also introduce an automated LLM-based evaluation framework that enables scalable assessment across thousands of samples without human annotation, measuring answer correctness, code execution success, visualization readability, and chart accuracy.We release Text2Vis at https: //github.com/vis-nlp/Text2Vis. Md. Tahmid Rahman Laskar, Shafiq R. Joty, Enamul Hoque Prince |
EMNLP | 3 |
| 2025 | Direct Judgement Preference OptimizationabstractTo meet the increasing need for timely and accurate evaluation of large language model (LLM) responses, training LLM-as-judges to evaluate and critique other model responses has emerged as a popular paradigm.However, existing judge models are largely trained with supervised finetuning (SFT) on small data scales to perform limited types of evaluation tasks, fundamentally limiting generalization.To meet the need for strong, generalized judge models, we explore training foundational judge models at large data scales (680K) with direct preference optimization (DPO).Using four training tasks, we form three types of DPO preference pairs targeting different aspects of evaluation: Generating meaningful critiques, making accurate judgements, and understanding what comprises good and bad responses.To demonstrate the effectiveness of our method, we train judge models of three sizes: 8B parameters, 12B, and 70B, and evaluate on a comprehensive suite of 13 benchmarks (7 pairwise, 4 single rating, and 2 classification).Our models achieve the best aggregate performance, with even our 8B model outperforming GPT-4o in pairwise benchmarks.Further analysis shows that our judge models produce factual and actionable critiques and serve as strong foundational judges for continued finetuning. Peifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong, Shafiq R. Joty |
EMNLP | 5 |
| 2025 | Preference Optimization for Reasoning with Pseudo FeedbackabstractPreference optimization techniques, such as Direct Preference Optimization (DPO), are frequently employed to enhance the reasoning capabilities of large language models (LLMs) in domains like mathematical reasoning and coding, typically following supervised fine-tuning. These methods rely on high-quality labels for reasoning tasks to generate preference pairs; however, the availability of reasoning datasets with human-verified labels is limited.
In this study, we introduce a novel approach to generate pseudo feedback for reasoning tasks by framing the labeling of solutions to reason problems as an evaluation against associated \emph{test cases}.
We explore two forms of pseudo feedback based on test cases: one generated by frontier LLMs and the other by extending self-consistency to multi-test-case.
We conduct experiments on both mathematical reasoning and coding tasks using pseudo feedback for preference optimization, and observe improvements across both tasks. Specifically, using Mathstral-7B as our base model, we improve MATH results from 58.3 to 68.6, surpassing both NuminaMath-72B and GPT-4-Turbo-1106-preview. In GSM8K and College Math, our scores increase from 85.6 to 90.3 and from 34.3 to 42.3, respectively. Building on Deepseek-coder-7B-v1.5, we achieve a score of 24.3 on LiveCodeBench (from 21.1), surpassing Claude-3-Haiku. Fangkai Jiao, Geyang Guo, Nancy F. Chen, Shafiq R. Joty, Furu Wei |
ICLR | 5 |
| 2025 | FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"abstractEnsuring faithfulness to context in large language models (LLMs) and retrieval-augmented generation (RAG) systems is crucial for reliable deployment in real-world applications, as incorrect or unsupported information can erode user trust. Despite advancements on standard benchmarks, faithfulness hallucination—where models generate responses misaligned with the provided context—remains a significant challenge. In this work, we introduce FaithEval, a novel and comprehensive benchmark tailored to evaluate the faithfulness of LLMs in contextual scenarios across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts. These tasks simulate real-world challenges where retrieval mechanisms may surface incomplete, contradictory, or fabricated information. FaithEval comprises 4.9K high-quality problems in total, validated through a rigorous four-stage context construction and validation framework, employing both LLM-based auto-evaluation and human validation. Our extensive study across a wide range of open-source and proprietary models reveals that even state-of-the-art models often struggle to remain faithful to the given context, and that larger models do not necessarily exhibit improved faithfulness. Code is available at: https://github.com/SalesforceAIResearch/FaithEval. Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq R. Joty |
ICLR | 7 |
| 2025 | Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsabstractScaling test-time computation, or affording a generator large language model (LLM) extra compute during inference, typically employs the help of external non-generative evaluators (i.e., reward models). Concurrently, LLM-judges, models trained to generate evaluations and critiques (explanations) in natural language, are becoming increasingly popular in automatic evaluation. Despite judge empirical successes, their effectiveness as evaluators in test-time scaling settings is largely unknown. In this paper, we introduce the Judge Evaluation for Test-Time Scaling (JETTS) benchmark, which evaluates judge performance in three domains (math reasoning, code generation, and instruction following) under three task settings: response reranking, step-level beam search, and critique-based response refinement. We evaluate 10 different judge models (7B-70B parameters) for 8 different base generator models (6.7B-72B parameters). Our benchmark shows that while judges are competitive with outcome reward models in reranking, they are consistently worse than process reward models in beam search procedures. Furthermore, though unique to LLM-judges, their natural language critiques are currently ineffective in guiding the generator towards better responses. Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, Shafiq R. Joty |
ICML | 5 |
| 2025 | ParaICL: Towards Parallel In-Context LearningabstractXingxuan Li, Xuan-Phi Nguyen, Shafiq Joty, Lidong Bing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xingxuan Li, Xuan-Phi Nguyen, Shafiq R. Joty, Lidong Bing |
NAACL (Long Papers) | 3 |
| 2025 | ReIFE: Re-evaluating Instruction-Following EvaluationabstractYixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixin Liu 0003, Kejian Shi, Alexander R. Fabbri, Yilun Zhao 0001, Peifeng Wang, Chien-Sheng Wu, Shafiq R. Joty, Arman Cohan |
NAACL (Long Papers) | 7 |
| 2025 | LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMsabstractDo Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq R. Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan |
NAACL (Long Papers) | 5 |
| 2025 | On Positional Bias of Faithfulness for Long-form SummarizationabstractDavid Wan, Jesse Vig, Mohit Bansal, Shafiq Joty. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. David Wan, Jesse Vig, Mohit Bansal, Shafiq R. Joty |
NAACL (Long Papers) | 4 |
| 2025 | The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageabstractAs large language models (LLMs) continue to advance, their capacity to function effectively across a diverse range of languages has shown marked improvement. Preliminary studies observe that the hidden activations of LLMs often resemble English, even when responding to non-English prompts. This has led to the widespread assumption that LLMs may ``think'' in English. However, more recent results showing strong multilingual performance, even surpassing English performance on specific tasks in other languages, challenge this view. In this work, we find that LLMs progressively develop a core language-agnostic parameter space—a remarkably small subset of parameters whose deactivation results in significant performance degradation across all languages. This compact yet critical set of parameters underlies the model’s ability to generalize beyond individual languages, supporting the emergence of abstract thought that is not tied to any specific linguistic system. Specifically, we identify language-related neurons—those are consistently activated during the processing of particular languages, and categorize them as either shared (active across multiple languages) or exclusive (specific to one). As LLMs undergo continued development over time, we observe a marked increase in both the proportion and functional importance of shared neurons, while exclusive neurons progressively diminish in influence. These shared neurons constitute the backbone of the core language-agnostic parameter space, supporting the emergence of abstract thought. Motivated by these insights, we propose neuron-specific training strategies tailored to LLMs' language-agnostic levels at different development stages. Experiments across diverse LLM families support our approach. Our codes are available at https://anonymous.4open.science/status/S-C393. Yiran Zhao 0006, Yang Zhang 0072, An Zhang 0003, Kenji Kawaguchi, Shafiq R. Joty, Junnan Li 0001, Tat-Seng Chua, Michael Shieh, Wenxuan Zhang 0001 |
NeurIPS | 6 |
| 2025 | Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement LearningabstractReinforcement learning (RL) has become the dominant paradigm for improving the performance of language models on complex reasoning tasks. Despite the substantial empirical gains demonstrated by RL-based training methods like GRPO, a granular understanding of why and how RL enhances performance is still lacking. To bridge this gap, we introduce SPARKLE, a fine-grained analytic framework to dissect the effects of RL across three key dimensions: (1) plan following and execution, (2) knowledge integration, and (3) chain of subproblems. Using this framework, we gain insights beyond mere accuracy. For instance, providing models with explicit human-crafted, step-by-step plans can surprisingly degrade performance on the most challenging benchmarks, yet RL-tuned models exhibit greater robustness, experiencing markedly smaller performance drops than base or SFT models. This suggests that RL may not primarily enhance the execution of external plans but rather empower models to formulate and follow internal strategies better suited to their reasoning processes. Conversely, we observe that RL enhances models' ability to integrate provided knowledge into their reasoning process, yielding consistent gains across diverse tasks. Finally, we study whether difficult problems---those yielding no RL signals and mixed-quality reasoning traces---can still be effectively used for training. We introduce SparkleRL-PSS, a multi-stage RL pipeline that reuses hard problems with partial step scaffolding, guiding exploration effectively without additional data generation. Together, our findings provide a principled foundation for understanding how RL shapes model behavior, offering practical insights for building more adaptive, data-efficient, and interpretable RL pipelines for reasoning tasks. Our code, data, and checkpoints are available at: https://sparkle-reasoning.github.io/. Yifei Ming, Zixuan Ke, Caiming Xiong, Shafiq R. Joty, Aws Albarghouthi, Frederic Sala |
NeurIPS | 5 |
| 2025 | From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation ModelsabstractData visualization in the form of charts plays a pivotal role in data analysis, offering critical insights and aiding in informed decision-making. Automatic chart understanding has witnessed significant advancements with the rise of large foundation models in recent years. Foundation models, such as large language models, have revolutionized various natural language processing tasks and are increasingly being applied to chart understanding tasks. This survey paper provides a comprehensive overview of the recent developments, challenges, and future directions in chart understanding within the context of these foundation models. We review fundamental building blocks crucial for studying chart understanding tasks. Additionally, we explore various tasks and their evaluation metrics and sources of both charts and textual inputs. Various modeling strategies are then examined, encompassing both classification-based and generation-based approaches, along with tool augmentation techniques that enhance chart understanding performance. Furthermore, we discuss the state-of-the-art performance of each task and discuss how we can improve the performance. Challenges and future directions are addressed, highlighting the importance of several topics, such as domain-specific charts, lack of efforts in developing evaluation metrics, and agent-oriented settings. This survey paper aims to provide valuable insights and directions for future research in chart understanding leveraging large foundation models. Kung-Hsiang Huang, Hou Pong Chan, May Fung, Haoyi Qiu, Shafiq R. Joty, Shih-Fu Chang, Heng Ji 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalabstractMohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, Shafiq Joty. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammad Abdullah Matin Khan, Saiful Bari, Xuan Do Long, Weishi Wang, Md. Rizwan Parvez, Shafiq R. Joty |
ACL (1) | 6 |
| 2024 | Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsabstractLarge language models (LLMs) are known to perform tasks by simply observing few exemplars.Moreover, competent generative capabilities of LLMs are observed mostly in highresource languages, while their performances among under-represented languages fall behind due to pre-training data imbalance.To elicit LLMs' ability onto low-resource languages without any supervised data, we propose to assemble synthetic exemplars from a diverse set of high-resource languages.These prompts can directly induce generative capabilities in lowresource languages and serve as intra-lingual exemplars to even improve tasks in these languages.Our unsupervised prompting method performs on par with supervised few-shot learning in LLMs of different sizes for translations between English and 34 Indic and African languages, and surpasses supervised prompting in non-English tasks.The method also significantly improves low-resource performances in many other intra-lingual tasks like summarization (XLSum), question answering (XQUAD & TydiQA) and conversational instruction following (Sea-Bench). Xuan-Phi Nguyen, Mahani Aljunied, Shafiq R. Joty, Lidong Bing |
ACL (1) | 3 |
| 2024 | On Context Utilization in Summarization with Large Language ModelsabstractLarge language models (LLMs) excel in abstractive summarization tasks, delivering fluent and pertinent summaries.Recent advancements have extended their capabilities to handle long-input contexts, exceeding 100k tokens.However, in question answering, language models exhibit uneven utilization of their input context.They tend to favor the initial and final segments, resulting in a U-shaped performance pattern concerning where the answer is located within the input.This bias raises concerns, particularly in summarization where crucial content may be dispersed throughout the source document(s).Besides, in summarization, mapping facts from the source to the summary is not trivial as salient content is usually re-phrased.In this paper, we conduct the first comprehensive study on context utilization and position bias in summarization.Our analysis encompasses 6 LLMs, 10 datasets, and 5 evaluation metrics.We introduce a new evaluation benchmark called MiddleSum on the which we benchmark two alternative inference methods to alleviate position bias: hierarchical summarization and incremental summarization 1 .Metric Model CNN/DM XSum Reddit SAMSum Multi-X AVG Arxiv PubMed GovReport SummScreenFD Multi-N AVG ROUGE-2 Flan-UL2 -0.296 -0.124 0.048 -0.069 -0.201 -0.128 _ _ _ _ _ _ Llama-2-7B -0.160 -0.023 0.063 -0.059 -0.100 -0.056 0.022 -0.113 -0.109 -0.079 -0.210 -0.098 Llama-2-13B -0.166 -0.086 0.031 -0.078 -0.039 -0.068 -0.017 -0.081 -0.166 -0.139 -0.213 -0.123 Xgen-7B -0.228 -0.042 0.066 -0.039 -0.041 -0.056 0.028 -0.091 -0.405 0.063 -0.283 -0.138 Mistral-7B -0.289 -0.031 0.006 -0.024 -0.052 -0.078 -0.270 -0.279 -0.585 -0.132 -0.324 -0.318 GPT-3.5 -0.323 -0.027 -0.031 -0.097 0.088 -0.078 0.026 -0.093 -0.123 -0.061 -0.233 -0.097 BERTScore Flan-UL2 -0.331 -0.185 0.062 -0.144 -0.399 -0. Mathieu Ravaut, Aixin Sun, Nancy F. Chen, Shafiq R. Joty |
ACL (1) | 4 |
| 2024 | Diffusion Model Alignment Using Direct Preference OptimizationabstractLarge language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; the best existing approach is to fine-tune a pretrained model using carefully curated high quality images and captions to improve visual appeal and text alignment. We propose Diffusion-DPO, a method to align diffusion models to human preferences by directly optimizing on human comparison data. Diffusion-DPO is adapted from the recently developed Direct Preference Optimization (DPO) [36], a simpler alternative to RLHF which directly optimizes a policy that best satisfies human preferences under a classification objective. We re-formulate DPO to account for a diffusion model notion of likelihood, utilizing the evidence lower bound to derive a differentiable objective. Using the Pick-a-Pic dataset of 851K crowdsourced pairwise preferences, we fine-tune the base model of the state-of-the-art Stable Diffusion XL (SDXL)-1.0 model with Diffusion-DPO. Our fine-tuned base model significantly outperforms both base SDXL-1.0 and the larger SDXL-1.0 model consisting of an additional refinement model in human evaluation, improving visual appeal and prompt alignment. We also develop a variant that uses AI feedback and has comparable performance to training on human preferences, opening the door for scaling of diffusion model alignment methods. Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq R. Joty |
CVPR | 9 |
| 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
Artemis Panagopoulou, Le Xue, Ning Yu 0006, Junnan Li 0001, Dongxu Li 0003, Shafiq R. Joty, Ran Xu 0001, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles |
ECCV (45) | 6 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 30 |
| 2024 | DataNarrative: Automated Data-Driven Storytelling with Visualizations and TextsabstractData-driven storytelling is a powerful method for conveying insights by combining narrative techniques with visualizations and text. These stories integrate visual aids, such as highlighted bars and lines in charts, along with textual annotations explaining insights. However, creating such stories requires a deep understanding of the data and meticulous narrative planning, often necessitating human intervention, which can be time-consuming and mentally taxing. While Large Language Models (LLMs) excel in various NLP tasks, their ability to generate coherent and comprehensive data stories remains underexplored. In this work, we introduce a novel task for data story generation and a benchmark containing 1,449 stories from diverse sources. To address the challenges of crafting coherent data stories, we propose a multi-agent framework employing two LLM agents designed to replicate the human storytelling process: one for understanding and describing the data (Reflection), generating the outline, and narration, and another for verification at each intermediary step. While our agentic framework generally outperforms non-agentic counterparts in both model-based and human evaluations, the results also reveal unique challenges in data story generation. Mohammed Saidul Islam, Md. Tahmid Rahman Laskar, Md. Rizwan Parvez, Enamul Hoque Prince, Shafiq R. Joty |
EMNLP | 5 |
| 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward SynthesizingabstractLarge Language Models (LLMs) have demonstrated significant potential in handling complex reasoning tasks through step-by-step rationale generation.However, recent studies have raised concerns regarding the hallucination and flaws in their reasoning process.Substantial efforts are being made to improve the reliability and faithfulness of the generated rationales.Some approaches model reasoning as planning, while others focus on annotating for process supervision.Nevertheless, the planning-based search process often results in high latency due to the frequent assessment of intermediate reasoning states and the extensive exploration space.Additionally, supervising the reasoning process with human annotation is costly and challenging to scale for LLM training.To address these issues, in this paper, we propose a framework to learn planning-based reasoning through Direct Preference Optimization (DPO) on collected trajectories, which are ranked according to our synthesized process rewards.Our results on challenging logical reasoning benchmarks demonstrate the effectiveness of our learning framework, showing that our 7B model can surpass the strong counterparts like GPT-3.5-Turbo. Fangkai Jiao, Chengwei Qin, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty |
EMNLP | 5 |
| 2024 | A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsabstractMd Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Md. Tahmid Rahman Laskar, Sawsan Alqahtani, Saiful Bari, Mohammad Abdullah Matin Khan, Haidar Khan, Amran Bhuiyan, Chee-Wei Tan 0001, Md. Rizwan Parvez, Enamul Hoque Prince, Shafiq R. Joty, Jimmy Huang 0001 |
EMNLP | 12 |
| 2024 | Evaluating Psychological Safety of Large Language ModelsabstractIn this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs).First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI).All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern.Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3.Then, we evaluated the LLMs in the GPT series by using well-being tests to study the impact of fine-tuning with more training data.We observed a continuous increase in the well-being scores of GPT models.Following these observations, we showed that finetuning Llama-2-chat-7B with responses from BFI using direct preference optimization could effectively reduce the psychological toxicity of the model.Based on the findings, we recommended the application of systematic and comprehensive psychological metrics to further evaluate and improve the safety of LLMs.Our code is available at https://github.com/DAMO- NLP-SG/PsychSafety.Warning: This paper contains examples with potentially harmful content. Xingxuan Li, Shafiq R. Joty, Lidong Bing |
EMNLP | 4 |
| 2024 | CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modulesabstractLarge Language Models (LLMs) have already become quite proficient at solving simpler programming tasks like those in HumanEval or MBPP benchmarks. However, solving more complex and competitive programming tasks is still quite challenging for these models - possibly due to their tendency to generate solutions as monolithic code blocks instead of decomposing them into logical sub-tasks and sub-modules. On the other hand, experienced programmers instinctively write modularized code with abstraction for solving complex tasks, often reusing previously developed modules. To address this gap, we propose CodeChain, a novel framework for inference that elicits modularized code generation through a chain of self-revisions, each being guided by some representative sub-modules generated in previous iterations. Concretely, CodeChain first instructs the LLM to generate modularized codes through chain-of-thought prompting. Then it applies a chain of self-revisions by iterating the two steps: 1) extracting and clustering the generated sub-modules and selecting the cluster representatives as the more generic and re-usable implementations, and 2) augmenting the original chain-of-thought prompt with these selected module-implementations and instructing the LLM to re-generate new modularized solutions. We find that by naturally encouraging the LLM to reuse the previously developed and verified sub-modules, CodeChain can significantly boost both modularity as well as correctness of the generated solutions, achieving relative pass@1 improvements of 35\% on APPS and 76\% on CodeContests. It is shown to be effective on both OpenAI LLMs as well as open-sourced LLMs like WizardCoder. We also conduct comprehensive ablation studies with different methods of prompting, number of clusters, model sizes, program qualities, etc., to provide useful insights that underpin CodeChain's success. Hung Le 0003, Hailin Chen, Amrita Saha, Akash Gokul, Doyen Sahoo, Shafiq R. Joty |
ICLR | 6 |
| 2024 | Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous SourcesabstractWe present chain-of-knowledge (CoK), a novel framework that augments large language models (LLMs) by dynamically incorporating grounding information from heterogeneous sources. It results in more factual rationales and reduced hallucination in generation.
Specifically, CoK consists of three stages: reasoning preparation, dynamic knowledge adapting, and answer consolidation.
Given a knowledge-intensive question, CoK first prepares several preliminary rationales and answers while identifying the relevant knowledge domains.
If there is no majority consensus among the answers from samples, CoK corrects the rationales step by step by adapting knowledge from the identified domains.
These corrected rationales can plausibly serve as a better foundation for the final answer consolidation.
Unlike prior studies that primarily use unstructured data, CoK also leverages structured knowledge sources such as Wikidata and tables that provide more reliable factual information.
To access both unstructured and structured knowledge sources in the dynamic knowledge adapting stage, we propose an adaptive query generator that allows the generation of queries for various types of query languages, including SPARQL, SQL, and natural sentences. Moreover, to minimize error propagation between rationales, CoK corrects the rationales progressively using preceding corrected rationales to generate and correct subsequent rationales.
Extensive experiments show that CoK consistently improves the performance of LLMs on knowledge-intensive tasks across different domains. Xingxuan Li, Yew Ken Chia, Bosheng Ding, Shafiq R. Joty, Soujanya Poria, Lidong Bing |
ICLR | 5 |
| 2024 | Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News ArticlesabstractKung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kung-Hsiang Huang, Philippe Laban, Alexander R. Fabbri, Prafulla Kumar Choubey, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
NAACL-HLT | 5 |
| 2024 | Exploring Self-supervised Logic-enhanced Training for Large Language ModelsabstractFangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy Chen, Shafiq Joty. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty |
NAACL-HLT | 6 |
| 2024 | L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language ModelsabstractAbstract Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models’ language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research. All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https://l2c-eval.github.io/. Ansong Ni, Yilun Zhao 0001, Martin Riddell, Troy Feng, Stephen Yin, Ye Liu 0006, Semih Yavuz, Caiming Xiong, Shafiq R. Joty, Yingbo Zhou 0002, Dragomir R. Radev, Arman Cohan |
Trans. Assoc. Comput. Linguistics | 11 |
| 2024 | Improving Conversational Recommender System Via Contextual and Time-Aware Modeling With Less Domain-Specific KnowledgeabstractConversational Recommender Systems (CRS) has become an emerging research topic seeking to perform recommendations through interactive conversations, which generally consist of generation and recommendation modules. Prior work on CRS tends to incorporate more external and domain-specific knowledge like item reviews to enhance performance. Despite the fact that the collection and annotation of theexternal domain-specificinformation needs much human effort and degenerates the generalizability, too much extra knowledge introduces more difficulty to balance among them. Therefore, we propose to fully discover and extract theinternalknowledge from the context. We capture both entity-level and contextual-level representations to jointly model user preferences for the recommendation, where a time-aware attention is designed to emphasize the recently appeared items in entity-level representations. We further use the pre-trained BART to initialize the generation module to alleviate the data scarcity and enhance the context modeling. In addition to conducting experiments on a popular dataset (ReDial), we also include a multi-domain dataset (OpenDialKG) to show the effectiveness of our model. Experiments on both datasets show that our model achieves better performance on most evaluation metrics with less external knowledge and generalizes well to other domains. Additional analyses on the recommendation and generation tasks demonstrate the effectiveness of our model in different scenarios. Lingzhi Wang 0001, Shafiq R. Joty, Wei Gao 0001, Xingshan Zeng, Kam-Fai Wong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Is GPT-3 a Good Data Annotator?abstractBosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Boyang Li, Shafiq Joty, Lidong Bing. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bosheng Ding, Chengwei Qin, Yew Ken Chia, Boyang Li 0001, Shafiq R. Joty, Lidong Bing |
ACL (1) | 6 |
| 2023 | Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question GenerationabstractXuan Long Do, Bowei Zou, Shafiq Joty, Tran Tai, Liangming Pan, Nancy Chen, Ai Ti Aw. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Do Xuan Long, Bowei Zou, Shafiq R. Joty, Anh Tran Tai, Liangming Pan, Nancy F. Chen, AiTi Aw |
ACL (1) | 3 |
| 2023 | SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesabstractPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 4 |
| 2023 | Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationabstractYixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
ACL (1) | 8 |
| 2023 | Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsabstractDue to the huge amount of parameters, finetuning of pretrained language models (PLMs) is prone to overfitting in the low resource scenarios.In this work, we present a novel method that operates on the hidden representations of a PLM to reduce overfitting.During fine-tuning, our method inserts random autoencoders between the hidden layers of a PLM, which transform activations from the previous layers into multi-view compressed representations before feeding them into the upper layers.The autoencoders are plugged out after fine-tuning, so our method does not add extra parameters or increase computation cost during inference.Our method demonstrates promising performance improvement across a wide range of sequenceand token-level low-resource NLP tasks.Our code is available at https://github.com/DAMO- NLP-SG/MVCR. Xingxuan Li, Megh Thakkar, Xin Li 0056, Shafiq R. Joty, Luo Si, Lidong Bing |
ACL (1) | 5 |
| 2023 | Randomized Smoothing with Masked Inference for Adversarially Robust Text ClassificationsabstractLarge-scale pre-trained language models have shown outstanding performance in a variety of NLP tasks.However, they are also known to be significantly brittle against specifically crafted adversarial examples, leading to increasing interest in probing the adversarial robustness of NLP systems.We introduce RSMI, a novel two-stage framework that combines randomized smoothing (RS) with masked inference (MI) to improve the adversarial robustness of NLP systems.RS transforms a classifier into a smoothed classifier to obtain robust representations, whereas MI forces a model to exploit the surrounding context of a masked token in an input sequence.RSMI improves adversarial robustness by 2 to 3 times over existing state-of-the-art methods on benchmark datasets.We also perform in-depth qualitative analysis to validate the effectiveness of the different stages of RSMI and probe the impact of its components through extensive ablations.By empirically proving the stability of RSMI, we put it forward as a practical method to robustly train large-scale NLP models.Our code and datasets are available at https://github.com/Han8931/rsmi_nlp. Han Cheol Moon, Shafiq R. Joty, Megh Thakkar |
ACL (1) | 2 |
| 2023 | Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?abstractPrompt tuning (PT) which only tunes the embeddings of an additional sequence of tokens per task, keeping the pre-trained language model (PLM) frozen, has shown remarkable performance in few-shot learning.Despite this, PT has been shown to rely heavily on good initialization of the prompt embeddings.In this work, we study meta prompt tuning (MPT) to systematically explore how meta-learning can help improve (if it can) cross-task generalization in PT through learning to initialize the prompt embeddings from other relevant tasks.We empirically analyze a representative set of meta learning algorithms in a wide range of adaptation settings with different source/target task configurations on a large set of few-shot tasks.With extensive experiments and analysis, we demonstrate the effectiveness of MPT.We find the improvement to be significant particularly on classification tasks.For other kinds of tasks such as question answering, we observe that while MPT can outperform PT in most cases, it does not always outperform multi-task learning.We further provide an in-depth analysis from the perspective of task similarity. Chengwei Qin, Shafiq R. Joty, Qian Li 0043 |
ACL (1) | 2 |
| 2023 | Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction LearningabstractLarge language models (LLMs) have shown impressive performance in following natural language instructions to solve unseen tasks.However, it remains unclear whether models truly understand task definitions and whether the human-written definitions are optimal.In this paper, we systematically study the role of task definitions in instruction learning.We first conduct an ablation analysis informed by human annotations to understand which parts of a task definition are most important, and find that model performance only drops substantially when removing contents describing the task output, in particular label information.Next, we propose an automatic algorithm to compress task definitions to a minimal supporting set of tokens, and find that 60% of tokens can be removed while maintaining or even improving model performance.Based on these results, we propose two strategies to help models better leverage task instructions: (1) providing only key information for tasks in a common structured format, and (2) adding a metatuning stage to help the model better understand the definitions.With these two strategies, we achieve a 4.2 Rouge-L improvement over 119 unseen test tasks. Fan Yin, Jesse Vig, Philippe Laban, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 4 |
| 2023 | Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought FrameworkabstractAs large language models (LLMs) have become the norm in NLP, demonstrating good performance in generation and reasoning tasks, one of its most fatal disadvantages is the lack of factual correctness.Generating unfactual texts not only leads to lower performances but also degrades the trust and validity of their applications.Chain-of-Thought (CoT) prompting improves trust and model performance on complex reasoning tasks by generating interpretable reasoning chains, but still suffers from factuality concerns in knowledge-intensive tasks.In this paper, we propose the Verify-and-Edit framework for CoT prompting, which seeks to increase prediction factuality by post-editing reasoning chains according to external knowledge.Building on top of GPT-3, our framework lead to accuracy improvements in multiple open-domain question-answering tasks.For reproducing our results and extending the framework further, we make our codebase available at https://github.com/RuochenZhao/Verify- and-Edit * Equal contribution. Xingxuan Li, Shafiq R. Joty, Chengwei Qin, Lidong Bing |
ACL (1) | 3 |
| 2023 | Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code GenerationabstractWith the rise of powerful closed-sourced LLMs (ChatGPT, GPT-4), there are increasing interests in distilling the capabilies of close-sourced LLMs to smaller open-sourced LLMs.Previous distillation methods usually prompt Chat-GPT to generate a set of instructions and answers, for the student model to learn.However, such standard distillation approach neglects the merits and conditions of the student model.Inspired by modern teaching principles, we design a personalised distillation process, in which the student attempts to solve a task first, then the teacher provides an adaptive refinement for the student to improve.Instead of feeding the student with teacher's prior, personalised distillation enables personalised learning for the student model, as it only learns on examples it makes mistakes upon and learns to improve its own solution.On code generation, personalised distillation consistently outperforms standard distillation with only one third of the data.With only 2.5-3K personalised examples that incur a data-collection cost of 4-6$, we boost CodeGen-mono-16B by 7% to achieve 36.4% pass@1 and StarCoder by 12.2% to achieve 45.8% pass@1 on HumanEval. Hailin Chen, Amrita Saha, Steven C. H. Hoi, Shafiq R. Joty |
EMNLP | 4 |
| 2023 | SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationabstractPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander Fabbri, Caiming Xiong, Shafiq Joty, Chien-Sheng Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri, Caiming Xiong, Shafiq R. Joty, Chien-Sheng Wu |
EMNLP | 6 |
| 2023 | Towards Interpretable and Efficient Automatic Reference-Based Summarization EvaluationabstractYixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
EMNLP | 5 |
| 2023 | UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningabstractCharts are widely used for data analysis, providing visual representations and insights into complex data.To facilitate chart-based data analysis using natural language, several downstream tasks have been introduced recently such as chart question answering and chart summarization.However, existing methods for these tasks often rely on pretraining on language or vision-language tasks, neglecting the explicit modeling of chart structures (e.g., how chart elements are related to each other).To address this, we first build a large corpus of charts covering diverse topics and visual styles.We then present UniChart, a pretrained model for chart comprehension and reasoning.UniChart encodes the relevant text, data, and visual elements of charts and then uses a chart-grounded text decoder for text generation.We propose several chart-specific pretraining tasks that include: (i) low-level tasks to extract the visual elements (e.g., bars, lines) and data from charts, and (ii) high-level tasks to acquire chart understanding and reasoning skills.Our experiments demonstrate that pretraining UniChart on a large corpus with chart-specific objectives, followed by fine-tuning, yields state-of-the-art performance on four downstream tasks.Moreover, our model exhibits superior generalizability to unseen chart corpus, surpassing previous approaches that lack chart-specific objectives and utilize limited chart resources. Ahmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque Prince, Shafiq R. Joty |
EMNLP | 5 |
| 2023 | Lifelong Sequence Generation with Dynamic Module Expansion and AdaptationabstractLifelong sequence generation (LSG), a problem in continual learning, aims to continually train a model on a sequence of generation tasks to learn constantly emerging new generation patterns while avoiding the forgetting of previous knowledge.Existing LSG methods mainly focus on maintaining old knowledge while paying little attention to knowledge transfer across tasks.In contrast, humans can better learn new tasks by leveraging previously acquired knowledge from similar tasks.Inspired by the learning paradigm of humans, we propose Dynamic Module Expansion and Adaptation (DMEA), which enables the model to dynamically determine the architecture for acquiring new knowledge based on task correlation and select the most similar previous tasks to facilitate adaptation to new tasks.In addition, as the learning process can easily be biased towards the current task which might cause more severe forgetting of previously learned knowledge, we propose dynamic gradient scaling to balance the learning of the current task and replayed tasks.With extensive experiments, we demonstrate that DMEA can consistently outperform existing methods in different LSG settings. Chengwei Qin, Chen Chen 0075, Shafiq R. Joty |
EMNLP | 3 |
| 2023 | Towards Low-Resource Automatic Program Repair with Meta-Learning and Pretrained Language ModelsabstractAutomatic program repair (APR) has gained increasing attention as an essential technique in software development to reduce manual debugging efforts and boost developers' productivity.Recent advances in deep learning (DL) based models have demonstrated promising results by learning from large-scale bug-fix examples in a data-driven manner.However, in practical scenarios, software bugs have an imbalanced distribution, and the fixing knowledge learned by APR models often only capture the patterns of frequent error types, making it inapplicable to handle the rare error types.To address this limitation, we investigate a novel task of low-resource APR, and propose Meta-APR, a new meta-learning framework integrated with code pretrained language models to generate fixes for low-resource bugs with limited training samples.Our Meta-APR learns better errorspecific knowledge from high-resource bugs through efficient first-order meta-learning optimization, which allows for a faster adaptation to the target low-resource bugs.Besides, while we adopt CodeT5, a pretrained code-aware encoder-decoder Transformer, as the backbone model for Meta-APR, it is a model-agnostic framework that can be integrated with any neural models.Extensive experimental results on three benchmarks in various programming languages verify the superiority of our method over existing DL-based APR approaches. Weishi Wang, Yue Wang 0034, Steven C. H. Hoi, Shafiq R. Joty |
EMNLP | 4 |
| 2023 | Efficient Text-to-Code Retrieval with Cascaded Fast and Slow Transformer ModelsabstractThe goal of semantic code search or text-to-code search is to retrieve a semantically relevant code snippet from an existing code database using a natural language query. When constructing a practical semantic code search system, existing approaches fail to provide an optimal balance between retrieval speed and the relevance of the retrieved results. We propose an efficient and effective text-to-code search framework with cascaded fast and slow models, in which a fast transformer encoder model is learned to optimize a scalable index for fast retrieval followed by learning a slow classification-based re-ranking model to improve the accuracy of the top K results from the fast retrieval. To further reduce the high memory cost of deploying two separate models in practice, we propose to jointly train the fast and slow model based on a single transformer encoder with shared parameters. Empirically our cascaded method is not only efficient and scalable, but also achieves state-of-the-art results with an average mean reciprocal ranking (MRR) score of 0.7795 (across 6 programming languages) on the CodeSearchNet benchmark as opposed to the prior state-of-the-art result of 0.744 MRR. Our codebase can be found at this link. Akhilesh Gotmare, Junnan Li 0001, Shafiq R. Joty, Steven C. H. Hoi |
ESEC/SIGSOFT FSE | 3 |
| 2023 | RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program RepairabstractAutomatic program repair (APR) is crucial to reduce manual debugging efforts for developers and improve software reliability. While conventional search-based techniques typically rely on heuristic rules or a redundancy assumption to mine fix patterns, recent years have witnessed the surge of deep learning (DL) based approaches to automate the program repair process in a data-driven manner. However, their performance is often limited by a fixed set of parameters to model the highly complex search space of APR. Weishi Wang, Yue Wang 0034, Shafiq R. Joty, Steven C. H. Hoi |
ESEC/SIGSOFT FSE | 3 |
| 2023 | PIANO: Influence Maximization Meets Deep Reinforcement LearningabstractSince its introduction in 2003, the influence maximization (IM) problem has drawn significant research attention in the literature. The aim of IM, which is NP-hard, is to select a set of$k$users known as seed users who can influence the most individuals in the social network. The state-of-the-art algorithms estimate the expected influence of nodes based on sampled diffusion paths. As the number of required samples has been recently proven to be lower bounded by a particular threshold that presets tradeoff between the accuracy and the efficiency, the result quality of these traditional solutions is hard to be further improved without sacrificing efficiency. In this article, we present an orthogonal and novel paradigm to address the IM problem by leveraging deep reinforcement learning (RL) to estimate the expected influence. In particular, we present a novel framework called deeP reInforcement leArning-based iNfluence maximizatiOn (PIANO) that incorporates network embedding and RL techniques to address this problem. In order to make it practical, we further present PIANO-E and PIANO$\text{@}\langle d\rangle $, both of which can be applied directly to answer IM without training the model from scratch. Experimental study on real-world networks demonstrates that PIANO achieves the best performance with respect to efficiency and influence spread quality compared to state-of-the-art classical solutions. We also demonstrate that the learned parametric models generalize well across different networks. Besides, we provide a pool of pretrained PIANO models such that any IM task can be addressed by directly applying a model from the pool without training over the targeted network. Hui Li 0005, Mengting Xu, Sourav S. Bhowmick, Shafiq R. Joty, Changsheng Sun, Jiangtao Cui |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | UNISON: Unpaired Cross-Lingual Image CaptioningabstractImage captioning has emerged as an interesting research field in recent years due to its broad application scenarios. The traditional paradigm of image captioning relies on paired image-caption datasets to train the model in a supervised manner. However, creating such paired datasets for every target language is prohibitively expensive, which hinders the extensibility of captioning technology and deprives a large part of the world population of its benefit. In this work, we present a novel unpaired cross-lingual method to generate image captions without relying on any caption corpus in the source or the target language. Specifically, our method consists of two phases: (1) a cross-lingual auto-encoding process, which utilizing a sentence parallel (bitext) corpus to learn the mapping from the source to the target language in the scene graph encoding space and decode sentences in the target language, and (2) a cross-modal unsupervised feature mapping, which seeks to map the encoded scene graph features from image modality to language modality. We verify the effectiveness of our proposed method on the Chinese image caption generation task. The comparisons against several existing methods demonstrate the effectiveness of our approach. Jiahui Gao 0002, Yi Zhou 0042, Philip L. H. Yu, Shafiq R. Joty, Jiuxiang Gu |
AAAI | 4 |
| 2022 | Weakly Supervised Neuro-Symbolic Module Networks for Numerical Reasoning over TextabstractNeural Module Networks (NMNs) have been quite successful in incorporating explicit reasoning as learnable modules in various question answering tasks, including the most generic form of numerical reasoning over text in Machine Reading Comprehension (MRC). However to achieve this, contemporary Neural Module Networks models obtain strong supervision in form of specialized program annotation from the QA pairs through various heuristic parsing and exhaustive computation of all possible discrete operations on discrete arguments. Consequently they fail to generalize to more open-ended settings without such supervision. Hence, we propose Weakly Supervised Neuro-Symbolic Module Network (WNSMN) trained with answers as the sole supervision for numerical reasoning based MRC. WNSMN learns to execute a noisy heuristic program obtained from the dependency parse of the query, as discrete actions over both neural and symbolic reasoning modules and trains it end-to-end in a reinforcement learning framework with discrete reward from answer matching. On the subset of DROP having numerical answers, WNSMN outperforms NMN by 32% and the reasoning-free generative language model GenBERT by 8% in exact match accuracy under comparable weakly supervised settings. This showcases the effectiveness of modular networks that can handle explicit discrete reasoning over noisy programs in an end-to-end manner. Amrita Saha, Shafiq R. Joty, Steven C. H. Hoi |
AAAI | 2 |
| 2022 | GlobalWoZ: Globalizing MultiWoZ to Develop Multilingual Task-Oriented Dialogue SystemsabstractBosheng Ding, Junjie Hu, Lidong Bing, Mahani Aljunied, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Bosheng Ding, Junjie Hu 0001, Lidong Bing, Sharifah Mahani Aljunied, Shafiq R. Joty, Luo Si, Chunyan Miao |
ACL (1) | 5 |
| 2022 | Rethinking Self-Supervision Objectives for Generalizable Coherence ModelingabstractGiven the claims of improved text generation quality across various pre-trained neural models, we consider the coherence evaluation of machine generated text to be one of the principal applications of coherence models that needs to be investigated.Prior work in neural coherence modeling has primarily focused on devising new architectures for solving the permuted document task.We instead use a basic model architecture and show significant improvements over state of the art within the same training regime.We then design a harder self-supervision objective by increasing the ratio of negative samples within a contrastive learning setup, and enhance the model further through automatic hard negative mining coupled with a large global negative queue encoded by a momentum encoder.We show empirically that increasing the density of negative samples improves the basic model, and using a global negative queue further improves and stabilizes the model while training with hard negative samples.We evaluate the coherence model on task-independent test sets that resemble real-world applications and show significant improvements in coherence evaluations of downstream tasks. 1 Prathyusha Jwalapuram, Shafiq R. Joty |
ACL (1) | 2 |
| 2022 | Chart-to-Text: A Large-Scale Benchmark for Chart SummarizationabstractShankar Kantharaj, Rixie Tiffany Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, Shafiq Joty. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shankar Kantharaj, Rixie Tiffany Ko Leong, Ahmed Masry, Megh Thakkar, Enamul Hoque Prince, Shafiq R. Joty |
ACL (1) | 7 |
| 2022 | Continual Few-shot Relation Learning via Embedding Space Regularization and Data AugmentationabstractExisting continual relation learning (CRL) methods rely on plenty of labeled training data for learning a new task, which can be hard to acquire in real scenario as getting large and representative labeled data is often expensive and time-consuming.It is therefore necessary for the model to learn novel relational patterns with very few labeled data while avoiding catastrophic forgetting of previous task knowledge.In this paper, we formulate this challenging yet practical problem as continual few-shot relation learning (CFRL).Based on the finding that learning for new emerging few-shot tasks often results in feature distributions that are incompatible with previous tasks' learned distributions, we propose a novel method based on embedding space regularization and data augmentation.Our method generalizes to new fewshot tasks and avoids catastrophic forgetting of previous tasks by enforcing extra constraints on the relational embeddings and by adding extra relevant data in a self-supervised manner.With extensive experiments we demonstrate that our method can significantly outperform previous state-of-the-art methods in CFRL task settings.1 Chengwei Qin, Shafiq R. Joty |
ACL (1) | 2 |
| 2022 | SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive SummarizationabstractSequence-to-sequence neural networks have recently achieved great success in abstractive summarization, especially through fine-tuning large pre-trained language models on the downstream dataset.These models are typically decoded with beam search to generate a unique summary.However, the search space is very large, and with the exposure bias, such decoding is not optimal.In this paper, we show that it is possible to directly train a secondstage model performing re-ranking on a set of summary candidates.Our mixture-of-experts SummaReranker learns to select a better candidate and consistently improves the performance of the base model.With a base PEGASUS, we push ROUGE scores by 5.44% on CNN-DailyMail (47.16 ROUGE-1), 1.31% on XSum (48.12 ROUGE-1) and 9.34% on Reddit TIFU (29.83 ROUGE-1), reaching a new state-of-theart.Our code and checkpoints will be available at https://github.com/ntunlp/ SummaReranker. Mathieu Ravaut, Shafiq R. Joty, Nancy F. Chen |
ACL (1) | 2 |
| 2022 | CoHS-CQG: Context and History Selection for Conversational Question GenerationabstractConversational question generation (CQG) serves as a vital task for machines to assist humans, such as interactive reading comprehension, through conversations. Compared to traditional single-turn question generation (SQG), CQG is more challenging in the sense that the generated question is required not only to be meaningful, but also to align with the provided conversation. Previous studies mainly focus on how to model the flow and alignment of the conversation, but do not thoroughly study which parts of the context and history are necessary for the model. We believe that shortening the context and history is crucial as it can help the model to optimise more on the conversational alignment property. To this end, we propose CoHS-CQG, a two-stage CQG framework, which adopts a novel CoHS module to shorten the context and history of the input. In particular, it selects the top-p sentences and history turns by calculating the relevance scores of them. Our model achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings. Do Xuan Long, Bowei Zou, Liangming Pan, Nancy F. Chen, Shafiq R. Joty, AiTi Aw |
COLING | 5 |
| 2022 | Towards Multi-Sense Cross-Lingual Alignment of Contextual EmbeddingsabstractCross-lingual word embeddings (CLWE) have been proven useful in many cross-lingual tasks. However, most existing approaches to learn CLWE including the ones with contextual embeddings are sense agnostic. In this work, we propose a novel framework to align contextual embeddings at the sense level by leveraging cross-lingual signal from bilingual dictionaries only. We operationalize our framework by first proposing a novel sense-aware cross entropy loss to model word senses explicitly. The monolingual ELMo and BERT models pretrained with our sense-aware cross entropy loss demonstrate significant performance improvement for word sense disambiguation tasks. We then propose a sense alignment objective on top of the sense-aware cross entropy loss for cross-lingual model pretraining, and pretrain cross-lingual models for several language pairs (English to German/Spanish/Japanese/Chinese). Compared with the best baseline results, our cross-lingual models achieve 0.52%, 2.09% and 1.29% average performance improvements on zero-shot cross-lingual NER, sentiment classification and XNLI tasks, respectively. Thien Hai Nguyen, Shafiq R. Joty, Lidong Bing, Luo Si |
COLING | 3 |
| 2022 | Learning Label Modular Prompts for Text Classification in the WildabstractMachine learning models usually assume i.i.d data during training and testing, but data and tasks in real world often change over time.To emulate the transient nature of real world, we propose a challenging but practical task: text classification in-the-wild, which introduces different non-stationary training/testing stages.Decomposing a complex task into modular components can enable robust generalisation under such non-stationary environment.However, current modular approaches in NLP do not take advantage of recent advances in parameter efficient tuning of pretrained language models.To close this gap, we propose MOD-ULARPROMPT, a label-modular prompt tuning framework for text classification tasks.In MOD-ULARPROMPT, the input prompt consists of a sequence of soft label prompts, each encoding modular knowledge related to the corresponding class label.In two of most formidable settings, MODULARPROMPT outperforms relevant baselines by a large margin demonstrating strong generalisation ability.We also conduct comprehensive analysis to validate whether the learned prompts satisfy properties of a modular representation. 1 Hailin Chen, Amrita Saha, Shafiq R. Joty, Steven C. H. Hoi |
EMNLP | 3 |
| 2022 | OpenCQA: Open-ended Question Answering with ChartsabstractCharts are very popular to analyze data and convey important insights. People often analyze visualizations to answer open-ended questions that require explanatory answers. Answering such questions are often difficult and time-consuming as it requires a lot of cognitive and perceptual efforts. To address this challenge, we introduce a new task called OpenCQA, where the goal is to answer an open-ended question about a chart with descriptive texts. We present the annotation process and an in-depth analysis of our dataset. We implement and evaluate a set of baselines under three practical settings. In the first setting, a chart and the accompanying article is provided as input to the model. The second setting provides only the relevant paragraph(s) to the chart instead of the entire article, whereas the third setting requires the model to generate an answer solely based on the chart. Our analysis of the results show that the top performing models generally produce fluent and coherent text while they struggle to perform complex logical and arithmetic reasoning. Shankar Kantharaj, Do Xuan Long, Rixie Tiffany Ko Leong, Jia Qing Tan, Enamul Hoque Prince, Shafiq R. Joty |
EMNLP | 6 |
| 2022 | Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesabstractKnowledge-enhanced language representation learning has shown promising results across various knowledge-intensive NLP tasks.However, prior methods are limited in efficient utilization of multilingual knowledge graph (KG) data for language model (LM) pretraining.They often train LMs with KGs in indirect ways, relying on extra entity/relation embeddings to facilitate knowledge injection.In this work, we explore methods to make better use of the multilingual annotation and language agnostic property of KG triples, and present novel knowledge based multilingual language models (KMLMs) trained directly on the knowledge triples.We first generate a large amount of multilingual synthetic sentences using the Wikidata KG triples.Then based on the intra-and inter-sentence structures of the generated data, we design pretraining tasks to enable the LMs to not only memorize the factual knowledge but also learn useful logical patterns.Our pretrained KMLMs demonstrate significant performance improvements on a wide range of knowledge-intensive crosslingual tasks, including named entity recognition (NER), factual knowledge retrieval, relation classification, and a newly designed logical reasoning task. 1 Xin Li 0056, Ruidan He, Lidong Bing, Shafiq R. Joty, Luo Si |
EMNLP | 5 |
| 2022 | Towards Summary Candidates FusionabstractSequence-to-sequence deep neural models fine-tuned for abstractive summarization can achieve great performance on datasets with enough human annotations.Yet, it has been shown that they have not reached their full potential, with a wide gap between the top beam search output and the oracle beam.Recently, re-ranking methods have been proposed, to learn to select a better summary candidate.However, such methods are limited by the summary quality aspects captured by the first-stage candidates.To bypass this limitation, we propose a new paradigm in secondstage abstractive summarization called Sum-maFusion that fuses several summary candiates to produce a novel abstractive secondstage summary.Our method works well on several summarization datasets, improving both the ROUGE scores and qualitative properties of fused summaries.It is especially good when the candidates to fuse are worse, such as in the few-shot setup where we set a new state-of-the-art.We will make our code and checkpoints available at https: //github.com/ntunlp/SummaFusion/. Mathieu Ravaut, Shafiq R. Joty, Nancy F. Chen |
EMNLP | 2 |
| 2022 | Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised Translation
Xuan-Phi Nguyen, Hongyu Gong, Yun Tang 0002, Changhan Wang, Philipp Koehn, Shafiq R. Joty |
ICLR | 6 |
| 2022 | LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5
Chengwei Qin, Shafiq R. Joty |
ICLR | 2 |
| 2022 | GradMask: Gradient-Guided Token Masking for Textual Adversarial Example DetectionabstractWe present GradMask, a simple adversarial example detection scheme for natural language processing (NLP) models. It uses gradient signals to detect adversarially perturbed tokens in an input sequence and occludes such tokens by a masking process. GradMask provides several advantages over existing methods including improved detection performance and an interpretation of its decision with a only moderate computational cost. Its approximated inference cost is no more than a single forward- and back-propagation through the target model without requiring any additional detection module. Extensive evaluation on widely adopted NLP benchmark datasets demonstrates the efficiency and effectiveness of GradMask. Code and models are available at https://github.com/Han8931/grad_mask_detection Han Cheol Moon, Shafiq R. Joty, Xu Chi |
KDD | 2 |
| 2022 | Refining Low-Resource Unsupervised Translation by Language Disentanglement of Multilingual Translation ModelabstractNumerous recent work on unsupervised machine translation (UMT) implies that competent unsupervised translations of low-resource and unrelated languages, such as Nepali or Sinhala, are only possible if the model is trained in a massive multilingual environment, where these low-resource languages are mixed with high-resource counterparts. Nonetheless, while the high-resource languages greatly help kick-start the target low-resource translation tasks, the language discrepancy between them may hinder their further improvement. In this work, we propose a simple refinement procedure to separate languages from a pre-trained multilingual UMT model for it to focus on only the target low-resource task. Our method achieves the state of the art in the fully unsupervised translation tasks of English to Nepali, Sinhala, Gujarati, Latvian, Estonian and Kazakh, with BLEU score gains of 3.5, 3.5, 3.3, 4.1, 4.2, and 3.3, respectively. Our codebase is available at https://github.com/nxphi47/refineunsupmultilingual_mt Xuan-Phi Nguyen, Shafiq R. Joty, Kui Wu 0004, AiTi Aw |
NeurIPS | 2 |
| 2022 | LANTERN: Boredom-conscious Natural Language Description Generation of Query Execution Plans for Database EducationabstractThe database systems course in an undergraduate computer science degree program is gaining increasing importance due to the continuous supply of database-related jobs as well as the rise of Data Science. A key learning goal of learners taking such a course is to understand how SQL queries are executed in an RDBMS in practice. An RDBMS typically exposes a query execution plan (QEP) in a visual or textual format, which describes the execution steps for a given query. However, it is often daunting for a learner to comprehend these QEPs containing vendor-specific implementation details. In this demonstration, we present a novel, generic, and portable system called LANTERN that generates a natural language (NL)-based description of the execution strategy chosen by the underlying RDBMS to process a query. It provides a declarative framework called POOL for subject matter experts (SME) to efficiently create and manipulate the NL descriptions of physical operators of any RDBMS. It then exploits POOL to generate the NL descriptions of QEPs by integrating a rule-based and a deep learning-based techniques to infuse language variability in the descriptions. Such an NL generation strategy mitigates the impact of boredom on learners caused by repeated exposure of similar text generated by a rule-based system. Hui Li 0005, Sourav S. Bhowmick, Shafiq R. Joty, Weiguo Wang |
SIGMOD Conference | 4 |
| 2021 | UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPabstractM Saiful Bari, Tasnim Mohiuddin, Shafiq Joty. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Saiful Bari, Tasnim Mohiuddin, Shafiq R. Joty |
ACL/IJCNLP (1) | 3 |
| 2021 | MulDA: A Multilingual Data Augmentation Framework for Low-Resource Cross-Lingual NERabstractLinlin Liu, Bosheng Ding, Lidong Bing, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bosheng Ding, Lidong Bing, Shafiq R. Joty, Luo Si, Chunyan Miao |
ACL/IJCNLP (1) | 4 |
| 2021 | A Conditional Splitting Framework for Efficient Constituency ParsingabstractThanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq Joty, Xiaoli Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Thanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli Li 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Reliability Testing for Natural Language Processing SystemsabstractSamson Tan, Shafiq Joty, Kathy Baxter, Araz Taeihagh, Gregory A. Bennett, Min-Yen Kan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Samson Tan, Shafiq R. Joty, Kathy Baxter, Araz Taeihagh, Gregory A. Bennett, Min-Yen Kan |
ACL/IJCNLP (1) | 2 |
| 2021 | Span-Level Emotion Cause Analysis by BERT-based Graph Attention NetworkabstractWe study the task of span-level emotion cause analysis (SECA), which is focused on identifying the specific emotion cause span(s) triggering a certain emotion in the text. Compared to the popular clause-level emotion cause analysis (CECA), it is a finer-grained emotion cause analysis (ECA) task. In this paper, we design a BERT-based graph attention network for emotion cause span(s) identification. The proposed model takes advantage of the structure of BERT to capture the relationship information between emotion and text, and utilizes graph attention network to model the structure information of the text. Our SECA method can be easily used for extracting clause-level emotion causes for CECA as well. Experimental results show that the proposed method consistently outperforms the state-of-the-art ECA methods on benchmark emotion cause dataset. Xiangju Li, Wei Gao 0001, Shi Feng 0001, Daling Wang, Shafiq R. Joty |
CIKM | 5 |
| 2021 | Span-level Emotion Cause Analysis with Neural Sequence TaggingabstractThis paper addresses the task of span-level emotion cause analysis (SECA). It is a finer-grained emotion cause analysis (ECA) task, which aims to identify the specific emotion cause span(s) behind certain emotions in text. In this paper, we formalize SECA as a sequence tagging task for which several variants of neural network-based sequence tagging models to extract specific emotion cause span(s) in the given context. These models combine different types of encoding and decoding approaches. Furthermore, to make our models more "emotionally sensitive'', we utilize the multi-head attention mechanism to enhance the representation of context. Experimental evaluations conducted on two benchmark datasets demonstrate the effectiveness of the proposed models. Xiangju Li, Wei Gao 0001, Shi Feng 0001, Daling Wang, Shafiq R. Joty |
CIKM | 5 |
| 2021 | Rethinking Coherence Modeling: Synthetic vs. Downstream TasksabstractAlthough coherence modeling has come a long way in developing novel models, their evaluation on downstream applications for which they are purportedly developed has largely been neglected.With the advancements made by neural approaches in applications such as machine translation (MT), summarization and dialog systems, the need for coherence evaluation of these tasks is now more crucial than ever.However, coherence models are typically evaluated only on synthetic tasks, which may not be representative of their performance in downstream applications.To investigate how representative the synthetic tasks are of downstream use cases, we conduct experiments on benchmarking well-known traditional and neural coherence models on synthetic sentence ordering tasks, and contrast this with their performance on three downstream applications: coherence evaluation for MT and summarization, and next utterance prediction in retrieval-based dialog.Our results demonstrate a weak correlation between the model performances in the synthetic tasks and the downstream applications, motivating alternate training and evaluation methods for coherence models. 1 Tasnim Mohiuddin, Prathyusha Jwalapuram, Shafiq R. Joty |
EACL | 4 |
| 2021 | CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationabstractPre-trained models for Natural Languages (NL) like BERT and GPT have been recently shown to transfer well to Programming Languages (PL) and largely benefit a broad set of code-related tasks.Despite their success, most current methods either rely on an encoder-only (or decoder-only) pre-training that is suboptimal for generation (resp.understanding) tasks or process the code snippet in the same way as NL, neglecting the special characteristics of PL such as token types.We present CodeT5, a unified pre-trained encoder-decoder Transformer model that better leverages the code semantics conveyed from the developer-assigned identifiers.Our model employs a unified framework to seamlessly support both code understanding and generation tasks and allows for multi-task learning.Besides, we propose a novel identifier-aware pre-training task that enables the model to distinguish which code tokens are identifiers and to recover them when they are masked.Furthermore, we propose to exploit the user-written code comments with a bimodal dual generation task for better NL-PL alignment.Comprehensive experiments show that CodeT5 significantly outperforms prior methods on understanding tasks such as code defect detection and clone detection, and generation tasks across various directions including PL-NL, NL-PL, and PL-PL.Further analysis reveals that our model can better capture semantic information from code.Our code and pre-trained models are released at https: //github.com/salesforce/CodeT5. Yue Wang 0034, Weishi Wang, Shafiq R. Joty, Steven C. H. Hoi |
EMNLP (1) | 3 |
| 2021 | Effective Fine-Tuning Methods for Cross-lingual AdaptationabstractLarge scale multilingual pre-trained language models have shown promising results in zeroand few-shot cross-lingual tasks.However, recent studies have shown their lack of generalizability when the languages are structurally dissimilar.In this work, we propose a novel fine-tuning method based on co-training that aims to learn more generalized semantic equivalences as complementary to multilingual language modeling using the unlabeled data in the target language.We also propose an adaption method based on contrastive learning to better capture the semantic relationship in the parallel data, when a few translation pairs are available.To show our method's effectiveness, we conduct extensive experiments on cross-lingual inference and review classification tasks across various languages.We report significant gains compared to directly finetuning multilingual pre-trained models and other semi-supervised alternatives.1 Tao Yu 0009, Shafiq R. Joty |
EMNLP (1) | 2 |
| 2021 | A Unified Speaker Adaptation Approach for ASRabstractTransformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a trained model with target speaker data is the most natural approach for adaptation, but it takes a lot of compute and may cause catastrophic forgetting to the existing speakers. In this work, we propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. For feature adaptation, we employ a speaker-aware persistent memory model which generalizes better to unseen test speakers by making use of speaker i-vectors to form a persistent memory. For model adaptation, we use a novel gradual pruning method to adapt to target speakers without changing the model architecture, which to the best of our knowledge, has never been explored in ASR. Specifically, we gradually prune less contributing parameters on model encoder to a certain sparsity level, and use the pruned parameters for adaptation, while freezing the unpruned parameters to keep the original model performance. We conduct experiments on the Librispeech dataset. Our proposed approach brings relative 2.74-6.52% word error rate (WER) reduction on general speaker adaptation. On target speaker adaptation, our method outperforms the baseline with up to 20.58% relative WER reduction, and surpasses the finetuning method by up to relative 2.54%. Besides, with extremely low-resource adaptation data (e.g., 1 utterance), our method could improve the WER by relative 6.53% with only a few epochs of training. Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
EMNLP (1) | 4 |
| 2021 | Preventing Early Endpointing for Online Automatic Speech RecognitionabstractWith the recent development of end-to-end models in speech recognition, there have been more interests in adapting these models for online speech recognition. However, using end-to-end models for online speech recognition is known to suffer from an early endpointing problem, which brings in many deletion errors. In this paper, we propose to address the early endpointing problem from the gradient perspective. Specifically, we leverage on the recently proposed ScaleGrad technique, which was proposed to mitigate the text degeneration issue. Different from ScaleGrad, we adapt it to discourage the early generation of the end-of-sentence () token. A scaling term is added to directly maneuver the gradient of the training loss to encourage the model to learn to keep generating non-tokens. Compared with previous approaches such as voice-activity-detection and end-of-query detection, the proposed method does not rely on various types of silence, and it also saves the trouble from obtaining the ground truth endpoint with forced alignment. Nevertheless, it can be jointly applied with other techniques. Experiments on AISHELL-1 dataset show that our model brings relative 5.4%-10.1% CER reductions over the baseline, and surpasses the unlikelihood training method which directly reduces the generation probability oftoken. Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 4 |
| 2021 | Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationabstractAdvanced large-scale neural language models have led to significant success in many language generation tasks. However, the most commonly used training objective, Maximum Likelihood Estimation (MLE), has been shown problematic, where the trained model prefers using dull and repetitive phrases. In this work, we introduce ScaleGrad, a modification straight to the gradient of the loss function, to remedy the degeneration issue of the standard MLE objective. By directly maneuvering the gradient information, ScaleGrad makes the model learn to use novel tokens. Empirical results show the effectiveness of our method not only in open-ended generation, but also in directed generation tasks. With the simplicity in architecture, our method can serve as a general training objective that is applicable to most of the neural text generation tasks. Simeng Han, Shafiq R. Joty |
ICML | 3 |
| 2021 | Cross-model Back-translated Distillation for Unsupervised Machine TranslationabstractRecent unsupervised machine translation (UMT) systems usually employ three main principles: initialization, language modeling and iterative back-translation, though they may apply them differently. Crucially, iterative back-translation and denoising auto-encoding for language modeling provide data diversity to train the UMT systems. However, the gains from these diversification processes has seemed to plateau. We introduce a novel component to the standard UMT framework called Cross-model Back-translated Distillation (CBD), that is aimed to induce another level of data diversification that existing principles lack. CBD is applicable to all previous UMT approaches. In our experiments, CBD achieves the state of the art in the WMT’14 English-French, WMT’16 English-German and English-Romanian bilingual unsupervised translation tasks, with 38.2, 30.1, and 36.3 BLEU respectively. It also yields 1.5–3.3 BLEU improvements in IWSLT English-French and English-German tasks. Through extensive experimental analyses, we show that CBD is effective because it embraces data diversity while other similar variants do not. Xuan-Phi Nguyen, Shafiq R. Joty, Thanh-Tung Nguyen, Kui Wu 0004, AiTi Aw |
ICML | 2 |
| 2021 | Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data AugmentationabstractAlexander Fabbri, Simeng Han, Haoyuan Li, Haoran Li, Marjan Ghazvininejad, Shafiq Joty, Dragomir Radev, Yashar Mehdad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Alexander R. Fabbri, Simeng Han, Haoran Li 0007, Marjan Ghazvininejad, Shafiq R. Joty, Dragomir R. Radev, Yashar Mehdad |
NAACL-HLT | 6 |
| 2021 | RST Parsing from ScratchabstractWe introduce a novel top-down end-to-end formulation of document level discourse parsing in the Rhetorical Structure Theory (RST) framework. In this formulation, we consider discourse parsing as a sequence of splitting decisions at token boundaries and use a seq2seq network to model the splitting decisions. Our framework facilitates discourse parsing from scratch without requiring discourse segmentation as a prerequisite; rather, it yields segmentation as part of the parsing process. Our unified parsing model adopts a beam search to decode the best tree structure by searching through a space of high scoring trees. With extensive experiments on the standard RST discourse treebank, we demonstrate that our parser outperforms existing methods by a good margin in both end-to-end parsing and parsing with gold segmentation. More importantly, it does so without using any handcrafted features, making it faster and easily adaptable to new languages and domains. Thanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli Li 0001 |
NAACL-HLT | 3 |
| 2021 | Code-Mixing on Sesame Street: Dawn of the Adversarial PolyglotsabstractMultilingual models have demonstrated impressive cross-lingual transfer performance.However, test sets like XNLI are monolingual at the example level.In multilingual communities, it is common for polyglots to code-mix when conversing with each other.Inspired by this phenomenon, we present two strong blackbox adversarial attacks (one word-level, one phrase-level) for multilingual models that push their ability to handle code-mixed sentences to the limit.The former uses bilingual dictionaries to propose perturbations and translations of the clean example for sense disambiguation.The latter directly aligns the clean example with its translations before extracting phrases as perturbations.Our phrase-level attack has a success rate of 89.75% against XLM-R large , bringing its average accuracy of 79.85 down to 8.18 on XNLI.Finally, we propose an efficient adversarial training scheme that trains in the same number of steps as the original model and show that it improves model accuracy.1 Original P: The girl that can help me is all the way across town.H: There is no one who can help me.Adversary P: olan girl that can help me is all the way across town.H: one who can help me. Samson Tan, Shafiq R. Joty |
NAACL-HLT | 2 |
| 2021 | Align before Fuse: Vision and Language Representation Learning with Momentum DistillationabstractLarge-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens. Because the visual tokens and word tokens are unaligned, it is challenging for the multimodal encoder to learn image-text interactions. In this paper, we introduce a contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning. Unlike most existing methods, our method does not require bounding box annotations nor high-resolution images. In order to improve learning from noisy web data, we propose momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model. We provide a theoretical analysis of ALBEF from a mutual information maximization perspective, showing that different training tasks can be interpreted as different ways to generate views for an image-text pair. ALBEF achieves state-of-the-art performance on multiple downstream vision-language tasks. On image-text retrieval, ALBEF outperforms methods that are pre-trained on orders of magnitude larger datasets. On VQA and NLVR$^2$, ALBEF achieves absolute improvements of 2.37% and 3.84% compared to the state-of-the-art, while enjoying faster inference speed. Code and models are available at https://github.com/salesforce/ALBEF. Junnan Li 0001, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, Steven C. H. Hoi |
NeurIPS | 4 |
| 2021 | Towards Enhancing Database Education: Natural Language Generation Meets Query Execution PlansabstractThe database systems course is offered as part of an undergraduate computer science degree program in many major universities. A key learning goal of learners taking such a course is to understand how sql queries are processed in a rdbms in practice. Since aquery execution plan (qep ) describes the execution steps of a query, learners can acquire the understanding by perusing the qep s generated by a rdbms. Unfortunately, in practice, it is often daunting for a learner to comprehend these qep s containing vendor-specific implementation details, hindering her learning process. In this paper, we present a novel, end-to-end,generic system called lantern that generates a natural language description of a qep to facilitate understanding of the query execution steps. It takes as input an sql query and its qep, and generates a natural language description of the execution strategy deployed by the underlying rdbms. Specifically, it deploys adeclarative framework called pool that enablessubject matter experts to efficiently create and maintain natural language descriptions of physical operators used in qep s. Arule-based framework called rule-lantern is proposed that exploits pool to generate natural language descriptions of qep s. Despite the high accuracy of rule-lantern, our engagement with learners reveal that, consistent with existing psychology theories, perusing such rule-based descriptions lead toboredom due to repetitive statements across different qep s. To address this issue, we present a noveldeep learning-based language generation framework called neural -lantern that infuses language variability in the generated description by exploiting a set ofparaphrasing tools andword embedding. Our experimental study with real learners shows the effectiveness of lantern in facilitating comprehension of qep s. Weiguo Wang, Sourav S. Bhowmick, Hui Li 0005, Shafiq R. Joty |
SIGMOD Conference | 4 |
| 2020 | Zero-Resource Cross-Lingual Named Entity RecognitionabstractRecently, neural methods have achieved state-of-the-art (SOTA) results in Named Entity Recognition (NER) tasks for many languages without the need for manually crafted features. However, these models still require manually annotated training data, which is not available for many languages. In this paper, we propose an unsupervised cross-lingual NER model that can transfer NER knowledge from one language to another in a completely unsupervised way without relying on any bilingual dictionary or parallel data. Our model achieves this through word-level adversarial learning and augmented fine-tuning with parameter sharing and feature augmentation. Experiments on five different languages demonstrate the effectiveness of our approach, outperforming existing models by a good margin and setting a new SOTA for each language pair. Saiful Bari, Shafiq R. Joty, Prathyusha Jwalapuram |
AAAI | 2 |
| 2020 | Explicit Memory Tracker with Coarse-to-Fine Reasoning for Conversational Machine ReadingabstractYifan Gao, Chien-Sheng Wu, Shafiq Joty, Caiming Xiong, Richard Socher, Irwin King, Michael Lyu, Steven C.H. Hoi. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Yifan Gao 0001, Chien-Sheng Wu, Shafiq R. Joty, Caiming Xiong, Richard Socher, Irwin King, Michael R. Lyu, Steven C. H. Hoi |
ACL | 3 |
| 2020 | Efficient Constituency Parsing by PointingabstractWe propose a novel constituency parsing model that casts the parsing problem into a series of pointing tasks.Specifically, our model estimates the likelihood of a span being a legitimate tree constituent via the pointing score corresponding to the boundary words of the span.Our parsing model supports efficient top-down decoding and our learning objective is able to enforce structural consistency without resorting to the expensive CKY inference.The experiments on the standard English Penn Treebank parsing task show that our method achieves 92.78 F1 without using pre-trained models, which is higher than all the existing methods with similar time complexity.Using pre-trained BERT, our model achieves 95.48 F1, which is competitive with the state-of-theart while being faster.Our approach also establishes new state-of-the-art in Basque and Swedish in the SPMRL shared tasks on multilingual constituency parsing. Thanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli Li 0001 |
ACL | 3 |
| 2020 | Differentiable Window for Dynamic Local AttentionabstractWe propose Differentiable Window, a new neural module and general purpose component for dynamic window selection.While universally applicable, we demonstrate a compelling use case of utilizing Differentiable Window to improve standard attention modules by enabling more focused attentions over the input regions.We propose two variants of Differentiable Window, and integrate them within the Transformer architecture in two novel ways.We evaluate our proposed approach on a myriad of NLP tasks, including machine translation, sentiment analysis, subject-verb agreement and language modeling.Our experimental results demonstrate consistent and sizable improvements across all tasks. Thanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli Li 0001 |
ACL | 3 |
| 2020 | It's Morphin' Time! Combating Linguistic Discrimination with Inflectional PerturbationsabstractTraining on only perfect Standard English corpora predisposes pre-trained neural networks to discriminate against minorities from nonstandard linguistic backgrounds (e.g., African American Vernacular English, Colloquial Singapore English, etc.).We perturb the inflectional morphology of words to craft plausible and semantically similar adversarial examples that expose these biases in popular NLP models, e.g., BERT and Transformer, and show that adversarially fine-tuning them for a single epoch significantly improves robustness without sacrificing performance on clean data. 1 Samson Tan, Shafiq R. Joty, Min-Yen Kan, Richard Socher |
ACL | 2 |
| 2020 | Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning
Xiangxi Shi, Xu Yang 0021, Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001 |
ECCV (14) | 4 |
| 2020 | DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksabstractBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Bosheng Ding, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq R. Joty, Luo Si, Chunyan Miao |
EMNLP (1) | 6 |
| 2020 | Discern: Discourse-Aware Entailment Reasoning Network for Conversational Machine ReadingabstractYifan Gao, Chien-Sheng Wu, Jingjing Li, Shafiq Joty, Steven C.H. Hoi, Caiming Xiong, Irwin King, Michael Lyu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yifan Gao 0001, Chien-Sheng Wu, Jingjing Li 0007, Shafiq R. Joty, Steven C. H. Hoi, Caiming Xiong, Irwin King, Michael R. Lyu |
EMNLP (1) | 4 |
| 2020 | Pronoun-Targeted Fine-tuning for NMT with Hybrid LossesabstractPopular Neural Machine Translation model training uses strategies like backtranslation to improve BLEU scores, requiring large amounts of additional data and training. We introduce a class of conditional generative-discriminative hybrid losses that we use to fine-tune a trained machine translation model. Through a combination of targeted fine-tuning objectives and intuitive re-use of the training data the model has failed to adequately learn from, we improve the model performance of both a sentence-level and a contextual model without using any additional data. We target the improvement of pronoun translations through our fine-tuning and evaluate our models on a pronoun benchmark testset. Our sentence-level model shows a 0.5 BLEU improvement on both the WMT14 and the IWSLT13 De-En testsets, while our contextual model achieves the best results, improving from 31.81 to 32 BLEU on WMT14 De-En testset, and from 32.10 to 33.13 on the IWSLT13 De-En testset, with corresponding improvements in pronoun translation. We further show the generalizability of our method by reproducing the improvements on two additional language pairs, Fr-En and Cs-En. Prathyusha Jwalapuram, Shafiq R. Joty, Youlin Shen |
EMNLP (1) | 2 |
| 2020 | LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent SpaceabstractMost of the successful and predominant methods for Bilingual Lexicon Induction (BLI) are mapping-based, where a linear mapping function is learned with the assumption that the word embedding spaces of different languages exhibit similar geometric structures (i.e., approximately isomorphic).However, several recent studies have criticized this simplified assumption showing that it does not hold in general even for closely related languages.In this work, we propose a novel semi-supervised method to learn cross-lingual word embeddings for BLI.Our model is independent of the isomorphic assumption and uses non-linear mapping in the latent space of two independently pre-trained autoencoders.Through extensive experiments on fifteen (15) different language pairs (in both directions) comprising resource-rich and low-resource languages from two different datasets, we demonstrate that our method outperforms existing models by a good margin.Ablation studies show the importance of different model components and the necessity of non-linear mapping. Tasnim Mohiuddin, Saiful Bari, Shafiq R. Joty |
EMNLP (1) | 3 |
| 2020 | Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection EncodingabstractInflectional variation is a common feature of World Englishes such as Colloquial Singapore English and African American Vernacular English.Although comprehension by human readers is usually unimpaired by nonstandard inflections, current NLP systems are not yet robust.We propose Base-Inflection Encoding (BITE), a method to tokenize English text by reducing inflected words to their base forms before reinjecting the grammatical information as special symbols.Fine-tuning pretrained NLP models for downstream tasks using our encoding defends against inflectional adversaries while maintaining performance on clean data.Models using BITE generalize better to dialects with non-standard inflections without explicit training and translation models converge faster when trained with BITE.Finally, we show that our encoding improves the vocabulary efficiency of popular data-driven subword tokenizers.Since there has been no prior work on quantitatively evaluating vocabulary efficiency, we propose metrics to do so. 1 Samson Tan, Shafiq R. Joty, Lav R. Varshney, Min-Yen Kan |
EMNLP (1) | 2 |
| 2020 | Response Selection for Multi-Party Conversations with Dynamic Topic TrackingabstractWhile participants in a multi-party multi-turn conversation simultaneously engage in multiple conversation topics, existing response selection methods are developed mainly focusing on a two-party single-conversation scenario.Hence, the prolongation and transition of conversation topics are ignored by current methods.In this work, we frame response selection as a dynamic topic tracking task to match the topic between the response and relevant conversation context.With this new formulation, we propose a novel multi-task learning framework that supports efficient encoding through large pretrained models with only two utterances at once to perform dynamic topic disentanglement and response selection.We also propose Topic-BERT an essential pretraining step to embed topic information into BERT with self-supervised learning.Experimental results on the DSTC-8 Ubuntu IRC dataset show state-of-the-art results in response selection and topic disentanglement tasks outperforming existing methods by a good margin.1 Weishi Wang, Steven C. H. Hoi, Shafiq R. Joty |
EMNLP (1) | 3 |
| 2020 | VD-BERT: A Unified Vision and Dialog Transformer with BERTabstractVisual dialog is a challenging vision-language task, where a dialog agent needs to answer a series of questions through reasoning on the image content and dialog history.Prior work has mostly focused on various attention mechanisms to model such intricate interactions.By contrast, in this work, we propose VD-BERT, a simple yet effective framework of unified vision-dialog Transformer that leverages the pretrained BERT language models for Visual Dialog tasks.The model is unified in that (1) it captures all the interactions between the image and the multi-turn dialog using a single-stream Transformer encoder, and (2) it supports both answer ranking and answer generation seamlessly through the same architecture.More crucially, we adapt BERT for the effective fusion of vision and dialog contents via visually grounded training.Without the need of pretraining on external vision-language data, our model yields new state of the art, achieving the top position in both single-model and ensemble settings (74.54 and 75.35NDCG scores) on the visual dialog leaderboard.Our code and pretrained models are released at https: //github.com/salesforce/VD-BERT. Yue Wang 0034, Shafiq R. Joty, Michael R. Lyu, Irwin King, Caiming Xiong, Steven C. H. Hoi |
EMNLP (1) | 2 |
| 2020 | Online Conversation Disentanglement with Pointer NetworksabstractHuge amounts of textual conversations occur online every day, where multiple conversations take place concurrently.Interleaved conversations lead to difficulties in not only following the ongoing discussions but also extracting relevant information from simultaneous messages.Conversation disentanglement aims to separate intermingled messages into detached conversations.However, existing disentanglement methods rely mostly on handcrafted features that are dataset specific, which hinders generalization and adaptability.In this work, we propose an end-to-end online framework for conversation disentanglement that avoids time-consuming domain-specific feature engineering.We design a novel way to embed the whole utterance that comprises timestamp, speaker, and message text, and propose a custom attention mechanism that models disentanglement as a pointing problem while effectively capturing inter-utterance interactions in an end-to-end fashion.We also introduce a joint-learning objective to better capture contextual information.Our experiments on the Ubuntu IRC dataset show that our method achieves state-of-the-art performance in both link and conversation prediction tasks. Tao Yu 0009, Shafiq R. Joty |
EMNLP (1) | 2 |
| 2020 | Tree-Structured Attention with Hierarchical Accumulation
Xuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard Socher |
ICLR | 2 |
| 2020 | Speech Transformer with Speaker Aware Persistent Memory
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 4 |
| 2020 | Universal Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 4 |
| 2020 | Cross Attention with Monotonic Alignment for Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 4 |
| 2020 | Self-Supervised Relationship ProbingabstractStructured representations of images that model visual relationships are beneficial for many vision and vision-language applications. However, current human-annotated visual relationship datasets suffer from the long-tailed predicate distribution problem which limits the potential of visual relationship models. In this work, we introduce a self-supervised method that implicitly learns the visual relationships without relying on any ground-truth visual relationship annotations. Our method relies on 1) intra- and inter-modality encodings to respectively model relationships within each modality separately and jointly, and 2) relationship probing, which seeks to discover the graph structure within each modality. By leveraging masked language modeling, contrastive learning, and dependency tree distances for self-supervision, our method learns better object features as well as implicit visual relationships. We verify the effectiveness of our proposed method on various vision-language tasks that benefit from improved visual relationship understanding. Jiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai 0001, Vlad I. Morariu, Handong Zhao, Tong Sun 0005 |
NeurIPS | 3 |
| 2020 | Data Diversification: A Simple Strategy For Neural Machine TranslationabstractWe introduce Data Diversification: a simple but effective strategy to boost neural machine translation (NMT) performance. It diversifies the training data by using the predictions of multiple forward and backward models and then merging them with the original dataset on which the final NMT model is trained. Our method is applicable to all NMT models. It does not require extra monolingual data like back-translation, nor does it add more computations and parameters like ensembles of models. Our method achieves state-of-the-art BLEU scores of 30.7 and 43.7 in the WMT'14 English-German and English-French translation tasks, respectively. It also substantially improves on 8 other translation tasks: 4 IWSLT tasks (English-German and English-French) and 4 low-resource translation tasks (English-Nepali and English-Sinhala). We demonstrate that our method is more effective than knowledge distillation and dual learning, it exhibits strong correlation with ensembles of models, and it trades perplexity off for better BLEU score. Xuan-Phi Nguyen, Shafiq R. Joty, Kui Wu 0004, AiTi Aw |
NeurIPS | 2 |
| 2020 | Unsupervised Word Translation with Adversarial AutoencoderabstractCrosslingual word embeddings learned from monolingual embeddings have a crucial role in many downstream tasks, ranging from machine translation to transfer learning. Adversarial training has shown impressive success in learning crosslingual embeddings and the associated word translation task without any parallel data by mapping monolingual embeddings to a shared space. However, recent work has shown superior performance for non-adversarial methods in more challenging language pairs. In this article, we investigate adversarial autoencoder for unsupervised word translation and propose two novel extensions to it that yield more stable training and improved results. Our method includes regularization terms to enforce cycle consistency and input reconstruction, and puts the target encoders as an adversary against the corresponding discriminator. We use two types of refinement procedures sequentially after obtaining the trained encoders and mappings from the adversarial training, namely, refinement with Procrustes solution and refinement with symmetric re-weighting. Extensive experimentations with high- and low-resource languages from two different data sets show that our method achieves better performance than existing adversarial and non-adversarial approaches and is also competitive with the supervised system. Along with performing comprehensive ablation studies to understand the contribution of different components of our adversarial model, we also conduct a thorough analysis of the refinement procedures to understand their effects. Tasnim Mohiuddin, Shafiq R. Joty |
Comput. Linguistics | 2 |
| 2020 | Video captioning with boundary-aware hierarchical language decoding and joint video prediction
Xiangxi Shi, Jianfei Cai 0001, Jiuxiang Gu, Shafiq R. Joty |
Neurocomputing | 4 |
| 2020 | An Attention-based Rumor Detection Model with Tree-structured Recursive Neural NetworksabstractRumor spread in social media severely jeopardizes the credibility of online content. Thus, automatic debunking of rumors is of great importance to keep social media a healthy environment. While facing a dubious claim, people often dispute its truthfulness sporadically in their posts containing various cues, which can form useful evidence with long-distance dependencies. In this work, we propose to learn discriminative features from microblog posts by following their non-sequential propagation structure and generate more powerful representations for identifying rumors. For modeling non-sequential structure, we first represent the diffusion of microblog posts with propagation trees, which provide valuable clues on how a claim in the original post is transmitted and developed over time. We then present a bottom-up and a top-down tree-structured models based on Recursive Neural Networks (RvNN) for rumor representation learning and classification, which naturally conform to the message propagation process in microblogs. To enhance the rumor representation learning, we reveal that effective rumor detection is highly related to finding evidential posts, e.g., the posts expressing specific attitude towards the veracity of a claim, as an extension of the previous RvNN-based detection models that treat every post equally. For this reason, we design discriminative attention mechanisms for the RvNN-based models to selectively attend on the subset of evidential posts during the bottom-up/top-down recursive composition. Experimental results on four datasets collected from real-world microblog platforms confirm that (1) our RvNN-based models achieve much better rumor detection and classification performance than state-of-the-art approaches; (2) the attention mechanisms for focusing on evidential posts can further improve the performance of our RvNN-based method; and (3) our approach possesses superior capacity on detecting rumors at a very early stage. Jing Ma 0004, Wei Gao 0001, Shafiq R. Joty, Kam-Fai Wong |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2019 | Adversarial Unsupervised Representation Learning for Activity Time-SeriesabstractSufficient physical activity and restful sleep play a major role in the prevention and cure of many chronic conditions. Being able to proactively screen and monitor such chronic conditions would be a big step forward for overall health. The rapid increase in the popularity of wearable devices pro-vides a significant new source, making it possible to track the user’s lifestyle real-time. In this paper, we propose a novel unsupervised representation learning technique called activ-ity2vecthat learns and “summarizes” the discrete-valued ac-tivity time-series. It learns the representations with three com-ponents: (i) the co-occurrence and magnitude of the activ-ity levels in a time-segment, (ii) neighboring context of the time-segment, and (iii) promoting subject-invariance with ad-versarial training. We evaluate our method on four disorder prediction tasks using linear classifiers. Empirical evaluation demonstrates that our proposed method scales and performs better than many strong baselines. The adversarial regime helps improve the generalizability of our representations by promoting subject invariant features. We also show that using the representations at the level of a day works the best since human activity is structured in terms of daily routines. Karan Aggarwal, Shafiq R. Joty, Luis Fernández-Luque, Jaideep Srivastava |
AAAI | 2 |
| 2019 | A Unified Linear-Time Framework for Sentence-Level Discourse ParsingabstractWe propose an efficient neural framework for sentence-level discourse analysis in accordance with Rhetorical Structure Theory (RST). Our framework comprises a discourse segmenter to identify the elementary discourse units (EDU) in a text, and a discourse parser that constructs a discourse tree in a top-down fashion. Both the segmenter and the parser are based on Pointer Networks and operate in linear time. Our segmenter yields an F1 score of 95.4%, and our parser achieves an F1 score of 81.7% on the aggregated labeled (relation) metric, surpassing previous approaches by a good margin and approaching human agreement on both tasks (98.3 and 83.0 F1). Shafiq R. Joty, Prathyusha Jwalapuram, Saiful Bari |
ACL (1) | 2 |
| 2019 | Sentence-Level Evidence Embedding for Claim Verification with Hierarchical Attention NetworksabstractClaim verification is generally a task of verifying the veracity of a given claim, which is critical to many downstream applications.It is cumbersome and inefficient for human fact-checkers to find consistent pieces of evidence, from which solid verdict could be inferred against the claim.In this paper, we propose a novel end-to-end hierarchical attention network focusing on learning to represent coherent evidence as well as their semantic relatedness with the claim.Our model consists of three main components: 1) A coherence-based attention layer embeds coherent evidence considering the claim and sentences from relevant articles; 2) An entailment-based attention layer attends on sentences that can semantically infer the claim on top of the first attention; and 3) An output layer predicts the verdict based on the embedded evidence.Experimental results on three public benchmark datasets show that our proposed model outperforms a set of state-of-the-art baselines. Jing Ma 0004, Wei Gao 0001, Shafiq R. Joty, Kam-Fai Wong |
ACL (1) | 3 |
| 2019 | Evaluating Pronominal Anaphora in Machine Translation: An Evaluation Measure and a Test SuiteabstractPrathyusha Jwalapuram, Shafiq Joty, Irina Temnikova, Preslav Nakov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Prathyusha Jwalapuram, Shafiq R. Joty, Irina P. Temnikova, Preslav Nakov |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Using Clinical Notes with Time Series Data for ICU ManagementabstractSwaraj Khadanga, Karan Aggarwal, Shafiq Joty, Jaideep Srivastava. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Swaraj Khadanga, Karan Aggarwal, Shafiq R. Joty, Jaideep Srivastava |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Hierarchical Pointer Net ParsingabstractLinlin Liu, Xiang Lin, Shafiq Joty, Simeng Han, Lidong Bing. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shafiq R. Joty, Simeng Han, Lidong Bing |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Unified Neural Coherence ModelabstractHan Cheol Moon, Tasnim Mohiuddin, Shafiq Joty, Chi Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Han Cheol Moon, Tasnim Mohiuddin, Shafiq R. Joty, Xu Chi |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from SpeechabstractWe present the Factorial Deep Markov Model (FDMM) for representation learning of speech. The FDMM learns disentangled, interpretable and lower dimensional latent representations from speech without supervision. We use a static and dynamic latent variable to exploit the fact that information in a speech signal evolves at different time scales. Latent representations learned by the FDMM outperform a baseline i-vector system on speaker verification and dialect identification while also reducing the error rate of a phone recognition system in a domain mismatch scenario. Sameer Khurana, Shafiq R. Joty, Ahmed Ali 0002, James R. Glass |
ICASSP | 2 |
| 2019 | Unpaired Image Captioning via Scene Graph AlignmentsabstractMost of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image scene graph generator, a sentence scene graph generator, a scene graph encoder, and a sentence decoder. Specifically, we first train the scene graph encoder and the sentence decoder on the text modality. To align the scene graphs between images and sentences, we propose an unsupervised feature alignment method that maps the scene graph features from the image to the sentence modality. Experimental results show that our proposed model can generate quite promising results without using any image-caption training pairs, outperforming existing methods by a wide margin. Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Handong Zhao, Xu Yang 0021, Gang Wang 0012 |
ICCV | 2 |
| 2019 | Watch It Twice: Video Captioning with a Refocused Video EncoderabstractWith the rapid growth of video data and the increasing demands of various crossmodal applications such as intelligent video search and assistance towards visually-impaired people, video captioning task has received a lot of attention recently in computer vision and natural language processing fields. The state-of-the-art video captioning methods focus more on encoding the temporal information, while lacking effective ways to remove irrelevant temporal information and also neglecting the spatial details. In particular, the current unidirectional video encoder can be negatively affected by irrelevant temporal information, especially the irrelevant information at the beginning and at the end of a video. In addition, disregarding detailed spatial features may lead to incorrect word choices in decoding. In this paper, we propose a novel recurrent video encoding method and a novel visual spatial feature for the video captioning task. The recurrent encoding module encodes the video twice with a predicted key frame to avoid irrelevant temporal information often occurring at the beginning and at the end of a video. The novel spatial features represent spatial information from different regions of a video and provide the decoder with more detailed information. Experiments on two benchmark datasets show superior performance of the proposed method. Xiangxi Shi, Jianfei Cai 0001, Shafiq R. Joty, Jiuxiang Gu |
ACM Multimedia | 3 |
| 2019 | NEURON: Query Execution Plan Meets Natural Language Processing For Augmenting DB EducationabstractA core component of a database systems course at the undergraduate level is the design and implementation of the query optimizer in an rdbms. The query optimization process produces aquery execution plan (qep ), which represents an execution strategy for an sql query. Unfortunately, in practice, it is often difficult for a student to comprehend a query execution strategy by perusing its qep, hindering her learning process. In this demonstration, we present a novel system called neuron that facilitates natural language interaction with qep s to enhance its understanding. neuron accepts an sql query (which may include joins, aggregation, nesting, among other things) as input, executes it, and generates a simplified natural language description (both in text and voice form) of the execution strategy deployed by the underlying rdbms. Furthermore, it facilitates understanding of various features related to a qep through anatural language question answering (nlqa ) framework. We advocate that such tool, world's first of its kind, can greatly enhance students' learning of the query optimization topic. Sourav S. Bhowmick, Wanlu Zhang, Wanyi Huang, Shafiq R. Joty |
SIGMOD Conference | 6 |
| 2018 | Domain Adaptation with Adversarial Training and Graph EmbeddingsabstractThe success of deep neural networks (DNNs) is heavily dependent on the availability of labeled data.However, obtaining labeled data is a big challenge in many real-world problems.In such scenarios, a DNN model can leverage labeled and unlabeled data from a related domain, but it has to deal with the shift in data distributions between the source and the target domains.In this paper, we study the problem of classifying social media posts during a crisis event (e.g., Earthquake).For that, we use labeled and unlabeled data from past similar events (e.g., Flood) and unlabeled data for the current event.We propose a novel model that performs adversarial learning based domain adaptation to deal with distribution drifts and graph based semi-supervised learning to leverage unlabeled data within a single unified deep learning framework.Our experiments with two real-world crisis datasets collected from Twitter demonstrate significant improvements over several baselines. Firoj Alam, Shafiq R. Joty, Muhammad Imran 0002 |
ACL (1) | 2 |
| 2018 | Coherence Modeling of Asynchronous Conversations: A Neural Entity Grid ApproachabstractWe propose a novel coherence model for written asynchronous conversations (e.g., forums, emails), and show its applications in coherence assessment and thread reconstruction tasks.We conduct our research in two steps.First, we propose improvements to the recently proposed neural entity grid model by lexicalizing its entity transitions.Then, we extend the model to asynchronous conversations by incorporating the underlying conversational structure in the entity grid representation and feature computation.Our model achieves state of the art results on standard coherence assessment tasks in monologue and conversations outperforming existing models.We also demonstrate its effectiveness in reconstructing thread structures. *All authors contributed equally.s0: LDI Corp., Cleveland, said it will offer $50 million in commercial paper backed by leaserental receivables.s1: Tasnim Mohiuddin, Shafiq R. Joty, Tien Dat Nguyen |
ACL (1) | 2 |
| 2018 | A Structured Learning Approach with Neural Conditional Random Fields for Sleep StagingabstractSleep plays a vital role in human health, both mental and physical. Sleep disorders like sleep apnea are increasing in prevalence, with the rapid increase in factors like obesity. Sleep apnea is most commonly treated with Continuous Positive Air Pressure (CPAP) therapy. Presently, however, there is no mechanism to monitor a patient's progress with CPAP. Accurate detection of sleep stages from CPAP flow signal is crucial for such a mechanism. We propose, for the first time, an automated sleep staging model based only on the flow signal. Deep neural networks have recently shown high accuracy on sleep staging by eliminating handcrafted features. However, these methods focus exclusively on extracting informative features from the input signal, without paying much attention to the dynamics of sleep stages in the output sequence. We propose an end-to-end framework that uses a combination of deep convolution and recurrent neural networks to extract high-level features from raw flow signal with a structured output layer based on a conditional random field to model the temporal transition structure of the sleep stages. We improve upon the previous methods by 10% using our model, that can be augmented to the previous sleep staging deep learning methods. We also show that our method can be used to accurately track sleep metrics like sleep efficiency calculated from sleep stages that can be deployed for monitoring the response of CPAP therapy on sleep apnea patients. Apart from the technical contributions, we expect this study to motivate new research questions in sleep science. Karan Aggarwal, Swaraj Khadanga, Shafiq R. Joty, Louis Kazaglis, Jaideep Srivastava |
IEEE BigData | 3 |
| 2018 | ANR: Aspect-based Neural RecommenderabstractTextual reviews, which are readily available on many e-commerce and review websites such as Amazon and Yelp, serve as an invaluable source of information for recommender systems. However, not all parts of the reviews are equally important, and the same choice of words may reflect a different meaning based on its context. In this paper, we propose a novel end-to-end Aspect-based Neural Recommender (ANR) to perform aspect-based representation learning for both users and items via an attention-based component. Furthermore, we model the multi-faceted process behind how users rate items by estimating the aspect-level user and item importance by adapting the neural co-attention mechanism. Our proposed model concurrently address several shortcomings of existing recommender systems, and a thorough experimental study on 25 benchmark datasets from Amazon and Yelp shows that ANR significantly outperforms recently proposed state-of-the-art baselines such as DeepCoNN, D-Attn and ALFM. Jin Yao Chin, Kaiqi Zhao 0001, Shafiq R. Joty, Gao Cong |
CIKM | 3 |
| 2018 | Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative ModelsabstractTextual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset. Jiuxiang Gu, Jianfei Cai 0001, Shafiq R. Joty, Li Niu 0002, Gang Wang 0012 |
CVPR | 3 |
| 2018 | Unpaired Image Captioning by Language Pivoting
Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Gang Wang 0012 |
ECCV (1) | 2 |
| 2018 | VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
Qing Li 0003, Qingyi Tao, Shafiq R. Joty, Jianfei Cai 0001, Jiebo Luo 0001 |
ECCV (7) | 3 |
| 2018 | Joint Multitask Learning for Community Question Answering Using Task-Specific EmbeddingsabstractWe address jointly two important tasks for Question Answering in community forums: given a new question, (i) find related existing questions, and (ii) find relevant answers to this new question.We further use an auxiliary task to complement the previous two, i.e., (iii) find good answers with respect to the thread question in a question-comment thread.We use deep neural networks (DNNs) to learn meaningful task-specific embeddings, which we then incorporate into a conditional random field (CRF) model for the multitask setting, performing joint learning over a complex graph structure.While DNNs alone achieve competitive results when trained to produce the embeddings, the CRF, which makes use of the embeddings and the dependencies between the tasks, improves the results significantly and consistently across a variety of evaluation metrics, thus showing the complementarity of DNNs and structured learning. Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
EMNLP | 1 |
| 2018 | Discourse Processing and Its Applications in Text MiningabstractDiscourse processing is a suite of Natural Language Processing (NLP) tasks to uncover linguistic structures from texts at several levels, which can support many text mining applications. This involves identifying the topic structure, the coherence structure, the coreference structure, and the conversation structure for conversational discourse. Taken together, these structures can inform text summarization, essay scoring, sentiment analysis, machine translation, information extraction, question answering, and thread recovery. The tutorial starts with an overview of basic concepts in discourse analysis - monologue vs. conversation, synchronous vs. asynchronous conversation, and key linguistic structures in discourse analysis. It then covers traditional machine learning methods along with the most recent works using deep learning, and compare their performances on benchmark datasets. For each discourse structure we describe, we show its applications in downstream text mining tasks. Methods and metrics for evaluation are discussed in detail. We conclude the tutorial with an interactive discussion of future challenges and opportunities. Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng, Gabriel Murray |
ICDM | 1 |
| 2018 | Graph Based Semi-Supervised Learning with Convolution Neural Networks to Classify Crisis Related Tweets
Firoj Alam, Shafiq R. Joty, Muhammad Imran 0002 |
ICWSM | 2 |
| 2018 | SegBot: A Generic Neural Text Segmentation Model with Pointer NetworkabstractText segmentation is a fundamental task in natural language processing that comes in two levels of granularity: (i) segmenting a document into a sequence of topical segments (topic segmentation), and (ii) segmenting a sentence into a sequence of elementary discourse units (EDU segmentation). Traditional solutions to the two tasks heavily rely on carefully designed features. The recently proposed neural models do not need manual feature engineering, but they either suffer from sparse boundary tags or they cannot well handle the issue of variable size output vocabulary. We propose a generic end-to-end segmentation model called SegBot. SegBot uses a bidirectional recurrent neural network to encode input text sequence. The model then uses another recurrent neural network together with a pointer network to select text boundaries in the input sequence. In this way, SegBot does not require hand-crafted features. More importantly, our model inherently handles the issue of variable size output vocabulary and the issue of sparse boundary tags. In our experiments, SegBot outperforms state-of-the-art models on both topic and EDU segmentation tasks. Jing Li 0034, Aixin Sun, Shafiq R. Joty |
IJCAI | 3 |
| 2018 | Modeling Speech Acts in Asynchronous Conversations: A Neural-CRF ApproachabstractParticipants in an asynchronous conversation (e.g., forum, e-mail) interact with each other at different times, performing certain communicative acts, called speech acts (e.g., question, request). In this article, we propose a hybrid approach to speech act recognition in asynchronous conversations. Our approach works in two main steps: a long short-term memory recurrent neural network (LSTM-RNN) first encodes each sentence separately into a task-specific distributed representation, and this is then used in a conditional random field (CRF) model to capture the conversational dependencies between sentences. The LSTM-RNN model uses pretrained word embeddings learned from a large conversational corpus and is trained to classify sentences into speech act types. The CRF model can consider arbitrary graph structures to model conversational dependencies in an asynchronous conversation. In addition, to mitigate the problem of limited annotated data in the asynchronous domains, we adapt the LSTM-RNN model to learn from synchronous conversations (e.g., meetings), using domain adversarial training of neural networks. Empirical evaluation shows the effectiveness of our approach over existing ones: (i) LSTM-RNNs provide better task-specific representations, (ii) conversational word embeddings benefit the LSTM-RNNs more than the off-the-shelf ones, (iii) adversarial training gives better domain-invariant representations, and (iv) the global CRF model improves over local models. Shafiq R. Joty, Tasnim Mohiuddin |
Comput. Linguistics | 1 |
| 2018 | Distributed Representations of Tuples for Entity ResolutionabstractDespite the efforts in 70+ years in all aspects of entity resolution (ER), there is still a high demand for democratizing ER - by reducing the heavy human involvement in labeling data, performing feature engineering, tuning parameters, and defining blocking functions. With the recent advances in deep learning, in particular distributed representations of words ( a.k.a . word embeddings), we present a novel ER system, called D eep ER, that achieves good accuracy, high efficiency, as well as ease-of-use ( i.e ., much less human efforts). We use sophisticated composition methods, namely uni- and bi-directional recurrent neural networks (RNNs) with long short term memory (LSTM) hidden units, to convert each tuple to a distributed representation ( i.e ., a vector), which can in turn be used to effectively capture similarities between tuples. We consider both the case where pre-trained word embeddings are available as well the case where they are not; we present ways to learn and tune the distributed representations that are customized for a specific ER task under different scenarios. We propose a locality sensitive hashing (LSH) based blocking approach that takes all attributes of a tuple into consideration and produces much smaller blocks, compared with traditional methods that consider only a few attributes. We evaluate our algorithms on multiple datasets (including benchmarks, biomedical data, as well as multi-lingual data) and the extensive experimental results show that D eep ER outperforms existing solutions. Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq R. Joty, Mourad Ouzzani, Nan Tang 0001 |
Proc. VLDB Endow. | 3 |
| 2017 | A Neural Local Coherence ModelabstractWe propose a local coherence model based on a convolutional neural network that operates over the entity grid representation of a text.The model captures long range entity transitions along with entity-specific features without loosing generalization, thanks to the power of distributed representation.We present a pairwise ranking method to train the model in an end-to-end fashion on a task and learn task-specific high level features.Our evaluation on three different coherence assessment tasks demonstrates that our model achieves state of the art results outperforming existing models by a good margin. Tien Dat Nguyen, Shafiq R. Joty |
ACL (1) | 2 |
| 2017 | Regularized and Retrofitted models for Learning Sentence Representation with ContextabstractVector representation of sentences is important for many text processing tasks that involve classifying, clustering, or ranking sentences. For solving these tasks, bag-of-word based representation has been used for a long time. In recent years, distributed representation of sentences learned by neural models from unlabeled data has been shown to outperform traditional bag-of-words representations. However, most existing methods belonging to the neural models consider only the content of a sentence, and disregard its relations with other sentences in the context. In this paper, we first characterize two types of contexts depending on their scope and utility. We then propose two approaches to incorporate contextual information into content-based models. We evaluate our sentence representation models in a setup, where context is available to infer sentence vectors. Experimental results demonstrate that our proposed models outshine existing models on three fundamental tasks, such as, classifying, clustering, and ranking sentences. Tanay Kumar Saha, Shafiq R. Joty, Naeemul Hassan, Mohammad Al Hasan |
CIKM | 2 |
| 2017 | Cross-language Learning with Adversarial Neural NetworksabstractWe address the problem of cross-language adaptation for question-question similarity reranking in community question answering, with the objective to port a system trained on one input language to another input language given labeled training data for the first language and only unlabeled data for the second language.In particular, we propose to use adversarial training of neural networks to learn high-level features that are discriminative for the main learning task, and at the same time are invariant across the input languages.The evaluation results show sizable improvements for our cross-language adversarial neural network (CLANN) model over a strong nonadversarial system. Shafiq R. Joty, Preslav Nakov, Lluís Màrquez, Israa Jaradat |
CoNLL | 1 |
| 2017 | Robust Classification of Crisis-Related Data on Social Networks Using Convolutional Neural Networks
Tien Dat Nguyen, Kamla Al-Mannai, Shafiq R. Joty, Hassan Sajjad 0001, Muhammad Imran 0002, Prasenjit Mitra 0001 |
ICWSM | 3 |
| 2017 | CQAVis: Visual Text Analytics for Community Question AnsweringabstractCommunity question answering (CQA) forums can provide effective means for sharing information and addressing a user's information needs about particular topics. However, many such online forums are not moderated, resulting in many low quality and redundant comments, which makes it very challenging for users to find the appropriate answers to their questions. In this paper, we apply a user-centered design approach to develop a system, CQAVis, which supports users in identifying high quality comments and get their questions answered. Informed by the user's requirements, the system combines both text analytics and interactive visualization techniques together in a synergistic way. Given a new question posed by the user, the text analytic module automatically finds relevant answers by exploring existing related questions and the comments within their threads. Then the visualization module presents the search results to the user and supports the exploration of related comments. We have evaluated the system in the wild by deploying it within a CQA forum among thousands of real users. Through the online study, we gained deeper insights about the potential utility of the system, as well as learned generalizable lessons for designing visual text analytics systems for the domain of CQA forums. Enamul Hoque Prince, Shafiq R. Joty, Lluís Màrquez, Giuseppe Carenini |
IUI | 2 |
| 2017 | Con-S2V: A Generic Framework for Incorporating Extra-Sentential Context into Sen2Vec
Tanay Kumar Saha, Shafiq R. Joty, Mohammad Al Hasan |
ECML/PKDD (1) | 2 |
| 2017 | Cross-Language Question Re-RankingabstractWe study how to find relevant questions in community forums when the language of the new questions is different from that of the existing questions in the forum. In particular, we explore the Arabic-English language pair. We compare a kernel-based system with a feed-forward neural network in a scenario where a large parallel corpus is available for training a machine translation system, bilingual dictionaries, and cross-language word embeddings. We observe that both approaches degrade the performance of the system when working on the translated text, especially the kernel-based system, which depends heavily on a syntactic kernel. We address this issue using a cross-language tree kernel, which compares the original Arabic tree to the English trees of the related questions. We show that this kernel almost closes the performance gap with respect to the monolingual system. On the neural network side, we use the parallel corpus to train cross-language embeddings, which we then use to represent the Arabic input and the English related questions in the same space. The results also improve to close to those of the monolingual neural network. Overall, the kernel system shows a better performance compared to the neural network in all cases. Giovanni Da San Martino, Salvatore Romeo, Alberto Barrón-Cedeño, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
SIGIR | 4 |
| 2017 | Discourse Structure in Machine Translation EvaluationabstractIn this article, we explore the potential of using sentence-level discourse structure for machine translation evaluation. We first design discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory (RST). Then, we show that a simple linear combination with these measures can help improve various existing machine translation evaluation metrics regarding correlation with human judgments both at the segment level and at the system level. This suggests that discourse information is complementary to the information used by many of the existing evaluation metrics, and thus it could be taken into account when developing richer evaluation metrics, such as the WMT-14 winning combined metric DiscoTKparty. We also provide a detailed analysis of the relevance of various discourse elements and relations from the RST parse trees for machine translation evaluation. In particular, we show that (i) all aspects of the RST tree are relevant, (ii) nuclearity is more useful than relation type, and (iii) the similarity of the translation RST tree to the reference RST tree is positively correlated with translation quality. Shafiq R. Joty, Francisco Guzmán, Lluís Màrquez, Preslav Nakov |
Comput. Linguistics | 1 |
| 2017 | Machine translation evaluation with neural networks
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
Comput. Speech Lang. | 2 |
| 2017 | Domain adaptation using neural network joint model
Shafiq R. Joty, Nadir Durrani, Hassan Sajjad 0001, Ahmed Abdelali |
Comput. Speech Lang. | 1 |
| 2016 | Speech Act Modeling of Written Asynchronous Conversations with Task-Specific Embeddings and Conditional Structured ModelsabstractThis paper addresses the problem of speech act recognition in written asynchronous conversations (e.g., fora, emails).We propose a class of conditional structured models defined over arbitrary graph structures to capture the conversational dependencies between sentences.Our models use sentence representations encoded by a long short term memory (LSTM) recurrent neural model.Empirical evaluation shows the effectiveness of our approach over existing ones: (i) LSTMs provide better task-specific representations, and (ii) the global joint model improves over local models. Shafiq R. Joty, Enamul Hoque Prince |
ACL (1) | 1 |
| 2016 | A Deep Fusion Model for Domain Adaptation in Phrase-based MTabstractWe present a novel fusion model for domain adaptation in Statistical Machine Translation. Our model is based on the joint source-target neural network Devlin et al., 2014, and is learned by fusing in- and out-domain models. The adaptation is performed by backpropagating errors from the output layer to the word embedding layer of each model, subsequently adjusting parameters of the composite model towards the in-domain data. On the standard tasks of translating English-to-German and Arabic-to-English TED talks, we observed average improvements of +0.9 and +0.7 BLEU points, respectively over a competition grade phrase-based system. We also demonstrate improvements over existing adaptation methods. Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali |
COLING | 3 |
| 2016 | Joint Learning with Global Inference for Comment Classification in Community Question AnsweringabstractThis paper addresses the problem of comment classification in community Question Answering.Following the state of the art, we approach the task with a global inference process to exploit the information of all comments in the answer-thread in the form of a fully connected graph.Our contribution comprises two novel joint learning models that are on-line and integrate inference within learning.The first one jointly learns two node-and edge-level MaxEnt classifiers with stochastic gradient descent and integrates the inference step with loopy belief propagation.The second model is an instance of fully connected pairwise CRFs (FCCRF).The FCCRF model significantly outperforms all other approaches and yields the best results on the task to date.Crucial elements for its success are the global normalization and an Ising-like edge potential. Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
HLT-NAACL | 1 |
| 2015 | Pairwise Neural Machine Translation EvaluationabstractFrancisco Guzmán, Shafiq Joty, Lluís Màrquez, Preslav Nakov. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
ACL (1) | 2 |
| 2015 | Global Thread-level Inference for Comment Classification in Community Question AnsweringabstractShafiq Joty, Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Shafiq R. Joty, Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
EMNLP | 1 |
| 2015 | How to Avoid Unwanted Pregnancies: Domain Adaptation using Neural Network ModelsabstractWe present novel models for domain adaptation based on the neural network joint model (NNJM).Our models maximize the cross entropy by regularizing the loss function with respect to in-domain model.Domain adaptation is carried out by assigning higher weight to out-domain sequences that are similar to the in-domain data.In our alternative model we take a more restrictive approach by additionally penalizing sequences similar to the outdomain data.Our models achieve better perplexities than the baseline NNJM models and give improvements of up to 0.5 and 0.6 BLEU points in Arabic-to-English and English-to-German language pairs, on a standard task of translating TED talks. Shafiq R. Joty, Hassan Sajjad 0001, Nadir Durrani, Kamla Al-Mannai, Ahmed Abdelali, Stephan Vogel |
EMNLP | 1 |
| 2015 | Fine-grained Opinion Mining with Recurrent Neural Networks and Word EmbeddingsabstractThe tasks in fine-grained opinion mining can be regarded as either a token-level sequence labeling problem or as a semantic compositional task.We propose a general class of discriminative models based on recurrent neural networks (RNNs) and word embeddings that can be successfully applied to such tasks without any taskspecific feature engineering effort.Our experimental results on the task of opinion target identification show that RNNs, without using any hand-crafted features, outperform feature-rich CRF-based models.Our framework is flexible, allows us to incorporate other linguistic features, and achieves results that rival the top performing systems in SemEval-2014. Pengfei Liu 0004, Shafiq R. Joty, Helen M. Meng |
EMNLP | 2 |
| 2015 | Using joint models or domain adaptation in statistical machine translation
Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali, Stephan Vogel |
MTSummit | 3 |
| 2015 | CODRA: A Novel Discriminative Framework for Rhetorical AnalysisabstractClauses and sentences rarely stand on their own in an actual discourse; rather, the relationship between them carries important information that allows the discourse to express a meaning as a whole beyond the sum of its individual parts. Rhetorical analysis seeks to uncover this coherence structure. In this article, we present CODRA— a COmplete probabilistic Discriminative framework for performing Rhetorical Analysis in accordance with Rhetorical Structure Theory, which posits a tree representation of a discourse. CODRA comprises a discourse segmenter and a discourse parser. First, the discourse segmenter, which is based on a binary classifier, identifies the elementary discourse units in a given text. Then the discourse parser builds a discourse tree by applying an optimal parsing algorithm to probabilities inferred from two Conditional Random Fields: one for intra-sentential parsing and the other for multi-sentential parsing. We present two approaches to combine these two stages of parsing effectively. By conducting a series of empirical evaluations over two different data sets, we demonstrate that CODRA significantly outperforms the state-of-the-art, often by a wide margin. We also show that a reranking of the k-best parse hypotheses generated by CODRA can potentially improve the accuracy even further. Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng |
Comput. Linguistics | 1 |
| 2014 | Using Discourse Structure Improves Machine Translation EvaluationabstractWe present experiments in using discourse structure for improving machine translation evaluation.We first design two discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory.Then, we show that these measures can help improve a number of existing machine translation evaluation metrics both at the segment-and at the system-level.Rather than proposing a single new metric, we show that discourse information is complementary to the state-of-the-art evaluation metrics, and thus should be taken into account in the development of future richer evaluation metrics. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
ACL (1) | 2 |
| 2014 | A Study of using Syntactic and Semantic Structures for Concept Segmentation and Labeling
Iman Saleh 0001, D. Scott Cyphers, James R. Glass, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
COLING | 4 |
| 2014 | Learning to Differentiate Better from Worse TranslationsabstractFrancisco Guzmán, Shafiq Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia |
EMNLP | 2 |
| 2014 | Discriminative Reranking of Discourse Parses Using Tree KernelsabstractIn this paper, we present a discrimina-tive approach for reranking discourse trees generated by an existing probabilistic dis-course parser. The reranker relies on tree kernels (TKs) to capture the global depen-dencies between discourse units in a tree. In particular, we design new computa-tional structures of discourse trees, which combined with standard TKs, originate novel discourse TKs. The empirical evalu-ation shows that our reranker can improve the state-of-the-art sentence-level parsing accuracy from 79.77 % to 82.15%, a rel-ative error reduction of 11.8%, which in turn pushes the state-of-the-art document-level accuracy from 55.8 % to 57.3%. 1 Shafiq R. Joty, Alessandro Moschitti |
EMNLP | 1 |
| 2014 | Semantic Kernels for Semantic ParsingabstractWe present an empirical study on the use of semantic information for Concept Seg-mentation and Labeling (CSL), which is an important step for semantic parsing. We represent the alternative analyses out-put by a state-of-the-art CSL parser with tree structures, which we rerank with a classifier trained on two types of seman-tic tree kernels: one processing structures built with words, concepts and Brown clusters, and another one using semantic similarity among the words composing the structure. The results on a corpus from the restaurant domain show that our semantic kernels exploiting similarity measures out-perform state-of-the-art rerankers. 1 Iman Saleh 0001, Alessandro Moschitti, Preslav Nakov, Lluís Màrquez, Shafiq R. Joty |
EMNLP | 5 |
| 2013 | Combining Intra- and Multi-sentential Rhetorical Parsing for Document-level Discourse Analysis
Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng, Yashar Mehdad |
ACL (1) | 1 |
| 2013 | Towards Topic Labeling with Phrase Entailment and Aggregation
Yashar Mehdad, Giuseppe Carenini, Raymond T. Ng, Shafiq R. Joty |
HLT-NAACL | 4 |
| 2013 | Dialogue Act Recognition in Synchronous and Asynchronous Conversations
Maryam Tavafi, Yashar Mehdad, Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng |
SIGDIAL Conference | 3 |
| 2013 | Topic Segmentation and Labeling in Asynchronous ConversationsabstractTopic segmentation and labeling is often considered a prerequisite for higher-level conversation analysis and has been shown to be useful in many Natural Language Processing (NLP) applications. We present two new corpora of email and blog conversations annotated with topics, and evaluate annotator reliability for the segmentation and labeling tasks in these asynchronous conversations. We propose a complete computational framework for topic segmentation and labeling in asynchronous conversations. Our approach extends state-of-the-art methods by considering a fine-grained structure of an asynchronous conversation, along with other conversational features by applying recent graph-based methods for NLP. For topic segmentation, we propose two novel unsupervised models that exploit the fine-grained conversational structure, and a novel graph-theoretic supervised model that combines lexical, conversational and topic features. For topic labeling, we propose two novel (unsupervised) random walk models that respectively capture conversation specific clues from two different sources: the leading sentences and the fine-grained conversational structure. Empirical evaluation shows that the segmentation and the labeling performed by our best models beat the state-of-the-art, and are highly correlated with human annotations. Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng |
J. Artif. Intell. Res. | 1 |
| 2012 | A Novel Discriminative Framework for Sentence-Level Discourse Analysis
Shafiq R. Joty, Giuseppe Carenini, Raymond T. Ng |
EMNLP-CoNLL | 1 |
| 2011 | Supervised Topic Segmentation of Email Conversations
Shafiq R. Joty, Giuseppe Carenini, Gabriel Murray, Raymond T. Ng |
ICWSM | 1 |
| 2011 | Unsupervised Modeling of Dialog Acts in Asynchronous Conversations
Shafiq R. Joty, Giuseppe Carenini, Chin-Yew Lin |
IJCAI | 1 |
| 2011 | Improving graph-based random walks for complex question answering using syntactic, shallow semantic and extended string subsequence kernels
Yllias Chali, Sadid A. Hasan, Shafiq R. Joty |
Inf. Process. Manag. | 3 |
| 2010 | Exploiting Conversation Structure in Unsupervised Topic Segmentation for Emails
Shafiq R. Joty, Giuseppe Carenini, Gabriel Murray, Raymond T. Ng |
EMNLP | 1 |
| 2009 | Complex Question Answering: Unsupervised Learning Approaches and ExperimentsabstractComplex questions that require inferencing and synthesizing information from multiple documents can be seen as a kind of topic-oriented, informative multi-document summarization where the goal is to produce a single text as a compressed version of a set of documents with a minimum loss of relevant information. In this paper, we experiment with one empirical method and two unsupervised statistical machine learning techniques: K-means and Expectation Maximization (EM), for computing relative importance of the sentences. We compare the results of these approaches. Our experiments show that the empirical approach outperforms the other two techniques and EM performs better than K-means. However, the performance of these approaches depends entirely on the feature set used and the weighting of these features. In order to measure the importance and relevance to the user query we extract different kinds of features (i.e. lexical, lexical semantic, cosine similarity, basic element, tree kernel based syntactic and shallow-semantic) for each of the document sentences. We use a local search technique to learn the weights of the features. To the best of our knowledge, no study has used tree kernel functions to encode syntactic/semantic information for more complex tasks such as computing the relatedness between the query sentences and the document sentences in order to generate query-focused summaries (or answers to complex questions). For each of our methods of generating summaries (i.e. empirical, K-means and EM) we show the effects of syntactic and shallow-semantic features over the bag-of-words (BOW) features. Yllias Chali, Shafiq R. Joty, Sadid A. Hasan |
J. Artif. Intell. Res. | 2 |
| 2008 | Selecting Sentences for Answering Complex Questions
Yllias Chali, Shafiq R. Joty |
EMNLP | 2 |
| 2008 | Exploiting Syntactic and Shallow Semantic Kernels to Improve Random Walks for Complex Question AnsweringabstractWe consider the problem of answering complex questions that require inferencing and synthesizing information from multiple documents and can be seen as a kind of topic-oriented, informative multi-document summarization. The stochastic, graph-based method for computing the relative importance of textual units (i.e. sentences) is very successful in generic summarization. In this method, a sentence is encoded as a vector in which each component represents the occurrence frequency (TF*IDF) of a word. However, the major limitation of the TF*IDF approach is that it only retains the frequency of the words and does not take into account the sequence, syntactic and semantic information. In this paper, we study the impact of syntactic and shallow semantic information in the graph-based method for answering complex questions. Experimental results show the effectiveness of the syntactic and shallow semantic information for this task. Yllias Chali, Shafiq R. Joty |
ICTAI (2) | 2 |
| 2008 | Answering Complex Questions Using Query-Focused Summarization TechniqueabstractUnlike simple questions, complex questions cannot be answered by simply extracting named entities. These questions require inferencing and synthesizing information from multiple documents that can be seen as a kind of topic-oriented, informative multi-document summarization. In this paper, we have experimented with one empirical and two unsupervised statistical machine learning techniques: k-means and Expectation Maximization (EM), for computing relative importance of the sentences. The feature set includes different kinds of features: lexical, lexical semantic, cosine similarity, basic element, tree kernel based syntactic and shallow-semantic. A gradient descent local search technique is used to learn the optimal weights of the features. The effects of the different features are also shown for all the methods of generating summaries. Yllias Chali, Shafiq R. Joty |
ICTAI (2) | 2 |