Gholamreza Haffari

dblp:57/5129 · also Reza Haffari · DBLP profile ↗
← Back
155ranked-venue papers
9as first author
82since 2021 · last 2026
0000-0001-7326-8380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 135 · 9 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 16 since 2021Databases, data management, data science and information retrieval · 17 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DIET: Machine Unlearning on a Data-Diet
abstract
Machine Unlearning (MU) aims to remove the influence of specific knowledge from a pretrained model. Existing methods often rely on retained training data to preserve utility; such dependence is impractical due to privacy and scalability constraints. A further complication arises when unlearning is applied to vision-language models (VLMs), where entangled multimodal representations make targeted forgetting especially challenging. We propose DIET, a principled retain-data-free unlearning method for VLMs that addresses these challenges by leveraging the geometry of hyperbolic space. The core idea is to push forget embeddings toward class-mismatched prototypes located at the boundary of the hyperbolic space. In hyperbolic geometry, points near the boundary become infinitely distant from interior points. As a result, moving forget embeddings to the boundary makes their influence on the model asymptotically negligible. To formalize this, we guide the forgetting process using the Busemann function, which quantifies directional distance to the boundary. We further develop an adaptive scheme based on optimal transport that selects mismatched prototypes for each forget embedding, enabling flexible unlearning dynamics. Extensive experiments on fine-grained datasets such as Flowers102, OxfordPets, and StanfordCars show that DIET achieves an average forget accuracy of 8.06%, while preserving 69.04% utility using only 16 samples per concept, significantly outperforming the best retain-free baselines with a 117.5% improvement in model utility, and showing competitive performance to retain-data baselines with only a 3.79% drop
Nilakshan Kunananthaseelan, Jing Wu 0021, Trung Le 0001, Gholamreza Haffari, Mehrtash Harandi
AAAI4
2026 LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations
abstract
Large language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability.We introduce LIVECULTUREBENCH 1 , a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms.The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles.Each episode assigns one resident a daily goal while others provide social context.An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty.Using LIVECULTUREBENCH across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed.
Viet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari, Dinh Q. Phung
ACL (1)4
2025 SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
abstract
Zhuang Li, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhuang Li 0001, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari
ACL (1)6
2025 On the Reliability of Large Language Models for Causal Discovery
abstract
This study investigates the efficacy of Large Language Models (LLMs) in causal discovery. Using newly available open-source LLMs, OLMo and BLOOM, which provide access to their pre-training corpora, we investigate how LLMs address causal discovery through three research questions. We examine: (i) the impact of memorization for accurate causal relation prediction, (ii) the influence of incorrect causal relations in pre-training data, and (iii) the contextual nuances that influence LLMs’ understanding of causal relations. Our findings indicate that while LLMs are effective in recognizing causal relations that occur frequently in pre-training data, their ability to generalize to new or rare causal relations is limited. Moreover, the presence of incorrect causal relations significantly undermines the confidence of LLMs in corresponding correct causal relations, and the contextual information critically affects the outcomes of LLMs to discern causal connections between random variables.
Tao Feng 0013, Lizhen Qu, Niket Tandon, Zhuang Li 0001, Xiaoxi Kang, Gholamreza Haffari
ACL (1)6
2025 IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data
abstract
Causal discovery is fundamental to scientific research, yet traditional statistical algorithms face significant challenges, including expensive data collection, redundant computation for known relations, and unrealistic assumptions.While recent LLM-based methods excel at identifying commonly known causal relations, they fail to uncover novel relations.We introduce IRIS (Iterative Retrieval and Integrated System for Real-Time Causal Discovery), a novel framework that addresses these limitations.Starting with a set of initial variables, IRIS automatically collects relevant documents, extracts variables, and uncovers causal relations.Our hybrid causal discovery method combines statistical algorithms and LLM-based methods to discover known and novel causal relations.In addition to causal discovery on initial variables, the missing variable proposal component of IRIS identifies and incorporates missing variables to expand the causal graphs.Our approach enables real-time causal discovery from only a set of initial variables without requiring pre-existing datasets.
Tao Feng 0013, Lizhen Qu, Niket Tandon, Gholamreza Haffari
ACL (1)4
2025 SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social Media
abstract
Opinion survey research is a crucial method used by social scientists for understanding societal beliefs and behaviors.Traditional methodologies often entail high costs and limited scalability, while current automated methods such as opinion synthesis exhibit severe biases and lack traceability.In this paper, we introduce SUR-VEYPILOT, a novel finite-state orchestrated agentic framework that automates the collection and analysis of human opinions from social media platforms.SURVEYPILOT addresses the limitations of pioneering approaches by (i) providing transparency and traceability in each state of opinion collection and (ii) incorporating several techniques for mitigating biases, notably with a novel genetic algorithm for improving result diversity.Our extensive experiments reveal that SURVEYPILOT achieves a close alignment with authentic survey results across multiple domains, observing average relative improvements of 68.98% and 51.37% when comparing to opinion synthesis and agent-based approaches.Implementation of SURVEYPILOT is available on https: //github.com/thanhpv2102/SurveyPilot
Viet Thanh Pham, Lizhen Qu, Zhuang Li 0001, Suraj Sharma, Gholamreza Haffari
ACL (1)5
2025 CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
abstract
Automatically evaluating the quality of responses in dialogue systems is a challenging yet crucial task. Current metrics often fail to align with human judgments, especially when assessing responses that are grammatically correct. To address this issue, we propose a novel metric, called CausalScore, which assesses the relevance of responses by measuring the causal strength between dialogue histories and responses. The causal strength is estimated by utilizing both unconditional dependence and conditional dependencies from dialogue histories to responses. We compare our metric with the existing competitive metrics in terms of their alignment with human judgements. Our experimental results demonstrate that CausalScore significantly surpasses existing state-of-the-art metrics by aligning better with human judgements. Additionally, we collect a dialogue dataset CGDIALOG+ with human-annotated causal relations and a set of pairwise human judgements to facilitate the development of automatic metrics.
Tao Feng 0013, Lizhen Qu, Xiaoxi Kang, Gholamreza Haffari
COLING4
2025 Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
abstract
Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explore adding a new language, i.e., Persian, to Llama (a model with a limited understanding of Persian) using parameter-efficient fine-tuning. We employ a multi-stage approach involving pretraining on monolingual Persian data, aligning representations through bilingual pretraining and instruction datasets, and instruction-tuning with task-specific datasets. We evaluate the model’s performance at each stage on generation and classification tasks. Our findings suggest that incorporating the Persian language, through bilingual data alignment, can enhance classification accuracy for Persian tasks, with no adverse impact and sometimes even improvements on English tasks. Additionally, the results highlight the model’s initial strength as a critical factor when working with limited training data, with cross-lingual alignment offering minimal benefits for the low-resource language. Knowledge transfer from English to Persian has a marginal effect, primarily benefiting simple classification tasks.
Samin Mahdizadeh Sani, Pouya Sadeghi, Thuy-Trang Vu, Yadollah Yaghoobzadeh, Gholamreza Haffari
COLING5
2025 Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
abstract
Large Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions.However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safetyalignment.Despite advances in defence measures for text and vision LLMs, effective safetyalignment strategies and audio-safety dataset specifically targeting LALMs are notably absent.Meanwhile defence measures based on Supervised Fine-tuning (SFT) struggle to address safety improvement while avoiding overrejection issues, significantly compromising helpfulness.In this work, we propose an unsupervised safety-fine-tuning strategy as remedy that reshapes model's representation space to enhance existing LALMs safety-alignment while balancing the risk of over-rejection.Our experiments, conducted across three generations of Qwen LALMs, demonstrate that our approach significantly improves LALMs safety under three modality input conditions (audiotext, text-only, and audio-only) while increasing over-rejection rate by only 0.88% on average. 1 Warning: this paper contains harmful examples.
Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari
EMNLP4
2025 Naver: a Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
abstract
Visual Grounding (VG) tasks, such as referring expression detection and segmentation tasks are important for linking visual entities to context, especially in complex reasoning tasks that require detailed query interpretation. This paper explores VG beyond basic perception, highlighting challenges for methods that require reasoning like human cognition. Recent advances in large language methods (LLMs) and Vision-Language methods (VLMs) have improved abilities for visual comprehension, contextual understanding, and reasoning. These methods are mainly split into end-to-end and compositional methods, with the latter offering more flexibility. Compositional approaches that integrate LLMs and foundation models show promising performance but still struggle with complex reasoning with language-based logical representations. To address these limitations, we propose NAVER, a compositional visual grounding method that integrates explicit probabilistic logic reasoning within a finite-state automaton, equipped with a self-correcting mechanism. This design improves robustness and interpretability in inference through explicit logic reasoning. Our results show that NAVER achieves SoTA performance comparing to recent end-to-end and compositional baselines. The code is available at https://github.com/ControlNet/NAVER .
Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Gholamreza Haffari, Peter J. Stuckey, Seyed Hamid Rezatofighi
ICCV5
2025 Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models
abstract
Large language models (LLMs) have demonstrated impressive reasoning abilities, but they still struggle with faithful reasoning due to knowledge gaps and hallucinations. To address these issues, knowledge graphs (KGs) have been utilized to enhance LLM reasoning through their structured knowledge. However, existing KG-enhanced methods, either retrieval-based or agent-based, encounter difficulties in accurately retrieving knowledge and efficiently traversing KGs at scale. In this work, we introduce graph-constrained reasoning (GCR), a novel framework that bridges structured knowledge in KGs with unstructured reasoning in LLMs. To eliminate hallucinations, GCR ensures faithful KG-grounded reasoning by integrating KG structure into the LLM decoding process through KG-Trie, a trie-based index that encodes KG reasoning paths. KG-Trie constrains the decoding process, allowing LLMs to directly reason on graphs and generate faithful reasoning paths grounded in KGs. Additionally, GCR leverages a lightweight KG-specialized LLM for graph-constrained reasoning alongside a powerful general LLM for inductive reasoning over multiple reasoning paths, resulting in accurate reasoning with zero reasoning hallucination. Extensive experiments on several KGQA benchmarks demonstrate that GCR achieves state-of-the-art performance and exhibits strong zero-shot generalizability to unseen KGs without additional training.
Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Yuan-Fang Li, Chen Gong 0002, Shirui Pan
ICML3
2025 The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
abstract
The performance of large language models (LLMs) is strongly influenced by the quality and diversity of data used during supervised fine-tuning (SFT). However, current data selection methods often prioritize one aspect over the other, resulting in suboptimal training outcomes. To address this, we formulate data selection as a set cover problem and present GraphFilter, a novel approach that balances both quality and diversity in data selection. GraphFilter models the dataset as a bipartite graph connecting sentences to their constituent n-grams, then employs a priority function that combines quality and diversity metrics multiplicatively. GraphFilter iteratively selects sentences with the highest priority, removes covered n-grams from the bipartite graph, and recomputes priorities to reflect the changing data landscape. We validate GraphFilter using three model backbones across six widely-used benchmarks, demonstrating that it outperforms nine existing baselines in both model performance and computational efficiency. Further analysis shows that our design choices lead to more effective subset selection, underscores the value of instruction diversity, and provides insights into how quality and diversity interact with different subset sizes.
Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Gholamreza Haffari
ICML4
2025 SpeechDialogueFactory: A Framework for Natural Speech Dialogue Generation
Minghan Wang, Ye Bai 0002, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari
INTERSPEECH6
2025 Continual Speech Learning with Fused Speech Features
Guitao Wang, Jinming Zhao, Guilin Qi, Tongtong Wu, Gholamreza Haffari
INTERSPEECH6
2025 MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation
abstract
Junqing He, Liang Zhu, Rui Wang, Xi Wang, Gholamreza Haffari, Jiaxing Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junqing He, Xi Wang 0012, Gholamreza Haffari
NAACL (Long Papers)5
2025 CultureInstruct: Curating Multi-Cultural Instructions at Scale
abstract
Viet Thanh Pham, Zhuang Li, Lizhen Qu, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Viet Thanh Pham, Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
NAACL (Long Papers)4
2025 ACCESS : A Benchmark for Abstract Causal Event Discovery and Reasoning
abstract
Vy Vo, Lizhen Qu, Tao Feng, Yuncheng Hua, Xiaoxi Kang, Songhai Fan, Tim Dwyer, Lay-Ki Soon, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Vy Vo, Lizhen Qu, Tao Feng 0013, Yuncheng Hua, Xiaoxi Kang, Songhai Fan, Tim Dwyer, Lay-Ki Soon, Gholamreza Haffari
NAACL (Long Papers)9
2025 Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
abstract
Hao Yang, Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari
NAACL (Long Papers)4
2025 FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
abstract
Diffusion generative models have become the standard for producing high-quality, coherent video content, yet their slow inference speeds and high computational demands hinder practical deployment. Although both quantization and sparsity can independently accelerate inference while maintaining generation quality, naively combining these techniques in existing training-free approaches leads to significant performance degradation, as they fail to achieve proper joint optimization. We introduce FPSAttention, a novel training-aware co-design of FP8 quantization and Sparsity for video generation, with a focus on the 3D bi-directional attention mechanism. Our approach features three key innovations: 1) A unified 3D tile-wise granularity that simultaneously supports both quantization and sparsity. 2) A denoising step-aware strategy that adapts to the noise schedule, addressing the strong correlation between quantization/sparsity errors and denoising steps. 3) A native, hardware-friendly kernel that leverages FlashAttention and is implemented with optimized Hopper architecture features, enabling highly efficient execution. Trained on Wan2.1's 1.3B and 14B models and evaluated on the vBench benchmark, FPSAttention achieves a 7.09$\times$ kernel speedup for attention operations and a 4.96$\times$ end-to-end speedup for video generation compared to the BF16 baseline at 720p resolution—without sacrificing generation quality.
Akide Liu, Zeyu Zhang 0006, Zhexin Li, Xuehai Bai, Yuanjie Xing, Yizeng Han, Jiasheng Tang, Jichao Wu, Mingyang Yang, Yuanyu He, Fan Wang 0019, Gholamreza Haffari, Bohan Zhuang
NeurIPS14
2025 GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) has proven effective in integrating knowledge into large language models (LLMs). However, conventional RAGs struggle to capture complex relationships between pieces of knowledge, limiting their performance in intricate reasoning that requires integrating knowledge from multiple sources. Recently, graph-enhanced retrieval augmented generation (GraphRAG) builds a graph structure to explicitly model these relationships, enabling more effective and efficient retrievers. Nevertheless, its performance is still hindered by the noise and incompleteness within the graph structure. To address this, we introduce GFM-RAG, a novel graph foundation model (GFM) for retrieval augmented generation. GFM-RAG is powered by an innovative graph neural network that reasons over graph structure to capture complex query-knowledge relationships. The GFM with 8M parameters undergoes a two-stage training process on large-scale datasets, comprising 60 knowledge graphs with over 14M triples and 700k documents. This results in impressive performance and generalizability for GFM-RAG, making it the first graph foundation model applicable to unseen datasets for retrieval without any fine-tuning required. Extensive experiments on three multi-hop QA datasets and seven domain-specific RAG datasets demonstrate that GFM-RAG achieves state-of-the-art performance while maintaining efficiency and alignment with neural scaling laws, highlighting its potential for further improvement.
Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Dinh Q. Phung, Chen Gong 0002, Shirui Pan
NeurIPS3
2025 Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning
abstract
Audio captioning systems face a fundamental challenge: teacher-forcing training creates exposure bias that leads to caption degeneration during inference. While contrastive methods have been proposed as solutions, they typically fail to capture the crucial temporal relationships between acoustic and linguistic modalities. We address this limitation by introducing the unbiased sliced Wasserstein RBF (USW-RBF) kernel with rotary positional embedding, specifically designed to preserve temporal information across modalities. Our approach offers a practical advantage: the kernel enables efficient stochastic gradient optimization, making it computationally feasible for real-world applications. Building on this foundation, we develop a complete audio captioning framework that integrates stochastic decoding to further mitigate caption degeneration. Extensive experiments on AudioCaps and Clotho datasets demonstrate that our method significantly improves caption quality, lexical diversity, and text-to-audio retrieval accuracy. Furthermore, we demonstrate the generalizability of our USW-RBF kernel by applying it to audio reasoning tasks, where it enhances the reasoning capabilities of large audio language models on the CompA-R in terms of correctness and quality. Our kernel also improves the reasoning accuracy of the MMAU-test-mini benchmarks by $4\%$. These results establish our approach as a powerful and generalizable solution for cross-modal alignment challenges in audio-language tasks.
Manh Luong, Dinh Q. Phung, Gholamreza Haffari, Lizhen Qu
NeurIPS4
2025 ChatRule: Mining Logical Rules with Large Language Models for Knowledge Graph Reasoning
Linhao Luo, Jiaxin Ju, Bo Xiong 0001, Yuan-Fang Li, Gholamreza Haffari, Shirui Pan
PAKDD (2)5
2025 (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
abstract
Abstract Literary translations remains one of the most challenging frontiers in machine translation due to the complexity of capturing figurative language, cultural nuances, and unique stylistic elements. In this work, we introduce TransAgents, a novel multi-agent framework that simulates the roles and collaborative practices of a human translation company, including a CEO, Senior Editor, Junior Editor, Translator, Localization Specialist, and Proofreader. The translation process is divided into two stages: a preparation stage where the team is assembled and comprehensive translation guidelines are drafted, and an execution stage that involves sequential translation, localization, proofreading, and a final quality check. Furthermore, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP), which evaluates translations based solely on target language quality and cultural appropriateness, and BLP, which leverages large language models like gpt-4 for direct text comparison. Although TransAgents achieves lower d-BLEU scores, due to the limited diversity of references, its translations are significantly better than those of other baselines and are preferred by both human evaluators and LLMs over traditional human references and gpt-4 translations. Our findings highlight the potential of multi-agent collaboration in enhancing translation quality, particularly for longer texts.1
Minghao Wu, Yulin Yuan, Gholamreza Haffari, Longyue Wang, Weihua Luo, Kaifu Zhang
Trans. Assoc. Comput. Linguistics4
2025 Scalable Frame-Based Construction of Sociocultural Norm Bases for Socially Aware Dialogues
abstract
Sociocultural norms serve as guiding principles for personal conduct in social interactions, emphasizing respect, cooperation, and appropriate behavior, which is able to benefit tasks including conversational information retrieval, contextual information retrieval, and retrieval-enhanced machine learning. We propose a scalable approach for constructing a Sociocultural Norm ( Scn ) Base using large language models (LLMs) for socially aware dialogues. We construct a comprehensive and publicly accessible Chinese Sociocultural NormBase ( ChineseNormBase ). Our approach utilizes socially aware dialogues, enriched with contextual frames, as the primary data source to constrain the generating process and reduce the hallucinations. This enables extracting of high-quality and nuanced natural-language norm statements, leveraging the pragmatic implications of utterances with respect to the situation. As real dialogue annotated with gold frames are not readily available, we propose using synthetic data. Our empirical results show (i) the quality of the Scn s derived from synthetic data is comparable to that from real dialogues annotated with gold frames, and (ii) the quality of the Scn s extracted from real data, annotated with either silver (predicted) or gold frames, surpasses that without the frame annotations. We further show the effectiveness of the extracted Scn s in a Retrieval-Augmented Generation (RAG)-based model to reason about multiple downstream dialogue tasks.
Shilin Qu, Weiqing Wang 0001, Xin Zhou 0023, Haolan Zhan, Zhuang Li 0001, Lizhen Qu, Linhao Luo, Yuan-Fang Li, Gholamreza Haffari
ACM Trans. Multim. Comput. Commun. Appl.9
2024 IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
abstract
Machine learning models have made incredible progress, but they still struggle when applied to examples from unseen domains.This study focuses on a specific problem of domain generalization, where a model is trained on one source domain and tested on multiple target domains that are unseen during training.We propose IMO: Invariant features Masks for Out-of-Distribution text classification, to achieve OOD generalization by learning domain-invariant features.During training, IMO employs a greedy algorithm to learn sparse representations for each layer in a top-down manner.It performs better than the opposite direction and learning of sparse representations for all layers simultaneously.Our comprehensive experiments show that IMO substantially outperforms strong baselines such as prompt-based methods and large language models, in terms of various evaluation metrics and settings.1
Tao Feng 0013, Lizhen Qu, Zhuang Li 0001, Haolan Zhan, Yuncheng Hua, Gholamreza Haffari
ACL (1)6
2024 Importance-Aware Data Augmentation for Document-Level Neural Machine Translation
abstract
Minghao Wu, Yufei Wang, George Foster, Lizhen Qu, Gholamreza Haffari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Minghao Wu, Yufei Wang 0003, George F. Foster, Lizhen Qu, Gholamreza Haffari
EACL (1)5
2024 Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs
abstract
Large Language Models (LLMs) demonstrate significant capabilities in processing natural language data, promising efficient knowledge extraction from diverse textual sources to enhance situational awareness and support decision-making. However, concerns arise due to their susceptibility to hallucination, resulting in contextually inaccurate content. This work focuses on harnessing LLMs for automated Event Extraction, introducing a new method to address hallucination by decomposing the task into Event Detection and Event Argument Extraction. Moreover, the proposed method integrates dynamic schema-aware augmented retrieval examples into prompts tailored for each specific inquiry, thereby extending and adapting advanced prompting techniques such as Retrieval-Augmented Generation. Evaluation findings on prominent event extraction benchmarks and results from a synthesized benchmark illustrate the method’s superior performance compared to baseline approaches.
Fatemeh Shiri, Farhad Moghimifar, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002, John Yoo
FUSION3
2024 Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
abstract
Large language models (LLMs) have demonstrated impressive reasoning abilities in complex tasks. However, they lack up-to-date knowledge and experience hallucinations during reasoning, which can lead to incorrect reasoning processes and diminish their performance and trustworthiness. Knowledge graphs (KGs), which capture vast amounts of facts in a structured format, offer a reliable source of knowledge for reasoning. Nevertheless, existing KG-based LLM reasoning methods only treat KGs as factual knowledge bases and overlook the importance of their structural information for reasoning. In this paper, we propose a novel method called reasoning on graphs (RoG) that synergizes LLMs with KGs to enable faithful and interpretable reasoning. Specifically, we present a planning-retrieval-reasoning framework, where RoG first generates relation paths grounded by KGs as faithful plans. These plans are then used to retrieve valid reasoning paths from the KGs for LLMs to conduct faithful reasoning. Furthermore, RoG not only distills knowledge from KGs to improve the reasoning ability of LLMs through training but also allows seamless integration with any arbitrary LLMs during inference. Extensive experiments on two benchmark KGQA datasets demonstrate that RoG achieves state-of-the-art performance on KG reasoning tasks and generates faithful and interpretable reasoning results.
Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui Pan
ICLR3
2024 Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation
abstract
The Learning-to-match (LTM) framework proves to be an effective inverse optimal transport approach for learning the underlying ground metric between two sources of data, facilitating subsequent matching. However, the conventional LTM framework faces scalability challenges, necessitating the use of the entire dataset each time the parameters of the ground metric are updated. In adapting LTM to the deep learning context, we introduce the mini-batch Learning-to-match (m-LTM) framework for audio-text retrieval problems. This framework leverages mini-batch subsampling and Mahalanobis-enhanced family of ground metrics. Moreover, to cope with misaligned training data in practice, we propose a variant using partial optimal transport to mitigate the harm of misaligned data pairs in training data. We conduct extensive experiments on audio-text matching problems using three datasets: AudioCaps, Clotho, and ESC-50. Results demonstrate that our proposed method is capable of learning rich and expressive joint embedding space, which achieves SOTA performance. Beyond this, the proposed m-LTM framework is able to close the modality gap across audio and text embedding, which surpasses both triplet and contrastive loss in the zero-shot sound event detection task on the ESC-50 dataset. Notably, our strategy of employing partial optimal transport with m-LTM demonstrates greater noise tolerance than contrastive loss, especially under varying noise ratios in training data on the AudioCaps dataset. Our code is available at https://github.com/v-manhlt3/m-LTM-Audio-Text-Retrieval
Manh Luong, Nhat Ho, Gholamreza Haffari, Dinh Q. Phung, Lizhen Qu
ICLR4
2024 Generating Simple, Conservative and Unifying Explanations for Logistic Regression Models
abstract
In this paper, we generate and compare three types of explanations of Machine Learning (ML) predictions: simple, conservative and unifying.Simple explanations are concise, conservative explanations address the surprisingness of a prediction, and unifying explanations convey the extent to which an ML model's predictions are applicable.The results of our user study show that (1) conservative and unifying explanations are liked equally and considered largely equivalent in terms of completeness, helpfulness for understanding the AI, and enticement to act, and both are deemed better than simple explanations; and (2) users' views about explanations are influenced by the (dis)agreement between the ML model's predictions and users' estimations of these predictions, and by the inclusion/omission of features users expect to see in explanations.
Sameen Maruf, Ingrid Zukerman, Xuelin Situ, Cécile Paris, Gholamreza Haffari
INLG5
2024 Measuring Affective and Motivational States as Conditions for Cognitive and Metacognitive Processing in Self-Regulated Learning
abstract
Even though the engagement in self-regulated learning (SRL) has been shown to boost academic performance, SRL skills of many learners remain underdeveloped. They often struggle to productively navigate multiple cognitive, affective, metacognitive and motivational (CAMM) processes in SRL. To provide learners with the required SRL support, it is essential to understand how learners enact CAMM processes as they study. More research is needed to advance the measurement of affective and motivational processes within SRL, and investigate how these processes influence learners’ cognition and metacognition. With this in mind, we conducted a lab study involving 22 university students who worked on a 45-minute reading and writing task in digital learning environment. We used a wearable electroencephalogram device to record learner academic emotional and motivational states, and digital trace data to record learner cognitive and metacognitive processes. We harnessed time series prediction and explainable artificial intelligence methods to examine how learner’s emotional and motivational states influence their choice of cognitive and metacognitive processes. Our results indicate that emotional and motivational states can predict learners’ use of low cognitive, high cognitive and metacognitive processes with considerable classification accuracy (F1 > 0.73), and that higher values of interest, engagement and excitement promote cognitive processing.
Mladen Rakovic, Navid Mohammadi Foumani, Mahsa Salehi, Levin Kuhlmann, Geoffrey Mackellar, Roberto Martínez-Maldonado, Gholamreza Haffari, Zach Swiecki, Xinyu Li 0004, Guanliang Chen, Dragan Gasevic
LAK8
2024 MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
abstract
A critical approach for efficiently deploying computationally demanding large language models (LLMs) is Key-Value (KV) caching. The KV cache stores key-value states of previously generated tokens, significantly reducing the need for repetitive computations and thereby lowering latency in autoregressive generation. However, the size of the KV cache grows linearly with sequence length, posing challenges for applications requiring long context input and extensive sequence generation. In this paper, we present a simple yet effective approach, called MiniCache, to compress the KV cache across layers from a novel depth perspective, significantly reducing the memory footprint for LLM inference. Our approach is based on the observation that KV cache states exhibit high similarity between the adjacent layers in the middle-to-deep portion of LLMs. To facilitate merging, we propose disentangling the states into the magnitude and direction components, interpolating the directions of the state vectors while preserving their lengths unchanged. Furthermore, we introduce a token retention strategy to keep highly distinct state pairs unmerged, thus preserving the information with minimal additional storage overhead. Our MiniCache is training-free and general, complementing existing KV cache compression strategies, such as quantization and sparsity. We conduct a comprehensive evaluation of MiniCache utilizing various models including LLaMA-2, LLaMA-3, Phi-3, Mistral, and Mixtral across multiple benchmarks, demonstrating its exceptional performance in achieving superior compression ratios and high throughput. On the ShareGPT dataset, LLaMA-2-7B with cross-layer merging achieves a compression ratio of $1.53\times$. Additionally, since MiniCache is orthogonal to existing quantization techniques, it can achieve a compression ratio of up to $5.02\times$ when combined with the 4-bit quantization technique, enhancing inference throughput by approximately $5\times$ and reducing the memory footprint by $41\%$ compared to the FP16 full cache baseline, all while maintaining near-lossless performance. Project is available at https://minicache.vmv.re .
Akide Liu, Jing Liu 0048, Zizheng Pan, Yefei He, Gholamreza Haffari, Bohan Zhuang
NeurIPS5
2024 Going beyond Imagination! Enhancing Multi-modal Dialogue Agents with Synthetic Visual Descriptions
abstract
Building a dialogue agent that can seamlessly interact with humans, in multi-modal regimes, requires two fundamental abilities: (1) understanding emotion and dialogue acts within situated user scenarios, and (2) grounding perceived visual cues to dialogue contexts.However, recent works have uncovered shortcomings of existing dialogue agents in understanding emotions and dialogue acts, and in grounding visual cues effectively.In this work, we investigate whether additional dialogue data with only visual descriptions can help dialogue agents effectively align visual and textual features, and enhance the ability of dialogue agents to ground perceived visual cues to dialogue contexts.To this end, in the absence of a suitable dataset, we propose a synthetic visual description generation pipeline, and contribute a large-scale synthetic visual description dataset.In addition, we propose a general training procedure for effectively leveraging these synthetic data.We conduct comprehensive analyses to evaluate the impact of synthetic data on two benchmarks: MELD and IEMO-CAP.Our findings suggest that synthetic visual descriptions can serve as an effective way to enhance a dialogue agents' grounding ability, and that the training scheme affects the extent to which these descriptions improve the agent's performance.
Haolan Zhan, Sameen Maruf, Ingrid Zukerman, Gholamreza Haffari
SIGDIAL4
2023 The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active Learning
abstract
Multilingual semantic parsing aims to leverage the knowledge from the high-resource languages to improve low-resource semantic parsing, yet commonly suffers from the data imbalance problem.Prior works propose to utilize the translations by either humans or machines to alleviate such issues.However, human translations are expensive, while machine translations are cheap but prone to error and bias.In this work, we propose an active learning approach that exploits the strengths of both human and machine translations by iteratively adding small batches of human translations into the machine-translated training set.Besides, we propose novel aggregated acquisition criteria that help our active learning method select utterances to be manually translated.Our experiments demonstrate that an ideal utterance selection can significantly reduce the error and bias in the translated data, resulting in higher parser accuracies than the parsers merely trained on the machine-translated data.
Zhuang Li 0001, Lizhen Qu, Phil Cohen 0001, Raj Tumuluri, Gholamreza Haffari
ACL (1)5
2023 A Minimal Approach for Natural Language Action Space in Text-based Games
abstract
Text-based games (TGs) are language-based interactive environments for reinforcement learning.While language models (LMs) and knowledge graphs (KGs) are commonly used for handling large action space in TGs, it is unclear whether these techniques are necessary or overused.In this paper, we revisit the challenge of exploring the action space in TGs and propose ϵ-admissible exploration, a minimal approach of utilizing admissible actions, for training phase.Additionally, we present a textbased actor-critic (TAC) agent that produces textual commands for game, solely from game observations, without requiring any KG or LM.Our method, on average across 10 games from Jericho, outperforms strong baselines and stateof-the-art agents that use LM and KG.Our approach highlights that a much lighter model design, with a fresh perspective on utilizing the information within the environments, suffices for an effective exploration of exponentially large action spaces. 1
Dongwon Ryu, Gholamreza Haffari, Shirui Pan, Ehsan Shareghi
CoNLL3
2023 MARLIN: Masked Autoencoder for facial video Representation LearnINg
abstract
This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our proposed framework, named MARLIN, is a facial video masked autoencoder, that learns highly robust and generic facial embeddings from abundantly available non-annotated web crawled facial videos. As a challenging auxiliary task, MARLIN reconstructs the spatio-temporal details of the face from the densely masked facial regions which mainly include eyes, nose, mouth, lips, and skin to capture local and global aspects that in turn help in encoding generic and transferable features. Through a variety of experiments on diverse downstream tasks, we demonstrate MARLIN to be an excellent facial video encoder as well as feature extractor, that performs consistently well across a variety of downstream tasks including FAR (1.13% gain over supervised benchmark), FER (2.64% gain over unsupervised benchmark), DFD (1.86% gain over unsupervised benchmark), LS (29.36% gain for Frechet Inception Distance), and even in low data regime. Our code and models are available at https://github.com/ControlNet/MARLIN.
Zhixi Cai, Shreya Ghosh 0001, Kalin Stefanov, Abhinav Dhall, Jianfei Cai 0001, Seyed Hamid Rezatofighi, Gholamreza Haffari, Munawar Hayat
CVPR7
2023 Protocon: Pseudo-Label Refinement via Online Clustering and Prototypical Consistency for Efficient Semi-Supervised Learning
abstract
Confidence-based pseudo-labeling is among the dominant approaches in semi-supervised learning (SSL). It relies on including high-confidence predictions made on unlabeled data as additional targets to train the model. We propose Protocon, a novel SSL method aimed at the less-explored label-scarce SSL where such methods usually underperform. Protocon refines the pseudolabels by lever-aging their nearest neighbours' information. The neighbours are identified as the training proceeds using an online clustering approach operating in an embedding space trained via a prototypical loss to encourage well-formed clusters. The online nature of Protocon allows it to utilise the label history of the entire dataset in one training cycle to refine labels in the following cycle without the need to store image embeddings. Hence, it can seamlessly scale to larger datasets at a low cost. Finally, Protocon addresses the poor training signal in the initial phase of training (due to fewer confident predictions) by introducing an auxiliary self-supervised loss. It delivers significant gains and faster convergence over state-of-the-art across 5 datasets, including CIFARs, ImageNet and DomainNet.
Islam Nassar, Munawar Hayat, Ehsan Abbasnejad, Seyed Hamid Rezatofighi, Gholamreza Haffari
CVPR5
2023 Document Flattening: Beyond Concatenating Context for Document-Level Neural Machine Translation
abstract
Existing work in document-level neural machine translation commonly concatenates several consecutive sentences as a pseudodocument, and then learns inter-sentential dependencies.This strategy limits the model's ability to leverage information from distant context.We overcome this limitation with a novel Document Flattening (DOCFLAT) technique that integrates FLAT-BATCH ATTEN-TION (FBA) and NEURAL CONTEXT GATE (NCG) into Transformer model to utilize information beyond the pseudo-document boundaries.FBA allows the model to attend to all the positions in the batch and learns the relationships between positions explicitly and NCG identifies the useful information from the distant context.We conduct comprehensive experiments and analyses on three benchmark datasets for English-German translation, and validate the effectiveness of two variants of DOCFLAT.Empirical results show that our approach outperforms strong baselines with statistical significance on BLEU, COMET and accuracy on the contrastive test set.The analyses highlight that DOCFLAT is highly effective in capturing the long-range information.
Minghao Wu, George F. Foster, Lizhen Qu, Gholamreza Haffari
EACL4
2023 On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on Codex
abstract
Terry Yue Zhuo, Zhuang Li, Yujin Huang, Fatemeh Shiri, Weiqing Wang, Gholamreza Haffari, Yuan-Fang Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Terry Yue Zhuo, Zhuang Li 0001, Yujin Huang, Fatemeh Shiri, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li
EACL6
2023 DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding
abstract
Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions.One representative benchmark for its study is Social Intelligence Queries (Social-IQ), a dataset of multiplechoice questions on videos of complex social interactions.We define a comprehensive methodology to study the soundness of Social-IQ, as the soundness of such benchmark datasets is crucial to the investigation of the underlying research problem.Our analysis reveals that Social-IQ contains substantial biases, which can be exploited by a moderately strong language model to learn spurious correlations to achieve perfect performance without being given the context or even the question.We introduce DeSIQ, a new challenging dataset, constructed by applying simple perturbations to Social-IQ.Our empirical analysis shows De-SIQ significantly reduces the biases in the original Social-IQ dataset.Furthermore, we examine and shed light on the effect of model size, model style, learning settings, commonsense knowledge, and multi-modality on the new benchmark performance.Our new dataset, observations and findings open up important research questions for the study of social intelligence.
Xiaoyu Guo 0009, Yuan-Fang Li, Gholamreza Haffari
EMNLP3
2023 Few-shot Domain-Adaptative Visually-fused Event Detection from Text
abstract
Incorporating auxiliary modalities such as images into event detection models has attracted increasing interest over the last few years. The complexity of natural language in describing situations has motivated researchers to leverage the related visual context to improve event detection performance. However, current approaches in this area suffer from data scarcity, where a large amount of labelled text-image pairs are required for model training. Furthermore, limited access to the visual context at inference time negatively impacts the performance of such models, which makes them practically ineffective in real-world scenarios. In this paper, we present a novel domain-adaptive visually-fused event detection approach that can be trained on a few labelled image-text paired data points. Specifically, we introduce a visual imaginator method that synthesises images from text in the absence of visual context. Moreover, the imaginator can be customised to a specific domain. In doing so, our model can leverage the capabilities of pre-trained vision-language models and can be trained in a few-shot setting. This also allows for effective inference where only single-modality data (i.e. text) is available. The experimental evaluation on the benchmark M2E2 dataset shows that our model outperforms existing state-of-the-art models, by up to 11 points.
Farhad Moghimifar, Fatemeh Shiri, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002
FUSION3
2023 Energy-based Self-Training and Normalization for Unsupervised Domain Adaptation
abstract
We propose an Unsupervised Domain Adaptation (UDA) method by making use of Energy-Based Learning (EBL) and demonstrate 1. EBL can be used to improve the instance selection for a self-training task on the unlabelled target domain, and 2. alignment and normalizing energy scores can learn domain-invariant representations. For the former, we show that an energy-based selection criterion can be used to model instance selections by mimicking the joint distribution between data and predictions in the target domain. As per learning domain invariant representations, we show that stable domain alignment can be achieved by a combined energy alignment and an energy normalization process. We implement our method in consistent with the vision-transformer (ViT) backbone and show that our proposed method can outperform state-of-the-art ViT based UDA methods on diverse benchmarks (DomainNet, Office-Home, and VISDA2017).
Samitha Herath, Basura Fernando, Ehsan Abbasnejad, Munawar Hayat, Shahram Khadivi, Mehrtash Harandi, Seyed Hamid Rezatofighi, Gholamreza Haffari
ICCV8
2023 Learning Object-Language Alignments for Open-Vocabulary Object Detection
Chuang Lin 0003, Peize Sun, Yi Jiang 0009, Ping Luo 0002, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan, Jianfei Cai 0001
ICLR6
2023 An Additive Instance-Wise Approach to Multi-class Model Interpretation
Vy Vo, Van Nguyen 0002, Trung Le 0001, Quan Hung Tran, Gholamreza Haffari, Seyit Ahmet Çamtepe, Dinh Q. Phung
ICLR5
2023 Reranking for Natural Language Generation from Logical Forms: A Study based on Large Language Models
abstract
Levon Haroutunian, Zhuang Li, Lucian Galescu, Philip Cohen, Raj Tumuluri, Gholamreza Haffari. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Levon Haroutunian, Zhuang Li 0001, Lucian Galescu, Phil Cohen 0001, Raj Tumuluri, Gholamreza Haffari
IJCNLP (1)6
2023 Investigating Pre-trained Audio Encoders in the Low-Resource Condition
abstract
Pre-trained speech encoders have been central to pushing state-of-the-art results across various speech understanding and generation tasks. Nonetheless, the capabilities of these encoders in low-resource settings are yet to be thoroughly explored. To address this, we conduct a comprehensive set of experiments using a representative set of 3 state-of-the-art encoders (Wav2vec2, WavLM, Whisper) in the low-resource setting across 7 speech understanding and generation tasks. We provide various quantitative and qualitative analyses on task performance, convergence speed, and representational properties of the encoders. We observe a connection between the pre-training protocols of these encoders and the way in which they capture information in their internal layers. In particular, we observe the Whisper encoder exhibits the greatest low-resource capabilities on content-driven tasks in terms of performance and convergence speed.
Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi
INTERSPEECH3
2023 Feature-based Learning for Diverse and Privacy-Preserving Counterfactual Explanations
abstract
Interpretable machine learning seeks to understand the reasoning process of complex black-box systems that are long notorious for lack of explainability. One flourishing approach is through counterfactual explanations, which provide suggestions on what a user can do to alter an outcome. Not only must a counterfactual example counter the original prediction from the black-box classifier but it should also satisfy various constraints for practical applications. Diversity is one of the critical constraints that however remains less discussed. While diverse counterfactuals are ideal, it is computationally challenging to simultaneously address some other constraints. Furthermore, there is a growing privacy concern over the released counterfactual data. To this end, we propose a feature-based learning framework that effectively handles the counterfactual constraints and contributes itself to the limited pool of private explanation models. We demonstrate the flexibility and effectiveness of our method in generating diverse counterfactuals of actionability and plausibility. Our counterfactual engine is more efficient than counterparts of the same capacity while yielding the lowest re-identification risks.
Vy Vo, Trung Le 0001, Van Nguyen 0002, He Zhao 0001, Edwin V. Bonilla, Gholamreza Haffari, Dinh Q. Phung
KDD6
2023 Normalizing Flow-based Neural Process for Few-Shot Knowledge Graph Completion
abstract
Knowledge graphs (KGs), as a structured form of knowledge representation, have been widely applied in the real world. Recently, few-shot knowledge graph completion (FKGC), which aims to predict missing facts for unseen relations with few-shot associated facts, has attracted increasing attention from practitioners and researchers. However, existing FKGC methods are based on metric learning or meta-learning, which often suffer from the out-of-distribution and overfitting problems. Meanwhile, they are incompetent at estimating uncertainties in predictions, which is critically important as model predictions could be very unreliable in few-shot settings. Furthermore, most of them cannot handle complex relations and ignore path information in KGs, which largely limits their performance. In this paper, we propose a normalizing flow-based neural process for few-shot knowledge graph completion (NP-FKGC). Specifically, we unify normalizing flows and neural processes to model a complex distribution of KG completion functions. This offers a novel way to predict facts for few-shot relations while estimating the uncertainty. Then, we propose a stochastic ManifoldE decoder to incorporate the neural process and handle complex relations in few-shot settings. To further improve performance, we introduce an attentive relation path-based graph neural network to capture path information in KGs. Extensive experiments on three public datasets demonstrate that our method significantly outperforms the existing FKGC methods and achieves state-of-the-art performance. Code is available at https://github.com/RManLuo/NP-FKGC.git.
Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui Pan
SIGIR3
2023 SocialDial: A Benchmark for Socially-Aware Dialogue Systems
abstract
Content Warning: this paper may contain content that is offensive or upsetting.
Haolan Zhan, Zhuang Li 0001, Yufei Wang 0003, Linhao Luo, Tao Feng 0013, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, Gholamreza Haffari
SIGIR13
2023 LAVA:Label-efficient Visual Learning and Adaptation
abstract
We present LAVA, a simple yet effective method for multi-domain visual transfer learning with limited data. LAVA builds on a few recent innovations to enable adapting to partially labelled datasets with class and domain shifts. First, LAVA learns self-supervised visual representations on the source dataset and ground them using class label semantics to overcome transfer collapse problems associated with supervised pretraining. Secondly, LAVA maximises the gains from unlabelled target data via a novel method which uses multi-crop augmentations to obtain highly robust pseudo-labels. By combining these ingredients, LAVA achieves a new state-of-the-art on ImageNet semi-supervised protocol, as well as on 7 out of 10 datasets in multi-domain few-shot learning on the Meta-dataset.1
Islam Nassar, Munawar Hayat, Ehsan Abbasnejad, Seyed Hamid Rezatofighi, Mehrtash Harandi, Gholamreza Haffari
WACV6
2023 Graph Sequential Neural ODE Process for Link Prediction on Dynamic and Sparse Graphs
abstract
Link prediction on dynamic graphs is an important task in graph mining. Existing approaches based on dynamic graph neural networks (DGNNs) typically require a significant amount of historical data (interactions over time), which is not always available in practice. The missing links over time, which is a common phenomenon in graph data, further aggravates the issue and thus creates extremely sparse and dynamic graphs. To address this problem, we propose a novel method based on the neural process, called Graph Sequential Neural ODE Process (GSNOP). Specifically, GSNOP combines the advantage of the neural process and neural ordinary differential equation that models the link prediction on dynamic graphs as a dynamic-changing stochastic process. By defining a distribution over functions, GSNOP introduces the uncertainty into the predictions, making it generalize to more situations instead of overfitting to the sparse data. GSNOP is also agnostic to model structures that can be integrated with any DGNN to consider the chronological and geometrical information for link prediction. Extensive experiments on three dynamic graph datasets show that GSNOP can significantly improve the performance of existing DGNNs and outperform other neural process variants.
Linhao Luo, Gholamreza Haffari, Shirui Pan
WSDM2
2023 Medical visual question answering: A survey
Donghao Zhang 0004, Qingyi Tao, Danli Shi, Gholamreza Haffari, Qi Wu 0001, Mingguang He, ZongYuan Ge
Artif. Intell. Medicine5
2023 Influence of context on users' views about explanations for decision-tree predictions
abstract
We consider the influence of two types of contextual information, background information available to users and users’ goals , on users’ views and preferences regarding textual explanations generated for the outcomes predicted by Decision Trees (DTs). To investigate the influence of background information, we generate contrastive explanations that address potential conflicts between aspects of DT predictions and plausible expectations licensed by background information. We define four types of conflicts, operationalize their identification, and specify explanatory schemas that address them. To investigate the influence of users’ goals, we employ an interactive setting where given a goal and an initial explanation for a predicted outcome, users select follow-up questions, and assess the explanations that answer these questions. Here, we offer algorithms to generate explanations that address six types of follow-up questions. The main result from both user studies is that explanations which have a contrastive aspect about a predicted class are generally preferred by users. In addition, the results from the first study indicate that these explanations are deemed especially valuable when users expectations differ from predicted outcomes; and the results from the second study indicate that contrastive explanations which describe how to change a predicted outcome are particularly well regarded in terms of helping users achieve this goal, and they are also popular in terms of helping users achieve other goals.
Sameen Maruf, Ingrid Zukerman, Ehud Reiter, Gholamreza Haffari
Comput. Speech Lang.4
2023 Less is More: Mitigate Spurious Correlations for Open-Domain Dialogue Response Generation Models by Causal Discovery
abstract
Abstract In this paper, we conduct the first study on spurious correlations for open-domain response generation models based on a corpus CGDialog curated by ourselves. The current models indeed suffer from spurious correlations and have a tendency to generate irrelevant and generic responses. Inspired by causal discovery algorithms, we propose a novel model-agnostic method for training and inference using a conditional independence classifier. The classifier is trained by a constrained self-training method, coined ConSTrain, to overcome data sparsity. The experimental results based on both human and automatic evaluation show that our method significantly outperforms the competitive baselines in terms of relevance, informativeness, and fluency.
Tao Feng 0013, Lizhen Qu, Gholamreza Haffari
Trans. Assoc. Comput. Linguistics3
2023 T3L: Translate-and-Test Transfer Learning for Cross-Lingual Text Classification
abstract
Abstract Cross-lingual text classification leverages text classifiers trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning (zero/ few-shots cross-lingual transfer). Nowadays, cross-lingual text classifiers are typically built on large-scale, multilingual language models (LMs) pretrained on a variety of languages of interest. However, the performance of these models varies significantly across languages and classification tasks, suggesting that the superposition of the language modelling and classification tasks is not always effective. For this reason, in this paper we propose revisiting the classic “translate-and-test” pipeline to neatly separate the translation and classification stages. The proposed approach couples 1) a neural machine translator translating from the targeted language to a high-resource language, with 2) a text classifier trained in the high-resource language, but the neural machine translator generates “soft” translations to permit end-to-end backpropagation during fine-tuning of the pipeline. Extensive experiments have been carried out over three cross-lingual text classification datasets (XNLI, MLDoc, and MultiEURLEX), with the results showing that the proposed approach has significantly improved performance over a competitive baseline.
Inigo Jauregi Unanue, Gholamreza Haffari, Massimo Piccardi
Trans. Assoc. Comput. Linguistics2
2023 KC-GEE: knowledge-based conditioning for generative event extraction
abstract
Abstract Event extraction is an important, but challenging task. Many existing techniques decompose it into event and argument detection/classification subtasks, which are complex structured prediction problems. Generation-based extraction techniques lessen the complexity of the problem formulation and are able to leverage the reasoning capabilities of large pretrained language models. However, they still suffer from poor zero-shot generalizability and are ineffective in handling long contexts such as documents. We propose a generative event extraction model, KC-GEE, that addresses these limitations. A key contribution of KC-GEE is a novel knowledge-based conditioning technique that injects the schema of candidate event types as the prefix into each layer of an encoder-decoder language model. This enables effective zero-shot learning and improves supervised learning. Our experiments on two benchmark datasets demonstrate the strong performance of our KC-GEE model. It achieves particularly strong results in the challenging document-level extraction task and in the zero-shot learning setting, outperforming state-of-the-art models by up to 5.4 absolute F1 points.
Tongtong Wu, Fatemeh Shiri, Jingqi Kang, Guilin Qi, Gholamreza Haffari, Yuan-Fang Li
World Wide Web (WWW)5
2022 Teaching Neural Module Networks to Do Arithmetic
abstract
Answering complex questions that require multi-step multi-type reasoning over raw text is challenging, especially when conducting numerical reasoning. Neural Module Networks (NMNs), follow the programmer-interpreter framework and design trainable modules to learn different reasoning skills. However, NMNs only have limited reasoning abilities, and lack numerical reasoning capability. We upgrade NMNs by: (a) bridging the gap between its interpreter and the complex questions; (b) introducing addition and subtraction modules that perform numerical reasoning over numbers. On a subset of DROP, experimental results show that our proposed methods enhance NMNs’ numerical reasoning skills by 17.7% improvement of F1 score and significantly outperform previous state-of-the-art models.
Xiaoyu Guo 0009, Yuan-Fang Li, Gholamreza Haffari
COLING4
2022 Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIs
abstract
Machine-learning-as-a-service (MLaaS) has attracted millions of users to their splendid large-scale models. Although published as black-box APIs, the valuable models behind these services are still vulnerable to imitation attacks. Recently, a series of works have demonstrated that attackers manage to steal or extract the victim models. Nonetheless, none of the previous stolen models can outperform the original black-box APIs. In this work, we conduct unsupervised domain adaptation and multi-victim ensemble to showing that attackers could potentially surpass victims, which is beyond previous understanding of model extraction. Extensive experiments on both benchmark datasets and real-world APIs validate that the imitators can succeed in outperforming the original black-box models on transferred domains. We consider our work as a milestone in the research of imitation attack, especially on NLP APIs, as the superior performance could influence the defense or even publishing strategy of API providers.
Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, Gholamreza Haffari
COLING5
2022 Active Learning by Feature Mixing
abstract
The promise of active learning (AL) is to reduce labelling costs by selecting the most valuable examples to annotate from a pool of unlabelled data. Identifying these examples is especially challenging with high-dimensional data (e.g. images, videos) and in low-data regimes. In this paper, we propose a novel method for batch AL called ALFA-Mix. We identify unlabelled instances with sufficiently-distinct features by seeking inconsistencies in predictions resulting from interventions on their representations. We construct interpolations between representations of labelled and unlabelled instances then examine the predicted labels. We show that inconsistencies in these predictions help discovering features that the model is unable to recognise in the unlabelled instances. We derive an efficient implementation based on a closed-form solution to the optimal interpolation causing changes in predictions. Our method outperforms all recent AL approaches in 30 different settings on 12 benchmarks of images, videos, and non-visual data. The improvements are especially significant in low-data regimes and on self-trained vision transformers, where ALFA-Mix outperforms the state-of-the-art in 59% and 43% of the experiments respectively11The code is available at https://github.com/aminparvaneh/alpha_mix_active_learning.
Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Gholamreza Haffari, Anton van den Hengel, Qinfeng Shi
CVPR4
2022 BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule
abstract
Differentiable Architecture Search (DARTS) has received massive attention in recent years, mainly because it significantly reduces the computational cost through weight sharing and continuous relaxation. However, more recent works find that existing differentiable NAS techniques struggle to outperform naive baselines, yielding deteriorative architectures as the search proceeds. Rather than directly optimizing the architecture parameters, this paper formulates the neural architecture search as a distribution learning problem through relaxing the architecture weights into Gaussian distributions. By leveraging the natural-gradient variational inference (NGVI), the architecture distribution can be easily optimized based on existing codebases without incurring more memory and computational consumption. We demonstrate how the differentiable NAS benefits from Bayesian principles, enhancing exploration and improving stability. The experimental results on NAS benchmark datasets confirm the significant improvements the proposed framework can make. In addition, instead of simply applying the argmax on the learned parameters, we further leverage the recently-proposed training-free proxies in NAS to select the optimal architecture from a group architectures drawn from the optimized distribution, where we achieve state-of-the-art results on the NAS-Bench-201 and NAS-Bench-1shot1 benchmarks. Our best architecture in the DARTS search space also obtains competitive test errors with 2.37%, 15.72%, and 24.2% on CIFAR-10, CIFAR-100, and ImageNet, respectively.
Miao Zhang 0022, Shirui Pan, Xiaojun Chang, Steven Su, Jilin Hu, Gholamreza Haffari, Bin Yang 0002
CVPR6
2022 Multimodal Transformer with Variable-Length Memory for Vision-and-Language Navigation
Chuang Lin 0003, Yi Jiang 0009, Jianfei Cai 0001, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan
ECCV (36)5
2022 Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation
abstract
In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of taskspecific labeled examples.In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space.Both types of priors satisfy a novel property called ϵ-disentangled.We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work.The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings.Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.The code is available at https://github. com/zhuang-li/VAE-DPrior.
Zhuang Li 0001, Lizhen Qu, Qiongkai Xu, Tongtong Wu, Tianyang Zhan, Gholamreza Haffari
EMNLP6
2022 Towards relation extraction from speech
abstract
Relation extraction has focused on extracting semantic relationships between entities from the unstructured written textual data.However, with the vast and rapidly increasing amounts of spoken data, relation extraction from speech is an important but under-explored problem. In this paper, we propose a new information extraction task, speech relation extraction (SpeechRE).To facilitate further research, we construct the first synthetic training datasets, as well as the first human-spoken test set with native English speakers.We establish strong baseline performance for SpeechRE via two approaches.The pipeline approach connects a pretrained ASR module with a text-based relation extraction module.The end-to-end approach employs a cross-modal encoder-decoder architecture.Our comprehensive experiments reveal the relative strengths and weaknesses of these approaches, and shed light on important future directions in SpeechRE research.We share the source code and datasets on https://github.com/ wutong8023/SpeechRE.* denotes the equal contribution.
Tongtong Wu, Guitao Wang, Jinming Zhao, Zhaoran Liu, Guilin Qi, Yuan-Fang Li, Gholamreza Haffari
EMNLP7
2022 Paraphrasing Techniques for Maritime QA system
Fatemeh Shiri, Terry Yue Zhuo, Zhuang Li 0001, Shirui Pan, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002
FUSION6
2022 Pretrained Language Model in Continual Learning: A Comparative Study
Tongtong Wu, Massimo Caccia, Zhuang Li 0001, Yuan-Fang Li, Guilin Qi, Gholamreza Haffari
ICLR6
2022 M-Adapter: Modality Adaptation for End-to-End Speech-to-Text Translation
abstract
End-to-end speech-to-text translation models are often initialized with pre-trained speech encoder and pre-trained text decoder. This leads to a significant training gap between pretraining and fine-tuning, largely due to the modality differences between speech outputs from the encoder and text inputs to the decoder. In this work, we aim to bridge the modality gap between speech and text to improve translation quality. We propose M-Adapter, a novel Transformer-based module, to adapt speech representations to text. While shrinking the speech sequence, M-Adapter produces features desired for speech-to-text translation via modelling global and local dependencies of a speech sequence. Our experimental results show that our model outperforms a strong baseline by up to 1 BLEU score on the Must-C En→DE dataset.
Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi
INTERSPEECH3
2022 Generate, Annotate, and Learn: NLP with Synthetic Text
abstract
Abstract This paper studies the use of language models as a source of synthetic unlabeled text for NLP. We formulate a general framework called “generate, annotate, and learn (GAL)” to take advantage of synthetic text within knowledge distillation, self-training, and few-shot learning applications. To generate high-quality task-specific text, we either fine-tune LMs on inputs from the task of interest, or prompt large LMs with few examples. We use the best available classifier to annotate synthetic text with soft pseudo labels for knowledge distillation and self-training, and use LMs to obtain hard labels for few-shot learning. We train new supervised models on the combination of labeled and pseudo-labeled data, which results in significant gains across several applications. We investigate key components of GAL and present theoretical and empirical arguments against the use of class-conditional LMs to generate synthetic labeled text instead of unlabeled text. GAL achieves new state-of-the-art knowledge distillation results for 6-layer transformers on the GLUE leaderboard.
Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, Mohammad Norouzi 0002
Trans. Assoc. Comput. Linguistics4
2021 Curriculum-Meta Learning for Order-Robust Continual Relation Extraction
abstract
Continual relation extraction is an important task that focuses on extracting new facts incrementally from unstructured text. Given the sequential arrival order of the relations, this task is prone to two serious challenges, namely catastrophic forgetting and order-sensitivity. We propose a novel curriculum-meta learning method to tackle the above two challenges in continual relation extraction. We combine meta learning and curriculum learning to quickly adapt model parameters to a new task and to reduce interference of previously seen tasks on the current task. We design a novel relation representation learning method through the distribution of domain and range types of relations. Such representations are utilized to quantify the difficulty of tasks for the construction of curricula. Moreover, we also present novel difficulty-based metrics to quantitatively measure the extent of order-sensitivity of a given model, suggesting new ways to evaluate model robustness. Our comprehensive experiments on three benchmark datasets show that our proposed method outperforms the state-of-the-art techniques. The code is available at the anonymous GitHub repository: https://github.com/wutong8023/AAAI_CML.
Tongtong Wu, Xuekai Li, Yuan-Fang Li, Gholamreza Haffari, Guilin Qi, Yujin Zhu
AAAI4
2021 Learning to Explain: Generating Stable Explanations Fast
abstract
Xuelin Situ, Ingrid Zukerman, Cecile Paris, Sameen Maruf, Gholamreza Haffari. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf, Gholamreza Haffari
ACL/IJCNLP (1)5
2021 All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-Training
abstract
Pseudo-labeling is a key component in semi-supervised learning (SSL). It relies on iteratively using the model to generate artificial labels for the unlabeled data to train against. A common property among its various methods is that they only rely on the model’s prediction to make labeling decisions without considering any prior knowledge about the visual similarity among the classes. In this paper, we demonstrate that this degrades the quality of pseudo-labeling as it poorly represents visually similar classes in the pool of pseudo-labeled data. We propose SemCo, a method which leverages label semantics and co-training to address this problem. We train two classifiers with two different views of the class labels: one classifier uses the one-hot view of the labels and disregards any potential similarity among the classes, while the other uses a distributed view of the labels and groups potentially similar classes together. We then co-train the two classifiers to learn based on their disagreements. We show that our method achieves state-of-the-art performance across various SSL tasks including 5.6% accuracy improvement on Mini-ImageNet dataset with 1000 labeled examples. We also show that our method requires smaller batch size and fewer training iterations to reach its best performance. We make our code available at https://github.com/islam-nassar/semco.
Islam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine, Gholamreza Haffari
CVPR5
2021 Leveraging Latent Economic Concepts and Sentiments in the News for Market Prediction
abstract
Most of the existing news-based market prediction techniques disregard conceptual and emotional relations in the news stream. In this work, we consider the conceptual relationship between news documents using contextualized latent concept modeling as well as leveraging news sentiment and technical indicators. We present our approach as an open-source RESTFul API. We build a corpus of financial news related to currency pairs in the Foreign Exchange and Cryptocurrencies markets. Next, we apply BERT-based embedding to generate word vectors, cluster the vectors to create latent economic concepts, and propose a document representation based on the distribution of words on these concepts as well as news sentiment. We use a recurrent convolutional neural network to jointly use BERT-based text representation and technical indicators embedding for market time series prediction. We further augment our model with technical indicators using another recurrent layer. The experimental results show the superiority of our method compared to the baselines. Our MarketNews dataset, news crawler, and MarketPredict APIs are available for public use.
Saeede Anbaee Farimani, Majid Vafaei Jahan, Amin Milani Fard, Gholamreza Haffari
DSAA4
2021 Learning Coupled Policies for Simultaneous Machine Translation using Imitation Learning
abstract
We present a novel approach to efficiently learn a simultaneous translation model with coupled programmer-interpreter policies. First, we present an algorithmic oracle to produce oracle READ/WRITE actions for training bilingual sentence-pairs using the notion of word alignments. This oracle actions are designed to capture enough information from the partial input before writing the output. Next, we perform a coupled scheduled sampling to effectively mitigate the exposure bias when learning both policies jointly with imitation learning. Experiments on six language-pairs show our method outperforms strong baselines in terms of translation quality while keeping the translation delay low.
Philip Arthur, Trevor Cohn, Gholamreza Haffari
EACL3
2021 Cognition-aware Cognate Detection
abstract
Diptesh Kanojia, Prashant Sharma, Sayali Ghodekar, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Diptesh Kanojia, Sayali Ghodekar, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
EACL5
2021 Few-Shot Semantic Parsing for New Predicates
abstract
In this work, we investigate the problems of semantic parsing in a few-shot learning setting.In this setting, we are provided with k utterance-logical form pairs per new predicate.The state-of-the-art neural semantic parsers achieve less than 25% accuracy on benchmark datasets when k = 1.To tackle this problem, we proposed to i) apply a designated metalearning method to train the model; ii) regularize attention scores with alignment statistics; iii) apply a smoothing technique in pretraining.As a result, our method consistently outperforms all the baselines in both one and two-shot settings.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
EACL4
2021 Total Recall: a Customized Continual Learning Method for Neural Semantic Parsers
abstract
This paper investigates continual learning for semantic parsing.In this setting, a neural semantic parser learns tasks sequentially without accessing full training data from previous tasks.Direct application of the SOTA continual learning algorithms to this problem fails to achieve comparable performance with retraining models with all seen tasks, because they have not considered the special properties of structured outputs, yielded by semantic parsers.Therefore, we propose TO-TAL RECALL, a continual learning method designed for neural semantic parsers from two aspects: i) a sampling method for memory replay that diversifies logical form templates and balances distributions of parse actions in a memory; ii) a two-stage training method that significantly improves generalization capability of the parsers across tasks.We conduct extensive experiments to study the research problems involved in continual semantic parsing, and demonstrate that a neural semantic parser trained with TOTAL RECALL achieves superior performance than the one trained directly with the SOTA continual learning algorithms, and achieve a 3-6 times speedup compared to retraining from scratch.Code and datasets are available at: https://github. com/zhuang-li/cl_nsp.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
EMNLP (1)3
2021 Lifelong Explainer for Lifelong Learners
abstract
Lifelong Learning (LL) black-box models are dynamic in that they keep learning from new tasks and constantly update their parameters.Owing to the need to utilize information from previously seen tasks, and capture commonalities in potentially diverse data, it is hard for automatic explanation methods to explain the outcomes of these models.In addition, existing explanation methods, e.g., LIME (Ribeiro et al., 2016), which are computationally expensive when explaining a static black-box model, are even more inefficient in the LL setting.In this paper, we propose a novel Lifelong Explanation (LLE) approach that continuously trains a student explainer under the supervision of a teacher -an arbitrary explanation algorithm -on different tasks undertaken in LL.We also leverage the Experience Replay (ER) mechanism to prevent catastrophic forgetting in the student explainer.Our experiments comparing LLE to three baselines on text classification tasks show that LLE can enhance the stability of the explanations for all seen tasks and maintain the same level of faithfulness to the black-box model as the teacher, while being up to 10 2 times faster at test time.Our ablation study shows that the ER mechanism in our LLE approach enhances the learning capabilities of the student explainer.Our code is available at https://github.com/situsnow/LLE.
Xuelin Situ, Sameen Maruf, Ingrid Zukerman, Cécile Paris, Gholamreza Haffari
EMNLP (1)5
2021 Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data Selection
abstract
This paper considers the unsupervised domain adaptation problem for neural machine translation (NMT), where we assume the access to only monolingual text in either the source or target language in the new domain.We propose a cross-lingual data selection method to extract in-domain sentences in the missing language side from a large generic monolingual corpus.Our proposed method trains an adaptive layer on top of multilingual BERT by contrastive learning to align the representation between the source and target language.This then enables the transferability of the domain classifier between the languages in a zero-shot manner.Once the in-domain data is detected by the classifier, the NMT model is then adapted to the new domain by jointly learning translation and domain discrimination tasks.We evaluate our cross-lingual data selection method on NMT across five diverse domains in three language pairs, as well as a real-world scenario of translation for COVID-19.The results show that our proposed method outperforms other selection baselines up to +1.5 BLEU score.
Thuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza Haffari
EMNLP (1)4
2021 Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation Training
abstract
Learning multilingual and multi-domain translation model is challenging as the heterogeneous and imbalanced data make the model converge inconsistently over different corpora in real world.One common practice is to adjust the share of each corpus in the training, so that the learning process is balanced and low-resource cases can benefit from the highresource ones.However, automatic balancing methods usually depend on the intra-and interdataset characteristics, which is usually agnostic or requires human priors.In this work, we propose an approach, MULTIUAT, that dynamically adjusts the training data usage based on the model's uncertainty on a small set of trusted clean data for multi-corpus machine translation.We experiment with two classes of uncertainty measures on multilingual (16 languages with 4 settings) and multi-domain settings (4 for in-domain and 2 for out-of-domain on English-German translation) and demonstrate our approach MULTIUAT substantially outperforms its baselines, including both static and dynamic strategies.We analyze the crossdomain transfer and show the deficiency of static and similarity based methods. 1
Minghao Wu, Meng Zhang 0019, Liangyou Li, Gholamreza Haffari, Qun Liu 0001
EMNLP (1)5
2021 It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation Data
abstract
Most existing simultaneous machine translation (SiMT) systems are trained and evaluated on offline translation corpora.We argue that SiMT systems should be trained and tested on real interpretation data.To illustrate this argument, we propose an interpretation test set and conduct a realistic evaluation of SiMT trained on offline translations.Our results, on our test set along with 3 existing smaller scale language pairs, highlight the difference of up-to 13.83 BLEU score when SiMT models are evaluated on translation vs interpretation data.In the absence of interpretation training data, we propose a translationto-interpretation (T2I) style transfer method which allows converting existing offline translations into interpretation-style data, leading to up-to 2.8 BLEU improvement.However, the evaluation gap remains notable, calling for constructing large-scale interpretation corpora better suited for evaluating and developing SiMT systems. 1
Jinming Zhao, Philip Arthur, Gholamreza Haffari, Trevor Cohn, Ehsan Shareghi
EMNLP (1)3
2021 Toward the Automated Construction of Probabilistic Knowledge Graphs for the Maritime Domain
Fatemeh Shiri, Teresa Wang, Shirui Pan, Xiaojun Chang, Yuan-Fang Li, Gholamreza Haffari, Van Nguyen 0002
FUSION6
2021 iDARTS: Differentiable Architecture Search with Stochastic Implicit Gradients
abstract
Differentiable ARchiTecture Search(DARTS) has recently become the mainstream in the neural architecture search (NAS) due to its efficiency and simplicity. With a gradient-based bi-level optimization, DARTS alternately optimizes the inner model weights and the outer architecture parameter in a weight-sharing supernet. A key challenge to the scalability and quality of the learned architectures is the need for differentiating through the inner-loop optimisation. While much has been discussed about several potentially fatal factors in DARTS, the architecture gradient, a.k.a. hypergradient, has received less attention. In this paper, we tackle the hypergradient computation in DARTS based on the implicit function theorem, making it only depends on the obtained solution to the inner-loop optimization and agnostic to the optimization path. To further reduce the computational requirements, we formulate a stochastic hypergradient approximation for differentiable NAS, and theoretically show that the architecture optimization with the proposed method is expected to converge to a stationary point. Comprehensive experiments on two NAS benchmark search spaces and the common NAS search space verify the effectiveness of our proposed method. It leads to architectures outperforming, with large margins, those learned by the baseline methods.
Miao Zhang 0022, Steven W. Su, Shirui Pan, Xiaojun Chang, Ehsan Abbasnejad, Gholamreza Haffari
ICML6
2021 Explaining Decision-Tree Predictions by Addressing Potential Conflicts between Predictions and Plausible Expectations
abstract
We offer an approach to explain Decision Tree (DT) predictions by addressing potential conflicts between aspects of these predictions and plausible expectations licensed by background information.We define four types of conflicts, operationalize their identification, and specify explanatory schemas that address them.Our human evaluation focused on the effect of explanations on users' understanding of a DT's reasoning and their willingness to act on its predictions.The results show that (1) explanations that address potential conflicts are considered at least as good as baseline explanations that just follow a DT path; and (2) the conflictbased explanations are deemed especially valuable when users' expectations disagree with the DT's predictions.
Sameen Maruf, Ingrid Zukerman, Ehud Reiter, Gholamreza Haffari
INLG4
2020 Reinforcement Learning Based Meta-Path Discovery in Large-Scale Heterogeneous Information Networks
abstract
Meta-paths are important tools for a wide variety of data mining and network analysis tasks in Heterogeneous Information Networks (HINs), due to their flexibility and interpretability to capture the complex semantic relation among objects. To date, most HIN analysis still relies on hand-crafting meta-paths, which requires rich domain knowledge that is extremely difficult to obtain in complex, large-scale, and schema-rich HINs. In this work, we present a novel framework, Meta-path Discovery with Reinforcement Learning (MPDRL), to identify informative meta-paths from complex and large-scale HINs. To capture different semantic information between objects, we propose a novel multi-hop reasoning strategy in a reinforcement learning framework which aims to infer the next promising relation that links a source entity to a target entity. To improve the efficiency, moreover, we develop a type context representation embedded approach to scale the RL framework to handle million-scale HINs. As multi-hop reasoning generates rich meta-paths with various length, we further perform a meta-path induction step to summarize the important meta-paths using Lowest Common Ancestor principle. Experimental results on two large-scale HINs, Yago and NELL, validate our approach and demonstrate that our algorithm not only achieves superior performance in the link prediction task, but also identifies useful meta-paths that would have been ignored by human experts.
Guojia Wan, Bo Du 0001, Shirui Pan, Gholamreza Haffari
AAAI4
2020 Dynamic Programming Encoding for Subword Segmentation in Neural Machine Translation
abstract
This paper introduces Dynamic Programming Encoding (DPE), a new segmentation algorithm for tokenizing sentences into subword units.We view the subword segmentation of output sentences as a latent variable that should be marginalized out for learning and inference.A mixed character-subword transformer is proposed, which enables exact log marginal likelihood estimation and exact MAP inference to find target segmentations with maximum posterior probability.DPE uses a lightweight mixed character-subword transformer as a means of pre-processing parallel data to segment output sentences using dynamic programming.Empirical results on machine translation suggest that DPE is effective for segmenting output sentences and can be combined with BPE dropout for stochastic segmentation of source sentences.DPE achieves an average improvement of 0.9 BLEU over BPE (Sennrich et al., 2016) and an average improvement of 0.55 BLEU over BPE dropout (Provilkov et al., 2019) on several WMT datasets including English ↔ (German, Romanian, Estonian, Finnish, Hungarian).
Xuanli He, Gholamreza Haffari, Mohammad Norouzi 0002
ACL2
2020 Contextual Neural Machine Translation Improves Translation of Cataphoric Pronouns
abstract
The advent of context-aware NMT has resulted in promising improvements in the overall translation quality and specifically in the translation of discourse phenomena such as pronouns.Previous works have mainly focused on the use of past sentences as context with a focus on anaphora translation.In this work, we investigate the effect of future sentences as context by comparing the performance of a contextual NMT model trained with the future context to the one trained with the past context.Our experiments and evaluation, using generic and pronoun-focused automatic metrics, show that the use of future context not only achieves significant improvements over the context-agnostic Transformer, but also demonstrates comparable and in some cases improved performance over its counterpart trained on past context.We also perform an evaluation on a targeted cataphora test suite and report significant gains over the contextagnostic Transformer in terms of BLEU.
KayYen Wong, Sameen Maruf, Gholamreza Haffari
ACL3
2020 Understanding Unnatural Questions Improves Reasoning over Text
abstract
Complex question answering (CQA) over raw text is a challenging task.A prominent approach to this task is based on the programmer-interpreter framework, where the programmer maps the question into a sequence of reasoning actions and the interpreter then executes these actions on the raw text.Learning an effective CQA model requires large amounts of human-annotated data, consisting of the ground-truth sequence of reasoning actions, which is time-consuming and expensive to collect at scale.In this paper, we address the challenge of learning a high-quality programmer (parser) by projecting natural human-generated questions into unnatural machinegenerated questions which are more convenient to parse.We firstly generate synthetic (question, action sequence) pairs by a data generator, and train a semantic parser that associates synthetic questions with their corresponding action sequences.To capture the diversity when applied to natural questions, we learn a projection model to map natural questions into their most similar unnatural questions for which the parser can work well.Without any natural training data, our projection model provides high-quality action sequences for the CQA task.Experimental results show that the QA model trained exclusively with synthetic data outperforms its state-of-the-art counterpart trained on human-labeled data.
Xiaoyu Guo 0009, Yuan-Fang Li, Gholamreza Haffari
COLING3
2020 Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages
abstract
Cognates are variants of the same lexical form across different languages; for example "fonema" in Spanish and "phoneme" in English are cognates, both of which mean "a unit of sound".The task of automatic detection of cognates among any two languages can help downstream NLP tasks such as Cross-lingual Information Retrieval, Computational Phylogenetics, and Machine Translation.In this paper, we demonstrate the use of cross-lingual word embeddings for detecting cognates among fourteen Indian Languages.Our approach introduces the use of context from a knowledge graph to generate improved feature representations for cognate detection.We then evaluate the impact of our cognate detection mechanism on neural machine translation (NMT), as a downstream task.We evaluate our methods to detect cognates on a challenging dataset of twelve Indian languages, namely, Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam.Additionally, we create evaluation datasets for two more Indian languages, Konkani and Nepali 1 .We observe an improvement of up to 18% points, in terms of F-score, for cognate detection.Furthermore, we observe that cognates extracted using our method help improve NMT quality by up to 2.76 BLEU.We also release 2 our code, newly constructed datasets and cross-lingual models publicly.
Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
COLING5
2020 Context Dependent Semantic Parsing: A Survey
abstract
Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations.Currently, most semantic parsing methods are not able to utilize contextual information (e.g.dialogue and comments history), which has a great potential to boost semantic parsing performance.To address this issue, context dependent semantic parsing has recently drawn a lot of attention.In this survey, we investigate progress on the methods for the context dependent semantic parsing, together with the current datasets and tasks.We then point out open problems and challenges for future research in this area.The collected resources for this topic are available at: https://github.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
COLING3
2020 CosMo: Conditional Seq2Seq-based Mixture Model for Zero-Shot Commonsense Question Answering
abstract
Commonsense reasoning refers to the ability of evaluating a social situation and acting accordingly.Identification of the implicit causes and effects of a social context is the driving capability which can enable machines to perform commonsense reasoning.The dynamic world of social interactions requires context-dependent on-demand systems to infer such underlying information.However, current approaches in this realm lack the ability to perform commonsense reasoning upon facing an unseen situation, mostly due to incapability of identifying a diverse range of implicit social relations.Hence they fail to estimate the correct reasoning path.In this paper, we present Conditional SEQ2SEQ-based Mixture model (COSMO), which provides us with the capabilities of dynamic and diverse content generation.We use COSMO to generate context-dependent clauses, which form a dynamic Knowledge Graph (KG) on-the-fly for commonsense reasoning.To show the adaptability of our model to context-dependant knowledge generation, we address the task of zero-shot commonsense question answering.The empirical results indicate an improvement of up to +5.2% over the state-of-the-art models.
Farhad Moghimifar, Lizhen Qu, Terry Yue Zhuo, Mahsa Baktash, Gholamreza Haffari
COLING5
2020 Collective Wisdom: Improving Low-resource Neural Machine Translation using Adaptive Knowledge Distillation
abstract
Scarcity of parallel sentence-pairs poses a significant hurdle for training high-quality Neural Machine Translation (NMT) models in bilingually low-resource scenarios.A standard approach is transfer learning, which involves taking a model trained on a high-resource language-pair and fine-tuning it on the data of the low-resource MT condition of interest.However, it is not clear generally which high-resource language-pair offers the best transfer learning for the target MT setting.Furthermore, different transferred models may have complementary semantic and/or syntactic strengths, hence using only one model may be sub-optimal.In this paper, we tackle this problem using knowledge distillation, where we propose to distill the knowledge of ensemble of teacher models to a single student model.As the quality of these teacher models varies, we propose an effective adaptive knowledge distillation approach to dynamically adjust the contribution of the teacher models during the distillation process.Experiments on transferring from a collection of six language pairs from IWSLT to five low-resource language-pairs from TED Talks demonstrate the effectiveness of our approach, achieving up to +0.9 BLEU score improvement compared to strong baselines.
Fahimeh Saleh, Wray L. Buntine, Gholamreza Haffari
COLING3
2020 Leveraging Discourse Rewards for Document-Level Neural Machine Translation
abstract
Document-level machine translation focuses on the translation of entire documents from a source to a target language.It is widely regarded as a challenging task since the translation of the individual sentences in the document needs to retain aspects of the discourse at document level.However, document-level translation models are usually not trained to explicitly ensure discourse quality.Therefore, in this paper we propose a training approach that explicitly optimizes two established discourse metrics, lexical cohesion (LC) and coherence (COH), by using a reinforcement learning objective.Experiments over four different language pairs and three translation domains have shown that our training approach has been able to achieve more cohesive and coherent document translations than other competitive approaches, yet without compromising the faithfulness to the reference translation.In the case of the Zh-En language pair, our method has achieved an improvement of 2.46 percentage points (pp) in LC and 1.17 pp in COH over the runner-up, while at the same time improving 0.63
Inigo Jauregi Unanue, Nazanin Esmaili, Gholamreza Haffari, Massimo Piccardi
COLING3
2020 Few-Shot Complex Knowledge Base Question Answering via Meta Reinforcement Learning
abstract
Complex question-answering (CQA) involves answering complex natural-language questions on a knowledge base (KB). However, the conventional neural program induction (NPI) approach exhibits uneven performance when the questions have different types, harboring inherently different characteristics, e.g., difficulty level. This paper proposes a meta-reinforcement learning approach to program induction in CQA to tackle the potential distributional bias in questions. Our method quickly and effectively adapts the meta-learned programmer to new questions based on the most similar questions retrieved from the training data. The meta-learned policy is then used to learn a good programming policy, utilizing the trial trajectories and their rewards for similar questions in the support set. Our method achieves state-of-the-art performance on the CQA dataset (Saha et al., 2018) while using only five trial trajectories for the top-5 retrieved questions in each support set, and meta-training on tasks constructed from only 1% of the training set. We have released our code at https://github.com/DevinJake/MRL-CQA.
Yuncheng Hua, Yuan-Fang Li, Gholamreza Haffari, Guilin Qi, Tongtong Wu
EMNLP (1)3
2020 Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
abstract
Recent work has shown the importance of adaptation of broad-coverage contextualised embedding models on the domain of the target task of interest.Current self-supervised adaptation methods are simplistic, as the training signal comes from a small percentage of randomly masked-out tokens.In this paper, we show that careful masking strategies can bridge the knowledge gap of masked language models (MLMs) about the domains more effectively by allocating self-supervision where it is needed.Furthermore, we propose an effective training strategy by adversarially masking out those tokens which are harder to reconstruct by the underlying MLM.The adversarial objective leads to a challenging combinatorial optimisation problem over subsets of tokens, which we tackle efficiently through relaxation to a variational lower-bound and dynamic programming.On six unsupervised domain adaptation tasks involving named entity recognition, our method strongly outperforms the random masking strategy and achieves up to +1.64 F1 score improvements.
Thuy-Trang Vu, Dinh Q. Phung, Gholamreza Haffari
EMNLP (1)3
2020 Personal Information Leakage Detection in Conversations
abstract
The global market size of conversational assistants (chatbots) is expected to grow to USD 9.4 billion by 2024, according to Marketsand-Markets.Despite the wide use of chatbots, leakage of personal information through chatbots poses serious privacy concerns for their users.In this work, we propose to protect personal information by warning users of detected suspicious sentences generated by conversational assistants.The detection task is formulated as an alignment optimization problem and a new dataset PERSONA-LEAKAGE is collected for evaluation.In this paper, we propose two novel constrained alignment models, which consistently outperform baseline methods on PERSONA-LEAKAGE 1 .Moreover, we conduct analysis on the behavior of recently proposed personalized chit-chat dialogue systems.The empirical results show that those systems suffer more from personal information disclosure than the widely used Seq2Seq model and the language model.In those cases, a significant number of information leaking utterances can be detected by our models with high precision.
Qiongkai Xu, Lizhen Qu, Gholamreza Haffari
EMNLP (1)4
2020 Decoding As Dynamic Programming For Recurrent Autoregressive Models
Najam Zaidi, Trevor Cohn, Gholamreza Haffari
ICLR3
2020 Retrieve, Program, Repeat: Complex Knowledge Base Question Answering via Alternate Meta-learning
abstract
A compelling approach to complex question answering is to convert the question to a sequence of actions, which can then be executed on the knowledge base to yield the answer, aka the programmer-interpreter approach. Use similar training questions to the test question, meta-learning enables the programmer to adapt to unseen questions to tackle potential distributional biases quickly. However, this comes at the cost of manually labeling similar questions to learn a retrieval model, which is tedious and expensive. In this paper, we present a novel method that automatically learns a retrieval model alternately with the programmer from weak supervision, i.e., the system’s performance with respect to the produced answers. To the best of our knowledge, this is the first attempt to train the retrieval model with the programmer jointly. Our system leads to state-of-the-art performance on a large-scale task for complex question answering over knowledge bases. We have released our code at https://github.com/DevinJake/MARL.
Yuncheng Hua, Yuan-Fang Li, Gholamreza Haffari, Guilin Qi
IJCAI3
2020 Reasoning Like Human: Hierarchical Reinforcement Learning for Knowledge Graph Reasoning
abstract
Knowledge Graphs typically suffer from incompleteness. A popular approach to knowledge graph completion is to infer missing knowledge by multihop reasoning over the information found along other paths connecting a pair of entities. However, multi-hop reasoning is still challenging because the reasoning process usually experiences multiple semantic issue that a relation or an entity has multiple meanings. In order to deal with the situation, we propose a novel Hierarchical Reinforcement Learning framework to learn chains of reasoning from a Knowledge Graph automatically. Our framework is inspired by the hierarchical structure through which human handle cognitionally ambiguous cases. The whole reasoning process is decomposed into a hierarchy of two-level Reinforcement Learning policies for encoding historical information and learning structured action space. As a consequence, it is more feasible and natural for dealing with the multiple semantic issue. Experimental results show that our proposed model achieves substantial improvements in ambiguous relation tasks.
Guojia Wan, Shirui Pan, Chen Gong 0002, Chuan Zhou 0001, Gholamreza Haffari
IJCAI5
2020 Caption Alignment for Low Resource Audio-Visual Data
abstract
Understanding videos via captioning has gained a lot of traction recently. While captions are provided alongside videos, the information about where a caption aligns within a video is missing, which could be particularly useful for indexing and retrieval. Existing work on learning to infer alignments has mostly exploited visual features and ignored the audio signal. Video understanding applications often underestimate the importance of the audio modality. We focus on how to make effective use of the audio modality for temporal localization of captions within videos. We release a new audio-visual dataset that has captions time-aligned by (i) carefully listening to the audio and watching the video, and (ii) watching only the video. Our dataset is audio-rich and contains captions in two languages, English and Marathi (a low-resource language). We further propose an attention-driven multimodal model, for effective utilization of both audio and video for temporal localization. We then investigate (i) the effects of audio in both data preparation and model design, and (ii) effective pretraining strategies (Audioset, ASR-bottleneck features, PASE, etc.) handling low-resource setting to help extract rich audio representations.
Vighnesh Reddy Konda, Mayur Warialani, Rakesh Prasanth Achari, Varad Bhatnagar, Jayaprakash Akula, Preethi Jyothi, Ganesh Ramakrishnan, Gholamreza Haffari, Pankaj Singh
INTERSPEECH8
2020 Challenge Dataset of Cognates and False Friend Pairs from Indian Languages
abstract
Cognates are present in multiple variants of the same text across different languages (e.g., “hund” in German and “hound” in the English language mean “dog”). They pose a challenge to various Natural Language Processing (NLP) applications such as Machine Translation, Cross-lingual Sense Disambiguation, Computational Phylogenetics, and Information Retrieval. A possible solution to address this challenge is to identify cognates across language pairs. In this paper, we describe the creation of two cognate datasets for twelve Indian languages namely Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam. We digitize the cognate data from an Indian language cognate dictionary and utilize linked Indian language Wordnets to generate cognate sets. Additionally, we use the Wordnet data to create a False Friends’ dataset for eleven language pairs. We also evaluate the efficacy of our dataset using previously available baseline cognate detection approaches. We also perform a manual evaluation with the help of lexicographers and release the curated gold-standard dataset with this paper.
Diptesh Kanojia, Malhar Kulkarni, Pushpak Bhattacharyya, Gholamreza Haffari
LREC4
2020 SummPip: Unsupervised Multi-Document Summarization with Sentence Graph Compression
abstract
Obtaining training data for multi-document Summarization (MDS) is time consuming and resource-intensive, so recent neural models can only be trained for limited domains. In this paper, we propose SummPip: an unsupervised method for multi-document summarization, in which we convert the original documents to a sentence graph, taking both linguistic and deep representation into account, then apply spectral clustering to obtain multiple clusters of sentences, and finally compress each cluster to generate the final summary. Experiments on Multi-News and DUC-2004 datasets show that our method is competitive to previous unsupervised methods and is even comparable to the neural supervised approaches. In addition, human evaluation shows our system produces consistent and complete summaries compared to human written ones.
Jinming Zhao, Ming Liu 0028, Longxiang Gao, Lan Du 0002, He Zhao 0001, He Zhang 0034, Gholamreza Haffari
SIGIR8
2020 A comparative study of data-dependent approaches without learning in measuring similarities of data objects
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari
Data Min. Knowl. Discov.4
2019 Learning How to Active Learn by Dreaming
abstract
Heuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.Recent data-driven AL policy learning methods are also restricted to learn from closely related domains.We introduce a new sample-efficient method that learns the AL policy directly on the target domain of interest by using wake and dream cycles.Our approach interleaves between querying the annotation of the selected datapoints to update the underlying student learner and improving AL policy using simulation where the current student learner acts as an imperfect annotator.We evaluate our method on cross-domain and cross-lingual text classification and named entity recognition tasks.Experimental results show that our dream-based AL policy training strategy is more effective than applying the pretrained policy without further fine-tuning, and better than the existing strong baseline methods that use heuristics or reinforcement learning.
Thuy-Trang Vu, Ming Liu 0028, Dinh Q. Phung, Gholamreza Haffari
ACL (1)4
2019 Utilizing Wordnets for Cognate Detection among Indian Languages
abstract
Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics.Unidentified cognate pairs can pose a challenge to these applications and result in a degradation of performance.In this paper, we detect cognate word pairs among ten Indian languages with Hindi and use deep learning methodologies to predict whether a word pair is cognate or not.We identify IndoWordnet as a potential resource to detect cognate word pairs based on orthographic similarity-based methods and train neural network models using the data obtained from it.We identify parallel corpora as another potential resource and perform the same experiments for them.We also validate the contribution of Wordnets through further experimentation and report improved performance of up to 26%.We discuss the nuances of cognate detection among closely related Indian languages and release the lists of detected cognates as a dataset.We also observe the behaviour of, to an extent, unrelated Indian language pairs and release the lists of detected cognates among them as well.
Diptesh Kanojia, Kevin Patel, Malhar Kulkarni, Pushpak Bhattacharyya, Gholamreza Haffari
GWC5
2019 Twenty years of bioinformatics research for protease-specific substrate and cleavage site prediction: a comprehensive revisit and benchmarking of existing methods
abstract
The roles of proteolytic cleavage have been intensively investigated and discussed during the past two decades. This irreversible chemical process has been frequently reported to influence a number of crucial biological processes (BPs), such as cell cycle, protein regulation and inflammation. A number of advanced studies have been published aiming at deciphering the mechanisms of proteolytic cleavage. Given its significance and the large number of functionally enriched substrates targeted by specific proteases, many computational approaches have been established for accurate prediction of protease-specific substrates and their cleavage sites. Consequently, there is an urgent need to systematically assess the state-of-the-art computational approaches for protease-specific cleavage site prediction to further advance the existing methodologies and to improve the prediction performance. With this goal in mind, in this article, we carefully evaluated a total of 19 computational methods (including 8 scoring function-based methods and 11 machine learning-based methods) in terms of their underlying algorithm, calculated features, performance evaluation and software usability. Then, extensive independent tests were performed to assess the robustness and scalability of the reviewed methods using our carefully prepared independent test data sets with 3641 cleavage sites (specific to 10 proteases). The comparative experimental results demonstrate that PROSPERous is the most accurate generic method for predicting eight protease-specific cleavage sites, while GPS-CCD and LabCaS outperformed other predictors for calpain-specific cleavage sites. Based on our review, we then outlined some potential ways to improve the prediction performance and ease the computational burden by applying ensemble learning, deep learning, positive unlabeled learning and parallel and distributed computing techniques. We anticipate that our study will serve as a practical and useful guide for interested readers to further advance next-generation bioinformatics tools for protease-specific cleavage site prediction.
Fuyi Li, Yanan Wang 0003, Chen Li 0021, Tatiana T. Marquez-Lago, André Leier, Neil D. Rawlings, Gholamreza Haffari, Jerico Revote, Tatsuya Akutsu, Kuo-Chen Chou, Anthony W. Purcell, Robert N. Pike, Geoffrey I. Webb, Alexander Ian Smith, Trevor Lithgow, Roger J. Daly, James C. Whisstock, Jiangning Song
Briefings Bioinform.7
2018 Graph-to-Sequence Learning using Gated Graph Neural Networks
abstract
Many NLP applications can be framed as a graph-to-sequence learning problem. Previous work proposing neural architectures on graph-to-sequence obtained promising results compared to grammar-based approaches but still rely on linearisation heuristics and/or standard recurrent networks to achieve the best performance. In this work propose a new model that encodes the full structural information contained in the graph. Our architecture couples the recently proposed Gated Graph Neural Networks with an input transformation that allows nodes and edges to have their own hidden representations, while tackling the parameter explosion problem present in previous work. Experimental results shows that our model outperforms strong baselines in generation from AMR graphs and syntax-based neural machine translation.
Daniel Beck, Gholamreza Haffari, Trevor Cohn
ACL (1)2
2018 Learning How to Actively Learn: A Deep Imitation Learning Approach
abstract
Heuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.We introduce a method that learns an AL policy using imitation learning (IL).Our IL-based approach makes use of an efficient and effective algorithmic expert, which provides the policy learner with good actions in the encountered AL situations.The AL strategy is then learned with a feedforward network, mapping situations to most informative query datapoints.We evaluate our method on two different tasks: text classification and named entity recognition.Experimental results show that our IL-based AL strategy is more effective than strong previous methods using heuristics and reinforcement learning.
Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari
ACL (1)3
2018 Document Context Neural Machine Translation with Memory Networks
abstract
We present a document-level neural machine translation model which takes both source and target document context into account using memory networks.We model the problem as a structured prediction problem with interdependencies among the observed and hidden variables, i.e., the source sentences and their unobserved target translations in the document.The resulting structured prediction problem is tackled with a neural translation model equipped with two memory components, one each for the source and target side, to capture the documental interdependencies.We train the model endto-end, and propose an iterative decoding algorithm based on block coordinate descent.Experimental results of English translations from French, German, and Estonian documents show that our model is effective in exploiting both source and target document context, and statistically significantly outperforms the previous work in terms of BLEU and METEOR.
Sameen Maruf, Gholamreza Haffari
ACL (1)2
2018 Incorporating Syntactic Uncertainty in Neural Machine Translation with a Forest-to-Sequence Model
abstract
Incorporating syntactic information in Neural Machine Translation (NMT) can lead to better reorderings, particularly useful when the language pairs are syntactically highly divergent or when the training bitext is not large. Previous work on using syntactic information, provided by top-1 parse trees generated by (inevitably error-prone) parsers, has been promising. In this paper, we propose a forest-to-sequence NMT model to make use of exponentially many parse trees of the source sentence to compensate for the parser errors. Our method represents the collection of parse trees as a packed forest, and learns a neural transducer to translate from the input forest to the target sentence. Experiments on English to German, Chinese and Farsi translation tasks show the superiority of our approach over the sequence-to-sequence and tree-to-sequence neural translation models.
Poorya ZareMoodi, Gholamreza Haffari
COLING2
2018 Sequence to Sequence Mixture Model for Diverse Machine Translation
abstract
Sequence to sequence (SEQ2SEQ) models often lack diversity in their generated translations.This can be attributed to the limitation of SEQ2SEQ models in capturing lexical and syntactic variations in a parallel corpus resulting from different styles, genres, topics, or ambiguity of the translation process.In this paper, we develop a novel sequence to sequence mixture (S2SMIX) model that improves both translation diversity and quality by adopting a committee of specialized translation models rather than a single translation model.Each mixture component selects its own training dataset via optimization of the marginal loglikelihood, which leads to a soft clustering of the parallel corpus.Experiments on four language pairs demonstrate the superiority of our mixture model compared to a SEQ2SEQ baseline with standard or diversity-boosted beam search.Our mixture model uses negligible additional parameters and incurs no extra computation cost during decoding.
Xuanli He, Gholamreza Haffari, Mohammad Norouzi 0002
CoNLL2
2018 Learning to Actively Learn Neural Machine Translation
abstract
Traditional active learning (AL) methods for machine translation (MT) rely on heuristics.However, these heuristics are limited when the characteristics of the MT problem change due to e.g. the language pair or the amount of the initial bitext.In this paper, we present a framework to learn sentence selection strategies for neural MT.We train the AL query strategy using a high-resource language-pair based on AL simulations, and then transfer it to the lowresource language-pair of interest.The learned query strategy capitalizes on the shared characteristics between the language pairs to make an effective use of the AL budget.Our experiments on three language-pairs confirms that our method is more effective than strong heuristic-based methods in various conditions, including cold-start and warm-start as well as small and extremely small data conditions.
Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari
CoNLL3
2018 Automatic Post-Editing of Machine Translation: A Neural Programmer-Interpreter Approach
abstract
Automated Post-Editing (PE) is the task of automatically correcting common and repetitive errors found in machine translation (MT) output.In this paper, we present a neural programmer-interpreter approach to this task, resembling the way that humans perform postediting using discrete edit operations, which we refer to as programs.Our model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores.
Thuy-Trang Vu, Gholamreza Haffari
EMNLP2
2018 Neural Machine Translation Advised by Statistical Machine Translation: The Case of Farsi-Spanish Bilingually Low-Resource Scenario
abstract
In this paper, we propose a sequence-to-sequence NMT model on Farsi-Spanish bilingually low-resource language pair. We apply effective preprocessing steps specific for Farsi language and optimize the model for both translation and transliteration. We also propose a loss function that enhances the word alignment and consequently improves translation quality.
Benyamin Ahmadnia, Parisa Kordjamshidi, Gholamreza Haffari
ICMLA3
2018 The Context-Dependent Additive Recurrent Neural Net
abstract
Quan Hung Tran, Tuan Lai, Gholamreza Haffari, Ingrid Zukerman, Trung Bui, Hung Bui. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Quan Hung Tran, Tuan Manh Lai, Gholamreza Haffari, Ingrid Zukerman, Trung Bui
NAACL-HLT3
2018 Neural Machine Translation for Bilingually Scarce Scenarios: a Deep Multi-Task Learning Approach
abstract
Poorya Zaremoodi, Gholamreza Haffari. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Poorya ZareMoodi, Gholamreza Haffari
NAACL-HLT2
2018 PROSPERous: high-throughput prediction of substrate cleavage sites for 90 proteases with improved accuracy
abstract
Summary: Proteases are enzymes that specifically cleave the peptide backbone of their target proteins. As an important type of irreversible post-translational modification, protein cleavage underlies many key physiological processes. When dysregulated, proteases' actions are associated with numerous diseases. Many proteases are highly specific, cleaving only those target substrates that present certain particular amino acid sequence patterns. Therefore, tools that successfully identify potential target substrates for proteases may also identify previously unknown, physiologically relevant cleavage sites, thus providing insights into biological processes and guiding hypothesis-driven experiments aimed at verifying protease-substrate interaction. In this work, we present PROSPERous, a tool for rapid in silico prediction of protease-specific cleavage sites in substrate sequences. Our tool is based on logistic regression models and uses different scoring functions and their pairwise combinations to subsequently predict potential cleavage sites. PROSPERous represents a state-of-the-art tool that enables fast, accurate and high-throughput prediction of substrate cleavage sites for 90 proteases. Availability and implementation: http://prosperous.erc.monash.edu/. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jiangning Song, Fuyi Li, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Gholamreza Haffari, Kuo-Chen Chou, Geoffrey I. Webb, Robert N. Pike
Bioinform.6
2017 Efficient Benchmarking of NLP APIs using Multi-armed Bandits
abstract
Comparing NLP systems to select the best one for a task of interest, such as named entity recognition, is critical for practitioners and researchers.A rigorous approach involves setting up a hypothesis testing scenario using the performance of the systems on query documents.However, often the hypothesis testing approach needs to send a large number of document queries to the systems, which can be problematic.In this paper, we present an effective alternative based on the multi-armed bandit (MAB).We propose a hierarchical generative model to represent the uncertainty in the performance measures of the competing systems, to be used by Thompson Sampling to solve the resulting MAB.Experimental results on both synthetic and real data show that our approach requires significantly fewer queries compared to the standard benchmarking technique to identify the best system according to Fmeasure.
Gholamreza Haffari, Tuan Dung Tran, Mark J. Carman
EACL (1)1
2017 A Hierarchical Neural Model for Learning Sequences of Dialogue Acts
abstract
We propose a novel hierarchical Recurrent Neural Network (RNN) for learning sequences of Dialogue Acts (DAs).The input in this task is a sequence of utterances (i.e., conversational contributions) comprising a sequence of tokens, and the output is a sequence of DA labels (one label per utterance).Our model leverages the hierarchical nature of dialogue data by using two nested RNNs that capture long-range dependencies at the dialogue level and the utterance level.This model is combined with an attention mechanism that focuses on salient tokens in utterances.Our experimental results show that our model outperforms strong baselines on two popular datasets, Switchboard and MapTask; and our detailed empirical analysis highlights the impact of each aspect of our model.
Quan Hung Tran, Ingrid Zukerman, Gholamreza Haffari
EACL (1)3
2017 Towards Decoding as Continuous Optimisation in Neural Machine Translation
abstract
We propose a novel decoding approach for neural machine translation (NMT) based on continuous optimisation.We reformulate decoding, a discrete optimization problem, into a continuous problem, such that optimization can make use of efficient gradient-based techniques.Our powerful decoding framework allows for more accurate decoding for standard neural machine translation models, as well as enabling decoding in intractable models such as intersection of several different NMT models.Our empirical results show that our decoding framework is effective, and can leads to substantial improvements in translations, especially in situations where greedy search and beam search are not feasible.Finally, we show how the technique is highly competitive with, and complementary to, reranking.
Cong Duy Vu Hoang, Gholamreza Haffari, Trevor Cohn
EMNLP2
2017 Preserving Distributional Information in Dialogue Act Classification
abstract
This paper introduces a novel training/decoding strategy for sequence labeling.Instead of greedily choosing a label at each time step, and using it for the next prediction, we retain the probability distribution over the current label, and pass this distribution to the next prediction.This approach allows us to avoid the effect of label bias and error propagation in sequence learning/decoding.Our experiments on dialogue act classification demonstrate the effectiveness of this approach.Even though our underlying neural network model is relatively simple, it outperforms more complex neural models, achieving state-of-the-art results on the MapTask and Switchboard corpora.
Quan Hung Tran, Ingrid Zukerman, Gholamreza Haffari
EMNLP3
2017 Compressed Nonparametric Language Modelling
abstract
Hierarchical Pitman-Yor Process priors are compelling for learning language models, outperforming point-estimate based methods. However, these models remain unpopular due to computational and statistical inference issues, such as memory and time usage, as well as poor mixing of sampler. In this work we propose a novel framework which represents the HPYP model compactly using compressed suffix trees. Then, we develop an efficient approximate inference scheme in this framework that has a much lower memory footprint compared to full HPYP and is fast in the inference time. The experimental results illustrate that our model can be built on significantly larger datasets compared to previous HPYP models, while being several orders of magnitudes smaller, fast for training and inference, and outperforming the perplexity of the state-of-the-art Modified Kneser-Ney count-based LM smoothing by up to 15%.
Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn
IJCAI2
2017 Multi-domain evaluation framework for named entity recognition tools
Zahraa Said Abdallah, Mark J. Carman, Gholamreza Haffari
Comput. Speech Lang.3
2017 Data-dependent dissimilarity measure: an effective alternative to geometric distance measures
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari
Knowl. Inf. Syst.4
2016 Improving Word Alignment of Rare Words with Word Embeddings
abstract
We address the problem of inducing word alignment for language pairs by developing an unsupervised model with the capability of getting applied to other generative alignment models. We approach the task by: i)proposing a new alignment model based on the IBM alignment model 1 that uses vector representation of words, and ii)examining the use of similar source words to overcome the problem of rare source words and improving the alignments. We apply our method to English-French corpora and run the experiments with different sizes of sentence pairs. Our results show competitive performance against the baseline and in some cases improve the results up to 6.9% in terms of precision.
Masoud Jalili Sabet, Heshaam Faili, Gholamreza Haffari
COLING3
2016 Richer Interpolative Smoothing Based on Modified Kneser-Ney Language Modeling
abstract
In this work we present a generalisation of the Modified Kneser-Ney interpolative smoothing for richer smoothing via additional discount parameters.We provide mathematical underpinning for the estimator of the new discount parameters, and showcase the utility of our rich MKN language models on several European languages.We further explore the interdependency among the training data size, language model order, and number of discount parameters.Our empirical results illustrate that larger number of discount parameters, i) allows for better allocation of mass in the smoothing process, particularly on small data regime where statistical sparsity is severe, and ii) leads to significant reduction in perplexity, particularly for out-of-domain test sets which introduce higher ratio of out-ofvocabulary words. 1
Ehsan Shareghi, Trevor Cohn, Gholamreza Haffari
EMNLP3
2016 Incorporating Structural Alignment Biases into an Attentional Neural Translation Model
abstract
Neural encoder-decoder models of machine translation have achieved impressive results, rivalling traditional translation models. However their modelling formulation is overly simplistic, and omits several key inductive biases built into traditional models. In this paper we extend the attentional neural translation model to include structural biases from word based alignment models, including positional bias, Markov conditioning, fertility and agreement over translation directions. We show improvements over a baseline attentional model and standard phrase-based model over several language pairs, evaluating on difficult languages in a low resource setting.
Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, Gholamreza Haffari
HLT-NAACL6
2016 Incorporating Side Information into Recurrent Neural Network Language Models
abstract
Cong Duy Vu Hoang, Trevor Cohn, Gholamreza Haffari. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Cong Duy Vu Hoang, Trevor Cohn, Gholamreza Haffari
HLT-NAACL3
2016 A Latent Variable Recurrent Neural Network for Discourse-Driven Language Models
abstract
This paper presents a novel latent variable recurrent neural network architecture for jointly modeling sequences of words and (possibly latent) discourse relations between adjacent sentences.A recurrent neural network generates individual words, thus reaping the benefits of discriminatively-trained vector representations.The discourse relations are represented with a latent variable, which can be predicted or marginalized, depending on the task.The resulting model can therefore employ a training objective that includes not only discourse relation classification, but also word prediction.As a result, it outperforms state-ofthe-art alternatives for two tasks: implicit discourse relation classification in the Penn Discourse Treebank, and dialog act classification in the Switchboard corpus.Furthermore, by marginalizing over latent discourse relations at test time, we obtain a discourse informed language model, which improves over a strong LSTM baseline.
Yangfeng Ji, Gholamreza Haffari, Jacob Eisenstein
HLT-NAACL2
2016 Inter-document Contextual Language model
abstract
Quan Hung Tran, Ingrid Zukerman, Gholamreza Haffari. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Quan Hung Tran, Ingrid Zukerman, Gholamreza Haffari
HLT-NAACL3
2016 Text mining electronic hospital records to automatically classify admissions against disease: Measuring the impact of linking data sources
Simon Kocbek, Lawrence Cavedon, David Martínez 0001, Christopher Bain, Chris Mac Manus, Gholamreza Haffari, Ingrid Zukerman, Karin Verspoor
J. Biomed. Informatics6
2016 Fast, Small and Exact: Infinite-order Language Modelling with Compressed Suffix Trees
abstract
Efficient methods for storing and querying are critical for scaling high-order m-gram language models to large corpora. We propose a language model based on compressed suffix trees, a representation that is highly compact and can be easily held in memory, while supporting queries needed in computing language model probabilities on-the-fly. We present several optimisations which improve query runtimes up to 2500×, despite only incurring a modest increase in construction time and memory usage. For large corpora and high Markov orders, our method is highly competitive with the state-of-the-art KenLM package. It imposes much lower memory requirements, often by orders of magnitude, and has runtimes that are either similar (for training) or comparable (for querying).
Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn
Trans. Assoc. Comput. Linguistics3
2015 Compact, Efficient and Unlimited Capacity: Language Modeling with Compressed Suffix Trees
abstract
Efficient methods for storing and querying language models are critical for scaling to large corpora and high Markov orders.In this paper we propose methods for modeling extremely large corpora without imposing a Markov condition.At its core, our approach uses a succinct index -a compressed suffix tree -which provides near optimal compression while supporting efficient search.We present algorithms for on-the-fly computation of probabilities under a Kneser-Ney language model.Our technique is exact and although slower than leading LM toolkits, it shows promising scaling properties, which we demonstrate through ∞-order modeling over the full Wikipedia collection.
Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn
EMNLP3
2015 Optimizing Multivariate Performance Measures for Learning Relation Extraction Models
abstract
Gholamreza Haffari, Ajay Nagesh, Ganesh Ramakrishnan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Gholamreza Haffari, Ajay Nagesh, Ganesh Ramakrishnan
HLT-NAACL1
2015 Structured Prediction of Sequences and Trees Using Infinite Contexts
Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn, Ann E. Nicholson
ECML/PKDD (2)2
2015 Novel Bernstein-like Concentration Inequalities for the Missing Mass
Bahman Yari Saeed Khanloo, Gholamreza Haffari
UAI2
2015 Half-space mass: a maximally robust and efficient data depth method
Bo Chen 0009, Kai Ming Ting, Takashi Washio, Gholamreza Haffari
Mach. Learn.4
2014 Noisy Or-based model for Relation Extraction using Distant Supervision
abstract
Distant supervision, a paradigm of relation extraction where training data is created by aligning facts in a database with a large unannotated corpus, is an attractive approach for training relation extractors.Various models are proposed in recent literature to align the facts in the database to their mentions in the corpus.In this paper, we discuss and critically analyse a popular alignment strategy called the "at least one" heuristic.We provide a simple, yet effective relaxation to this strategy.We formulate the inference procedures in training as integer linear programming (ILP) problems and implement the relaxation to the "at least one " heuristic via a soft constraint in this formulation.Empirically, we demonstrate that this simple strategy leads to a better performance under certain settings over the existing approaches.
Ajay Nagesh, Gholamreza Haffari, Ganesh Ramakrishnan
EMNLP2
2014 Mp-Dissimilarity: A Data Dependent Dissimilarity Measure
abstract
Nearest neighbour search is a core process in many data mining algorithms. Finding reliable closest matches of a query in a high dimensional space is still a challenging task. This is because the effectiveness of many dissimilarity measures, that are based on a geometric model, such as lp-norm, decreases as the number of dimensions increases. In this paper, we examine how the data distribution can be exploited to measure dissimilarity between two instances and propose a new data dependent dissimilarity measure called 'mp-dissimilarity'. Rather than relying on geometric distance, it measures the dissimilarity between two instances in each dimension as a probability mass in a region that encloses the two instances. It deems the two instances in a sparse region to be more similar than two instances in a dense region, though these two pairs of instances have the same geometric distance. Our empirical results show that the proposed dissimilarity measure indeed provides a reliable nearest neighbour search in high dimensional spaces, particularly in sparse data. Mp-dissimilarity produced better task specific performance than lp-norm and cosine distance in classification and information retrieval tasks.
Sunil Aryal, Kai Ming Ting, Gholamreza Haffari, Takashi Washio
ICDM3
2014 HIT'nDRIVE: Multi-driver Gene Prioritization Based on Hitting Time
Raunak Shrestha, Ermin Hodzic, Jake Yeung, Kendric Wang, Thomas Sauerwald, Phuong Dao, Shawn Anderson, Himisha Beltran, Mark A. Rubin, Colin C. Collins, Gholamreza Haffari, Süleyman Cenk Sahinalp
RECOMB11
2013 An Infinite Hierarchical Bayesian Model of Phrasal Translation
Trevor Cohn, Gholamreza Haffari
ACL (1)2
2013 The Haves and the Have-Nots: Leveraging Unlabelled Corpora for Sentiment Analysis
Kashyap Popat, A. R. Balamurali, Pushpak Bhattacharyya, Gholamreza Haffari
ACL (1)4
2013 Graph Propagation for Paraphrasing Out-of-Vocabulary Words in Statistical Machine Translation
Majid Razmara, Maryam Siahbani, Gholamreza Haffari, Anoop Sarkar
ACL (1)3
2013 Scalable Variational Inference for Extracting Hierarchical Phrase-based Translation Rules
Baskaran Sankaran, Gholamreza Haffari, Anoop Sarkar
IJCNLP2
2012 Feature-based classifiers for somatic mutation detection in tumour-normal paired sequencing data
abstract
MOTIVATION: The study of cancer genomes now routinely involves using next-generation sequencing technology (NGS) to profile tumours for single nucleotide variant (SNV) somatic mutations. However, surprisingly few published bioinformatics methods exist for the specific purpose of identifying somatic mutations from NGS data and existing tools are often inaccurate, yielding intolerably high false prediction rates. As such, the computational problem of accurately inferring somatic mutations from paired tumour/normal NGS data remains an unsolved challenge. RESULTS: We present the comparison of four standard supervised machine learning algorithms for the purpose of somatic SNV prediction in tumour/normal NGS experiments. To evaluate these approaches (random forest, Bayesian additive regression tree, support vector machine and logistic regression), we constructed 106 features representing 3369 candidate somatic SNVs from 48 breast cancer genomes, originally predicted with naive methods and subsequently revalidated to establish ground truth labels. We trained the classifiers on this data (consisting of 1015 true somatic mutations and 2354 non-somatic mutation positions) and conducted a rigorous evaluation of these methods using a cross-validation framework and hold-out test NGS data from both exome capture and whole genome shotgun platforms. All learning algorithms employing predictive discriminative approaches with feature selection improved the predictive accuracy over standard approaches by statistically significant margins. In addition, using unsupervised clustering of the ground truth 'false positive' predictions, we noted several distinct classes and present evidence suggesting non-overlapping sources of technical artefacts illuminating important directions for future study. AVAILABILITY: Software called MutationSeq and datasets are available from http://compbio.bccrc.ca.
Jiarui Ding, Ali Bashashati, Andrew Roth, Arusha Oloumi, Kane Tse, Thomas Zeng 0002, Gholamreza Haffari, Martin Hirst, Marco A. Marra, Anne Condon, Samuel Aparicio, Sohrab P. Shah
Bioinform.7
2011 Modeling the temporal dynamics of social rating networks using bidirectional effects of social relations and rating patterns
abstract
A social rating network (SRN) is a social network in which edges represent social relationships and users (nodes) express ratings on some of the given items. Such networks play an increasingly important role in reviewing websites such as Epinions.com or online sharing websites like Flickr.com. In this paper, we first observe and analyze the temporal behavior of users in a social rating network, who express ratings and create social relations. Then, we model the temporal dynamics of an SRN based on our observations, using the bidirectional effects of ratings and social relations. While existing models for other types of social networks have captured some of the effects, our model is the first one to represent all four effects, i.e. social relations-on-ratings (social influence), social relations-on-social relations (transitivity), ratings-on-social relations (selection), and ratings-on-ratings (correlational influence). Existing works consider these effects as static and constant throughout the evolution of an SRN, however our observations reveal that these effects are actually dynamic. We propose a probabilistic generative model for SRNs, which models the strength and dynamics of each effect throughout the network evolution. This model can serve for the prediction of future links, ratings or community structures. Due to the sensitive nature of SRNs, another motivation for our work is the generation of synthetic SRN data sets for research purposes. Our experimental studies on two real life datasets (Epinions and Flickr) demonstrate that the proposed model produces social rating networks that agree with real world data on a comprehensive set of evaluation criteria.
Mohsen Jamali, Gholamreza Haffari, Martin Ester
WWW2
2009 Active Learning for Multilingual Statistical Machine Translation
Gholamreza Haffari, Anoop Sarkar
ACL/IJCNLP1
2009 Active Learning for Statistical Phrase-based Machine Translation
Gholamreza Haffari, Maxim Roy, Anoop Sarkar
HLT-NAACL1
2009 Hierarchical Dirichlet Trees for Information Retrieval
Gholamreza Haffari, Yee Whye Teh
HLT-NAACL1
2009 A Rate Distortion Approach for Semi-Supervised Conditional Random Fields
abstract
We propose a novel information theoretic approach for semi-supervised learning of conditional random fields. Our approach defines a training objective that combines the conditional likelihood on labeled data and the mutual information on unlabeled data. Different from previous minimum conditional entropy semi-supervised discriminative learning methods, our approach can be naturally cast into the rate distortion theory framework in information theory. We analyze the tractability of the framework for structured prediction and present a convergent variational training algorithm to defy the combinatorial explosion of terms in the sum over label configurations. Our experimental results show that the rate distortion approach outperforms standard $l_2$ regularization and minimum conditional entropy regularization on both multi-class classification and sequence labeling problems.
Yang Wang 0003, Gholamreza Haffari, Greg Mori
NIPS2
2008 Homotopy-Based Semi-Supervised Hidden Markov Models for Sequence Labeling
Gholamreza Haffari, Anoop Sarkar
COLING1
2008 Boosting with incomplete information
abstract
In real-world machine learning problems, it is very common that part of the input feature vector is incomplete: either not available, missing, or corrupted. In this paper, we present a boosting approach that integrates features with incomplete information and those with complete information to form a strong classifier. By introducing hidden variables to model missing information, we form loss functions that combine fully labeled data with partially labeled data to effectively learn normalized and unnormalized models. The primal problems of the proposed optimization problems with these loss functions are provided to show their close relationship and the motivations behind them. We use auxiliary functions to bound the change of the loss functions and derive explicit parameter update rules for the learning algorithms. We demonstrate encouraging results on two real-world problems --- visual object recognition in computer vision and named entity recognition in natural language processing --- to show the effectiveness of the proposed boosting approach.
Gholamreza Haffari, Yang Wang 0003, Greg Mori, Feng Jiao
ICML1
2007 Transductive learning for statistical machine translation
Nicola Ueffing, Gholamreza Haffari, Anoop Sarkar
ACL2
2007 Analysis of Semi-Supervised Learning with the Yarowsky Algorithm
Gholamreza Haffari, Anoop Sarkar
UAI1
2007 Semi-supervised model adaptation for statistical machine translation
Nicola Ueffing, Gholamreza Haffari, Anoop Sarkar
Mach. Transl.2
2006 Tutorial on Inductive Semi-supervised Learning Methods: with Applicability to Natural Language Processing
Anoop Sarkar, Gholamreza Haffari
HLT-NAACL2
2002 Analysis of STAGE Algorithm Based on Solving Bin Packing Problem
Gholamreza Haffari, Saeed Bagheri Shouraki
ICMLA1