VLDB 2026 Research / reviewers in the wild / expert
Hadi Amiri
dblp:41/7403
· DBLP profile ↗
30ranked-venue papers
12as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FUTURE: Flexible Unlearning for Tree EnsembleabstractTree ensembles are widely recognized for their effectiveness in classification tasks, achieving state-of-the-art performance across diverse domains, including bioinformatics, finance, and medical diagnosis. With increasing emphasis on data privacy and the right to be forgotten, several unlearning algorithms have been proposed to enable tree ensembles to forget sensitive information. However, existing methods are often tailored to a particular model or rely on the discrete tree structure, making them difficult to generalize to complex ensembles and inefficient for large-scale datasets. To address these limitations, we propose FUTURE, a novel unlearning algorithm for tree ensembles. Specifically, we formulate the problem of forgetting samples as a gradient-based optimization task. In order to accommodate non-differentiability of tree ensembles, we adopt the probabilistic model approximations within the optimization framework. This enables end-to-end unlearning in an effective and efficient manner. Extensive experiments on real-world datasets show that FUTURE yields significant and successful unlearning performance. Ziheng Chen 0002, Jin Huang 0010, Jiali Cheng, Yuchan Guo, Lalitesh Morishetti, Kaushiki Nag, Hadi Amiri |
CIKM | 8 |
| 2025 | FROG: Fair Removal on GraphabstractWith growing emphasis on privacy regulations, machine unlearning has become increasingly critical in real-world applications such as social networks and recommender systems, many of which are naturally represented as graphs. However, existing graph unlearning methods often modify nodes or edges indiscriminately, overlooking their impact on fairness. For instance, forgetting links between users of different genders may inadvertently exacerbate group disparities. To address this issue, we propose a novel framework that jointly optimizes both the graph structure and the model to achieve fair unlearning. Our method rewires the graph by removing redundant edges that hinder forgetting while preserving fairness through targeted edge augmentation. We further introduce a worst-case evaluation mechanism to assess robustness under challenging scenarios. Experiments on real-world datasets show that our approach achieves more effective and fair unlearning than existing baselines. Ziheng Chen 0002, Jiali Cheng, Hadi Amiri, Kaushiki Nag, Lu Lin 0001, Sijia Liu 0001, Gabriele Tolomei, Xiangguo Sun |
CIKM | 3 |
| 2025 | Unveil Multi-Picture Descriptions for Multilingual Mild Cognitive Impairment Detection via Contrastive LearningabstractDetecting Mild Cognitive Impairment from picture descriptions is critical yet challenging, especially in multilingual and multiple picture settings. Prior work has primarily focused on English speakers describing a single picture (e.g., the 'Cookie Theft'). The TAUKDIAL-2024 challenge expands this scope by introducing multilingual speakers and multiple pictures, which presents new challenges in analyzing picture-dependent content. To address these challenges, we propose a framework with three components: (1) enhancing discriminative representation learning via supervised contrastive learning, (2) involving image modality rather than relying solely on speech and text modalities, and (3) applying a Product of Experts (PoE) strategy to mitigate spurious correlations and overfitting. Our framework improves MCI detection performance, achieving a +7.1% increase in Unweighted Average Recall (UAR) (from 68.1% to 75.2%) and a +2.9% increase in F1 score (from 80.6% to 83.5%) compared to the text unimodal baseline. Notably, the contrastive learning component yields greater gains for the text modality compared to speech. These results highlight our framework's effectiveness in multilingual and multi-picture MCI detection. Kristin Qi, Jiali Cheng, Youxiang Zhu, Hadi Amiri, Xiaohui Liang 0002 |
GLOBECOM | 4 |
| 2025 | Tool Unlearning for Tool-Augmented LLMsabstractTool-augmented large language models (LLMs) may need to forget learned tools due to security concerns, privacy restrictions, or deprecated tools. However, “tool unlearning” has not been investigated in machine unlearning literature. We introduce this novel task, which requires addressing distinct challenges compared to traditional unlearning: knowledge removal rather than forgetting individual samples, the high cost of optimizing LLMs, and the need for principled evaluation metrics. To bridge these gaps, we propose ToolDelete , the first approach for unlearning tools from tool-augmented LLMs which implements three properties for effective tool unlearning, and a new membership inference attack (MIA) model for evaluation. Experiments on three tool learning datasets and tool-augmented LLMs show that ToolDelete effectively unlearns both randomly selected and category-specific tools, while preserving the LLM’s knowledge on non-deleted tools and maintaining performance on general tasks. Jiali Cheng, Hadi Amiri |
ICML | 2 |
| 2025 | Speech Unlearning
Jiali Cheng, Hadi Amiri |
INTERSPEECH | 2 |
| 2024 | MultiDelete for Multimodal Machine Unlearning
Jiali Cheng, Hadi Amiri |
ECCV (41) | 2 |
| 2024 | FairFlow: Mitigating Dataset Biases through Undecided Learning for Natural Language UnderstandingabstractLanguage models are prone to dataset biases, known as shortcuts and spurious correlations in data, which often result in performance drop on new data.We present a new debiasing framework called "FAIRFLOW" that mitigates dataset biases by learning to be undecided in its predictions for data samples or representations associated with known or unknown biases.The framework introduces two key components: a suite of data and model perturbation operations that generate different biased views of input samples, and a contrastive objective that learns debiased and robust representations from the resulting biased views of samples.Experiments show that FAIRFLOW outperforms existing debiasing methods, particularly against out-ofdomain and hard test samples without compromising the in-domain performance 1 . Jiali Cheng, Hadi Amiri |
EMNLP | 2 |
| 2024 | CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
Jiali Cheng, Mohamed Elgaar, Nidhi Vakil, Hadi Amiri |
INTERSPEECH | 4 |
| 2023 | HuCurl: Human-induced Curriculum DiscoveryabstractWe introduce the problem of curriculum discovery and describe a curriculum learning framework capable of discovering effective curricula in a curriculum space based on prior knowledge about sample difficulty.Using annotation entropy and loss as measures of difficulty, we show that (i): the top-performing discovered curricula for a given model and dataset are often non-monotonic as apposed to monotonic curricula in existing literature, (ii): the prevailing easy-to-hard or hard-to-easy transition curricula are often at the risk of underperforming, and (iii): the curricula discovered for smaller datasets and models perform well on larger datasets and models respectively.The proposed framework encompasses some of the existing curriculum learning approaches and can discover curricula that outperform them across several NLP tasks. Mohamed Elgaar, Hadi Amiri |
ACL (1) | 2 |
| 2023 | Curriculum Learning for Graph Neural Networks: A Multiview Competence-based ApproachabstractA curriculum is a planned sequence of learning materials and an effective one can make learning efficient and effective for both humans and machines.Recent studies developed effective data-driven curriculum learning approaches for training graph neural networks in language applications.However, existing curriculum learning approaches often employ a single criterion of difficulty in their training paradigms.In this paper, we propose a new perspective on curriculum learning by introducing a novel approach that builds on graph complexity formalisms (as difficulty criteria) and model competence during training.The model consists of a scheduling scheme which derives effective curricula by accounting for different views of sample difficulty and model competence during training.The proposed solution advances existing research in curriculum learning for graph neural networks with the ability to incorporate a fine-grained spectrum of graph difficulty criteria in their training paradigms.Experimental results on real-world link prediction and node classification tasks illustrate the effectiveness of the proposed approach.1 Nidhi Vakil, Hadi Amiri |
ACL (1) | 2 |
| 2023 | Ling-CL: Understanding NLP Models through Linguistic CurriculaabstractWe employ a characterization of linguistic complexity from psycholinguistic and language acquisition research to develop data-driven curricula to understand the underlying linguistic knowledge that models learn to address NLP tasks.The novelty of our approach is in the development of linguistic curricula derived from data, existing knowledge about linguistic complexity, and model behavior during training.By analyzing several benchmark NLP datasets, our curriculum learning approaches identify sets of linguistic metrics (indices) that inform the challenges and reasoning required to address each task.Our work will inform future research in all NLP areas, allowing linguistic complexity to be considered early in the research and development process.In addition, our work prompts an examination of gold standards and fair evaluation in NLP. Mohamed Elgaar, Hadi Amiri |
EMNLP | 2 |
| 2022 | Generic and Trend-aware Curriculum Learning for Relation ExtractionabstractWe present a generic and trend-aware curriculum learning approach for graph neural networks.It extends existing approaches by incorporating sample-level loss trends to better discriminate easier from harder samples and schedule them for training.The model effectively integrates textual and structural information for relation extraction in text graphs.Experimental results show that the model provides robust estimations of sample difficulty and shows sizable improvement over the stateof-the-art approaches across several datasets. Nidhi Vakil, Hadi Amiri |
NAACL-HLT | 2 |
| 2022 | A Vector Quantization-Based Spike Compression Approach Dedicated to Multichannel Neural Recording MicrosystemsabstractImplantable high-density multichannel neural recording microsystems provide simultaneous recording of brain activities. Wireless transmission of the entire recorded data causes high bandwidth usage, which is not tolerable for implantable applications. As a result, a hardware-friendly compression module is required to reduce the amount of data before it is transmitted. This paper presents a novel compression approach that utilizes a spike extractor and a vector quantization (VQ)-based spike compressor. In this approach, extracted spikes are vector quantized using an unsupervised learning process providing a high spike compression ratio (CR) of 10–80. A combination of extracting and compressing neural spikes results in a significant data reduction as well as preserving the spike waveshapes. The compression performance of the proposed approach was evaluated under variant conditions. We also developed new architectures such that the hardware blocks of our approach can be implemented more efficiently. The compression module was implemented in a 180-nm standard CMOS process achieving a SNDR of 14.49[Formula: see text]dB and a classification accuracy (CA) of 99.62% at a CR of 20, while consuming 4[Formula: see text][Formula: see text]W power and 0.16[Formula: see text]mm2 chip area per channel. Nazanin Ahmadi-Dastgerdi, Hossein Hosseini-Nejad, Hadi Amiri, Afshin Shoeibi, Juan Manuel Górriz |
Int. J. Neural Syst. | 3 |
| 2021 | A novel bearing-only localization for generalized Gaussian noise
Mostafa Naseri, Hadi Amiri |
Signal Process. | 2 |
| 2019 | Learning to Estimate Nutrition Facts from Food Descriptions
Hadi Amiri, Andrew L. Beam, Isaac S. Kohane |
AMIA | 1 |
| 2018 | Toward Large-scale and Multi-facet Analysis of First Person Alcohol Drinking
Hadi Amiri, Kara M. Magane, Lauren E. Wisk, Guergana K. Savova, Elissa R. Weitzman |
AMIA | 1 |
| 2018 | Spotting Spurious Data with Neural NetworksabstractHadi Amiri, Timothy Miller, Guergana Savova. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
NAACL-HLT | 1 |
| 2017 | Repeat before Forgetting: Spaced Repetition for Efficient and Effective Training of Neural NetworksabstractWe present a novel approach for training artificial neural networks.Our approach is inspired by broad evidence in psychology that shows human learners can learn efficiently and effectively by increasing intervals of time between subsequent reviews of previously learned materials (spaced repetition).We investigate the analogy between training neural models and findings in psychology about human memory model and develop an efficient and effective algorithm to train neural models.The core part of our algorithm is a cognitively-motivated scheduler according to which training instances and their "reviews" are spaced over time.Our algorithm uses only 34-50% of data per epoch, is 2.9-4.8 times faster than standard training, and outperforms competing state-of-the-art baselines.1 Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
EMNLP | 1 |
| 2016 | Short Text Representation for Detecting Churn in MicroblogsabstractChurn happens when a customer leaves a brand or stop using its services. Brands reduce their churn rates by identifying and retaining potential churners through customer retention campaigns. In this paper, we consider the problem of classifying micro-posts as churny or non-churny with respect to a given brand. Motivated by the recent success of recurrent neural networks (RNNs) in word representation, we propose to utilize RNNs to learn micro-post and churn indicator representations. We show that such representations improve the performance of churn detection in microblogs and lead to more accurate ranking of churny contents. Furthermore, in this researchwe show that state-of-the-art sentiment analysis approaches fail to identify churny contents. Experiments on Twitter data about three telco brands show the utility of our approach for this task. Hadi Amiri, Hal Daumé III |
AAAI | 1 |
| 2016 | Learning Text Pair Similarity with Context-sensitive Autoencoders
Hadi Amiri, Philip Resnik, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 1 |
| 2015 | Target-Dependent Churn Classification in MicroblogsabstractWe consider the problem of classifying micro-posts as churny or non-churny with respect to a given brand. Using Twitter data about three brands, we find that standard machine learning techniques clearly outperform keyword based approaches. However, the three machine learning techniques we employed (linear classification, support vector machines, and logistic regression) do not perform as well on churn classification as on other text classification problems. We investigate demographic, content, and context churn indicators in microblogs and examine factors that make this problem more challenging. Experimental results show an average F1 performance of 75% for target-dependent churn classification in microblogs. Hadi Amiri, Hal Daumé III |
AAAI | 1 |
| 2013 | A Pattern Matching Based Model for Implicit Opinion Question IdentificationabstractThis paper presents the results of developing subjectivity classifiers for Implicit Opinion Question (IOQ) identification. IOQs are defined as opinion questions with no opinion words. An IOQ example is "will the U.S. government pay more attention to the Pacific Rim?" Our analysis on community questions of Yahoo! Answers shows that a large proportion of opinion questions are IOQs. It is thus important to develop techniques to identify such questions. In this research, we first propose an effective framework based on mutual information and sequential pattern mining to construct an opinion lexicon that not only contains opinion words but also patterns. The discovered words and patterns are then combined with a machine learning technique to identify opinion questions. The experimental results on two datasets demonstrate the effectiveness of our approach. Hadi Amiri, Zhengjun Zha, Tat-Seng Chua |
AAAI | 1 |
| 2013 | Emerging topic detection for organizations from microblogsabstractMicroblog services have emerged as an essential way to strengthen the communications among individuals and organizations. These services promote timely and active discussions and comments towards products, markets as well as public events, and have attracted a lot of attentions from organizations. In particular, emerging topics are of immediate concerns to organizations since they signal current concerns of, and feedback by their users. Two challenges must be tackled for effective emerging topic detection. One is the problem of real-time relevant data collection and the other is the ability to model the emerging characteristics of detected topics and identify them before they become hot topics. To tackle these challenges, we first design a novel scheme to crawl the relevant messages related to the designated organization by monitoring multi-aspects of microblog content, including users, the evolving keywords and their temporal sequence. We then develop an incremental clustering framework to detect new topics, and employ a range of content and temporal features to help in promptly detecting hot emerging topics. Extensive evaluations on a representative real-world dataset based on Twitter data demonstrate that our scheme is able to characterize emerging topics well and detect them before they become hot topics. Yan Chen 0019, Hadi Amiri, Zhoujun Li 0001, Tat-Seng Chua |
SIGIR | 2 |
| 2012 | Sense Sentiment Similarity: An AnalysisabstractThis paper describes an emotion-based approach to acquire sentiment similarity of word pairs with respect to their senses. Sentiment similarity indicates the similarity between two words from their underlying sentiments. Our approach is built on a model which maps from senses of words to vectors of twelve basic emotions. The emotional vectors are used to measure the sentiment similarity of word pairs. We show the utility of measuring sentiment similarity in two main natural language processing tasks, namely, indirect yes/no question answer pairs (IQAP) Inference and sentiment orientation (SO) prediction. Extensive experiments demonstrate that our approach can effectively capture the sentiment similarity of word pairs and utilize this information to address the above mentioned tasks. Mitra Mohtarami, Hadi Amiri, Man Lan, Thanh Phu Tran, Chew Lim Tan |
AAAI | 2 |
| 2012 | Mining sentiment terminology through timeabstractThe correspondence between sentiment terminology and the active language used for expressing opinions is a crucial prerequisite for effective sentiment analysis. Mining sentiment terminology includes the detection of new opinion words as well as inferring their polarities. In this paper, we first propose a novel approach based on the interchangeability characteristic of words to detect new opinion words through time. We then show that the current non-time-based polarity inference approaches may assign opposite polarity to the same opinion word at different times. To tackle this issue, we consider the polarity scores computed at different times as polarity evidences (with the possibility of flawed evidences) and combine them to compute a globally correct polarity score for each opinion word. The experiments show that our approach is effective both in terms of the quality of the discovered new opinion words as well as its ability in inferring their polarities through time. Furthermore, we show the application of mining sentiment terminology through time in the sentiment classification (SC) task. The experiments show that mining more recent new opinion words leads to greater improvement in the performance of SC. To the best of our knowledge, this is the first work that investigates "time" as an important factor in mining sentiment terminology. Hadi Amiri, Tat-Seng Chua |
CIKM | 1 |
| 2012 | Mining slang and urban opinion words and phrases from cQA services: an optimization approachabstractCurrent opinion lexicons contain most of the common opinion words, but they miss slang and so-called urban opinion words and phrases (e.g. delish, cozy, yummy, nerdy, and yuck). These subjectivity clues are frequently used in community questions and are useful for opinion question analysis. This paper introduces a principled approach to constructing an opinion lexicon for community-based question answering (cQA) services. We formulate the opinion lexicon induction as a semi-supervised learning task in the graph context. Our method makes use of existing opinion words to extract new opinion entities (slang and urban words/phrases) from community questions. It then models the opinion entities in a graph context to learn the polarity of the new opinion entities based on the graph connectivity information. In contrast to previous approaches, our method not only learns such polarities from the labeled data but also from the unlabeled data and is more feasible in the web context where the dictionary-based relations (such as synonym, antonym, or hyponym) between most words are not available for constructing a high quality graph. The experiments show that our approach is effective both in terms of the quality of the discovered new opinion entities as well as its ability in inferring their polarity. Furthermore, since the value of opinion lexicons lies in their usefulness in applications, we show the utility of the constructed lexicon in the sentiment classification task. Hadi Amiri, Tat-Seng Chua |
WSDM | 1 |
| 2011 | Predicting the uncertainty of sentiment adjectives in indirect answersabstractOpinion question answering (QA) requires automatic and correct interpretation of an answer relative to its question. However, the ambiguity that often exists in the question-answer pairs causes complexity in interpreting the answers. This paper aims to infer yes/no answers from indirect yes/no question-answer pairs (IQAPs) that are ambiguous due to the presence of ambiguous sentiment adjectives. We propose a method to measure the uncertainty of the answer in an IQAP relative to its question. In particular, to infer the yes or no response from an IQAP, our method employs antonyms, synonyms, word sense disambiguation as well as the semantic association between the sentiment adjectives that appear in the IQAP. Extensive experiments demonstrate the effectiveness of our method over the baseline. Mitra Mohtarami, Hadi Amiri, Man Lan, Chew Lim Tan |
CIKM | 2 |
| 2009 | Hamshahri: A standard Persian text collection
Abolfazl AleAhmad, Hadi Amiri, Ehsan Darrudi, Masoud Rahgozar, Farhad Oroumchian |
Knowl. Based Syst. | 2 |
| 2006 | Underwater Noise Modeling and Direction-Finding Based on Conditional Heteroscedastic Time SeriesabstractIn this paper, we propose a new method for practical non-Gaussian and non-stationary underwater ambient noise modeling and direction-finding approach. In this application, measurement of ambient noise in natural environment shows that noise can sometimes be significantly non-Gaussian and time-varying features such as variances. Therefore, signal processing algorithms such as direction-finding that are optimized for Gaussian noise, may degrade significantly in this environment. Generalized autoregressive conditional heteroscedasticity (GARCH) models are feasible for heavy tailed PDFs and time-varying variances of stochastic process and also has flexible forms. We use a more realistic GARCH (1,1) based noise model in the maximum likelihood approach for the estimation of direction-of-arrivals (DOAs) of impinging sources and show using experimental data that this model is suitable for the additive noise in an underwater environment Hadi Amiri, Hamidreza Amindavar, Mahmoud Kamarei |
ICASSP (4) | 1 |
| 2004 | Array signal processing using GARCH noise modelingabstractWe propose a new method for modeling practical non-Gaussian and non-stationary noise in array signal processing. GARCH (generalized autoregressive conditional heteroscedasticity) models are introduced as the feasible model for the heavy tailed probability density functions (PDFs) and time varying variances of stochastic processes. We use the GARCH noise model in the maximum likelihood approach for the estimation of directions-of-arrival (DOAs). Our analysis exploits time varying variance and spatially non-uniform noise in sensor array signal processing. We show through simulations that this GARCH modeling is suitable for high-resolution source separation and noise suppression in a non-Gaussian environment. Hadi Amiri, Hamidreza Amindavar, R. Lynn Kirlin |
ICASSP (2) | 1 |