Mahdieh Soleymani Baghshah

dblp:21/473 · also Mahdieh Soleymani · DBLP profile ↗
← Back
50ranked-venue papers
10as first author
24since 2021 · last 2025
0000-0002-1971-6231ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CER: Confidence Enhanced Reasoning in LLMs
abstract
Ensuring the reliability of Large Language Models (LLMs) in complex reasoning tasks remains a formidable challenge, particularly in scenarios that demand precise mathematical calculations and knowledge-intensive open-domain generation. In this work, we introduce an uncertainty-aware framework designed to enhance the accuracy of LLM responses by systematically incorporating model confidence at critical decision points. We propose an approach that encourages multi-step reasoning in LLMs and quantify the confidence of intermediate answers such as numerical results in mathematical reasoning and proper nouns in open-domain generation. Then, the overall confidence of each reasoning chain is evaluated based on confidence of these critical intermediate steps. Finally, we aggregate the answer of generated response paths in a way that reflects the reliability of each generated content (as opposed to self-consistency in which each generated chain contributes equally to majority voting). We conducted extensive experiments in five datasets, three mathematical datasets and two open-domain datasets, using four LLMs. The results consistently validate the effectiveness of our novel confidence-aggregation method, leading to an accuracy improvement of up to 7.4% and 5.8% over baseline approaches in math and open-domain generation tasks, respectively. Code is publicly available at https://github.com/sharif-ml-lab/CER.
Ali Razghandi, Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani Baghshah
ACL (1)3
2025 CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
abstract
Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios. This study offers a comprehensive analysis of CLIP’s limitations in these contexts using a specialized dataset, ComCO, designed to evaluate CLIP’s encoders in diverse multi-object scenarios. Our findings reveal significant biases: the text encoder prioritizes first-mentioned objects, and the image encoder favors larger objects. Through retrieval and classification tasks, we quantify these biases across multiple CLIP variants and trace their origins to CLIP’s training process, supported by analyses of the LAION dataset and training progression. Our image-text matching experiments show substantial performance drops when object size or token order changes, underscoring CLIP’s instability with rephrased but semantically similar captions. Extending this to longer captions and text-to-image models like Stable Diffusion, we demonstrate how prompt order influences object prominence in generated images. For more details and access to our dataset and analysis code, visit our project repository: https://clip-oscope.github.io/.
Reza Abbasi, Aminreza Sefid, Mohammadali Banayeeanzade, Mohammad H. Rohban, Mahdieh Soleymani Baghshah
CVPR6
2025 LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
abstract
Why do gradient-based explanations struggle with Transformers, and how can we improve them? We identify gradientflow imbalances in Transformers that violate FullGradcompleteness, a critical property for attribution faithfulness that CNNs naturally possess. To address this issue, we introduce LibraGrad—a theoretically grounded post-hoc approach that corrects gradient imbalances through pruning and scaling of backward paths, without changing the forward pass or adding computational overhead. We evaluate LibraGrad using three metric families: Faithfulness, which quantifies prediction changes under perturbations of the most and least relevant features; Completeness Error, which measures attribution conservation relative to model outputs; and Segmentation AP, which assesses alignment with human perception. Extensive experiments across 8 architectures, 4 model sizes, and 5 datasets show that LibraGrad universally enhances gradient-based methods, outperforming existing white-box methods—including Transformer-specific approaches—across all metrics. We demonstrate superior qualitative results through two complementary evaluations: precise text-prompted region highlighting on CLIP models and accurate class discrimination between co-occurring animals on ImageNet-finetuned models—two settings on which existing methods often struggle. Libra-Grad is effective even on the attention-free MLP-Mixer architecture, indicating potential for extension to other modern architectures. Our code is freely available at https://nightmachinery.github.io/LibraGrad/.
Faridoun Mehri, Mahdieh Soleymani Baghshah, Mohammad Taher Pilehvar
CVPR2
2025 Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs
abstract
Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, visual search, scene description, and spatial relationship understanding. A key factor is that current LVLMs process visual features largely in parallel, lacking mechanisms for spatially grounded, serial attention. This paper introduces Visual Input Structure for Enhanced Reasoning (VISER), a simple, effective method that augments visual inputs with low-level spatial structures and pairs them with a textual prompt that encourages sequential, spatially-aware parsing. We empirically demonstrate substantial performance improvements across core visual reasoning tasks, using only a single-query inference. Specifically, VISER improves GPT-4o performance on visual search, counting, and spatial relationship tasks by 25.0%, 26.8%, and 9.5%, respectively, and reduces edit distance error in scene description by 0.32 on 2D datasets. Furthermore, we find that the visual modification is essential for these gains; purely textual strategies, including Chain-of-Thought prompting, are insufficient and can even degrade performance. VISER underscores the importance of visual input design over purely linguistically based reasoning strategies and suggests that visual structuring is a powerful and general approach for enhancing compositional and spatial reasoning in LVLMs.
Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, Ali Rahimiakbar, Mohammad Mahdi Vahedi, Hosein Hasani, Mahdieh Soleymani Baghshah
NeurIPS7
2025 Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications, where they frequently face data distributions unseen during training. Despite progress, existing methods are often vulnerable to spurious correlations that mislead models and compromise robustness. To address this, we propose SPROD, a novel prototype-based OOD detection approach that explicitly addresses the challenge posed by unknown spurious correlations. Our post-hoc method refines class prototypes to mitigate bias from spurious features without additional data or hyperparameter tuning, and is broadly applicable across diverse backbones and OOD detection settings. We conduct a comprehensive spurious correlation OOD detection benchmarking, comparing our method against existing approaches and demonstrating its superior performance across challenging OOD datasets, such as CelebA, Waterbirds, UrbanCars, Spurious Imagenet, and the newly introduced Animals MetaCoCo. On average, SPROD improves AUROC by 4.8% and FPR@95 by 9.4% over the second best.
Reihaneh Zohrabi, Hosein Hasani, Mahdieh Soleymani Baghshah, Anna Rohrbach, Marcus Rohrbach, Mohammad H. Rohban
NeurIPS3
2025 Predicting risk/reward ratio in financial markets for asset management using machine learning
Reza Yarbakhsh, Mahdieh Soleymani Baghshah, Hamidreza Karimaghaie
Expert Syst. Appl.2
2025 Self-supervised 3D medical image segmentation by flow-guided mask propagation learning
Adeleh Bitarafan, Mohammad Mozafari, Mohammad Farid Azampour, Mahdieh Soleymani Baghshah, Nassir Navab, Azade Farshad
Medical Image Anal.4
2025 Inductive biases for zero-shot systematic generalization in language-informed reinforcement learning
Negin Hashemi Dijujin, Seyed Roozbeh Razavi Rohani, Mohammad Mahdi Samiei Paqaleh, Mahdieh Soleymani Baghshah
Mach. Learn.4
2024 Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
abstract
While standard Empirical Risk Minimization (ERM) training is proven effective for image classification on in-distribution data, it fails to perform well on out-of-distribution samples. One of the main sources of distribution shift for image classification is the compositional nature of images. Specifically, in addition to the main object or component(s) determining the label, some other image components usually exist, which may lead to the shift of input distribution between train and test environments. More importantly, these components may have spurious corre-lations with the label. To address this issue, we propose Decompose-and-Compose (DaC), which improves robustness to correlation shift by a compositional approach based on combining elements of images. Based on our observations, models trained with ERM usually highly attend to either the causal components or the components having a high spurious correlation with the label (especially in datapoints on which models have a high confidence). In fact, according to the amount of spurious correlation and the easiness of classification based on the causal or non-causal components, the model usually attends to one of these more (on samples with high confidence). Following this, we first try to identify the causal components of images using class activation maps of models trained with ERM. Afterwards, we intervene on images by combining them and retraining the model on the augmented data, including the counterfac-tual ones. This work proposes a group-balancing method by intervening on images without requiring group labels or information regarding the spurious features during training. The method has an overall better worst group accuracy compared to previous methods with the same amount of supervision on the group labels in correlation shift. Our code is available at https://github.com/fhn98/DaC.
Fahimeh Hosseini Noohdani, Parsa Hosseini, Aryan Yazdan Parast, Hamidreza Yaghoubi Araghi, Mahdieh Soleymani Baghshah
CVPR5
2024 GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
abstract
Vision-language models (VLMs) are intensively used in many downstream tasks, including those requiring assessments of individuals appearing in the images. While VLMs perform well in simple single-person scenarios, in real-world applications, we often face complex situations in which there are persons of different genders doing different activities. We show that in such cases, VLMs are biased towards identifying the individual with the expected gender (according to ingrained gender stereotypes in the model or other forms of sample selection bias) as the performer of the activity. We refer to this bias in associating an activity with the gender of its actual performer in an image or text as the Gender-Activity Binding (GAB) bias and analyze how this bias is internalized in VLMs. To assess this bias, we have introduced the GAB dataset with approximately 5500 AI-generated images that represent a variety of activities, addressing the scarcity of real-world images for some scenarios. To have extensive quality control, the generated images are evaluated for their diversity, quality, and realism. We have tested 12 renowned pre-trained VLMs on this dataset in the context of text-to-image and image-to-text retrieval to measure the effect of this bias on their predictions. Additionally, we have carried out supplementary experiments to quantify the bias in VLMs’ text encoders and to evaluate VLMs’ capability to recognize activities. Our experiments indicate that VLMs experience an average performance decline of about 13.2% when confronted with gender-activity binding bias.
Ali Abdollahi, Mahdi Ghaznavi, Mohammad Reza Karimi Nejad, Arash Mari Oriyad, Reza Abbasi, Ali Salesi, Melika Behjati, Mohammad H. Rohban, Mahdieh Soleymani Baghshah
ECAI9
2024 Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
Reza Abbasi, Mohammad H. Rohban, Mahdieh Soleymani Baghshah
ECCV (89)3
2024 SOInter: A Novel Deep Energy-Based Interpretation Method for Explaining Structured Output Models
abstract
This paper proposes a novel interpretation technique to explain the behavior of structured output models, which simultaneously learn mappings between an input vector and a set of output variables. As a result of the complex relationships between the computational path of output variables in structured models, a feature may impact an output value via other output variables. We focus on one of the outputs as the target and try to find the most important features adopted by the structured model to decide on the target in each locality of the input space. We consider an arbitrary structured output model available as a black-box and argue that considering correlations among output variables can improve explanation quality. The goal is to train a function as an interpreter for the target output variable over the input space. We introduce an energy-based training process for the interpreter function, which effectively considers the structural information incorporated into the model to be explained. The proposed method's effectiveness is confirmed using various simulated and real data sets.
Seyyede Fatemeh Seyyedsalehi, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001
ICLR2
2024 RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
abstract
In recent years, there have been significant improvements in various forms of image outlier detection. However, outlier detection performance under adversarial settings lags far behind that in standard settings. This is due to the lack of effective exposure to adversarial scenarios during training, especially on unseen outliers, leading detection models failing to learn robust features. To bridge this gap, we introduce RODEO, a data-centric approach that generates effective outliers for robust outlier detection. More specifically, we show that incorporating outlier exposure (OE) and adversarial training could be an effective strategy for this purpose, as long as the exposed training outliers meet certain characteristics, including diversity, and both conceptual differentiability and analogy to the inlier samples. We leverage a text-to-image model to achieve this goal. We demonstrate both quantitatively and qualitatively that our adaptive OE method effectively generates ”diverse” and ”near-distribution” outliers, leveraging information from both text and image domains. Moreover, our experimental results show that utilizing our synthesized outliers significantly enhances the performance of the outlier detector, particularly in adversarial settings.
Hossein Mirzaei, Hamid Reza Dehbashi, Ali Ansari 0001, Sepehr Ghobadi, Masoud Hadi, Arshia Soltani Moakhar, Mohammad Azizmalayeri, Mahdieh Soleymani Baghshah, Mohammad H. Rohban
ICML9
2023 VISA-FSS: A Volume-Informed Self Supervised Approach for Few-Shot 3D Segmentation
Mohammad Mozafari, Adeleh Bitarafan, Mohammad Farid Azampour, Azade Farshad, Mahdieh Soleymani Baghshah, Nassir Navab
MICCAI (2)5
2022 DMNP: A Deep Learning Approach for Missing Node Prediction in Partially Observed Graphs
abstract
Missing data is unavoidable in graphs, which can significantly affect the accuracy of downstream tasks. Many methods have been proposed to mitigate missing data in partially observed graphs. Most of these approaches assume they have complete access to graph nodes and only focus on recovering missing links, while in practice a part of the graph nodes can also be out of access. This work presents Deep Missing Node Predictor (DMNP), a novel deep learning-based approach to recovering missing nodes in partly observed graphs. Our proposed approach does not rely on additional information that in many cases does not exist. We compare our model with graph completion and deep graph generation baselines. The experimental results show that the DMNP model outperforms previous state-of-the-art approaches.
Faezeh Faez, Ali Akhoondian Amiri, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001
ASONAM3
2022 BIMRL: Brain Inspired Meta Reinforcement Learning
abstract
Sample efficiency has been a key issue in reinforcement learning (RL). An efficient agent must be able to leverage its prior experiences to quickly adapt to similar, but new tasks and situations. Meta-RL is one attempt at formalizing and ad-dressing this issue. Inspired by recent progress in meta-RL, we introduce BIMRL, a novel multi-layer architecture along with a novel brain-inspired memory module that will help agents quickly adapt to new tasks within a few episodes. We also utilize this memory module to design a novel intrinsic reward that will guide the agent's exploration. Our architecture is inspired by findings in cognitive neuroscience and is compatible with the knowledge on connectivity and functionality of different regions in the brain. We empirically validate the effectiveness of our proposed method by competing with or surpassing the performance of some strong baselines on multiple MiniGrid environments.
Seyed Roozbeh Razavi Rohani, Saeed Hedayatian, Mahdieh Soleymani Baghshah
IROS3
2022 Vol2Flow: Segment 3D Volumes Using a Sequence of Registration Flows
Adeleh Bitarafan, Mohammad Farid Azampour, Kian Bakhtari, Mahdieh Soleymani Baghshah, Matthias Keicher, Nassir Navab
MICCAI (4)4
2022 RA-GCN: Graph convolutional network for disease prediction problems with imbalanced data
Mahsa Ghorbani, Anees Kazi, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001, Nassir Navab
Medical Image Anal.3
2021 Rate-Distortion Analysis of Minimum Excess Risk in Bayesian Learning
abstract
In parametric Bayesian learning, a prior is assumed on the parameter $W$ which determines the distribution of samples. In this setting, Minimum Excess Risk (MER) is defined as the difference between the minimum expected loss achievable when learning from data and the minimum expected loss that could be achieved if $W$ was observed. In this paper, we build upon and extend the recent results of (Xu & Raginsky, 2020) to analyze the MER in Bayesian learning and derive information-theoretic bounds on it. We formulate the problem as a (constrained) rate-distortion optimization and show how the solution can be bounded above and below by two other rate-distortion functions that are easier to study. The lower bound represents the minimum possible excess risk achievable by \emph{any} process using $R$ bits of information from the parameter $W$. For the upper bound, the optimization is further constrained to use $R$ bits from the training set, a setting which relates MER to information-theoretic bounds on the generalization gap in frequentist learning. We derive information-theoretic bounds on the difference between these upper and lower bounds and show that they can provide order-wise tight rates for MER under certain conditions. This analysis gives more insight into the information-theoretic nature of Bayesian learning as well as providing novel bounds.
Hassan Hafez-Kolahi, Behrad Moniri, Shohreh Kasaei, Mahdieh Soleymani Baghshah
ICML4
2021 GKD: Semi-supervised Graph Knowledge Distillation for Graph-Independent Inference
Mahsa Ghorbani, Mojtaba Bahrami, Anees Kazi, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001, Nassir Navab
MICCAI (5)4
2021 Generative vs. Discriminative: Rethinking The Meta-Continual Learning
abstract
Deep neural networks have achieved human-level capabilities in various learning tasks. However, they generally lose performance in more realistic scenarios like learning in a continual manner. In contrast, humans can incorporate their prior knowledge to learn new concepts efficiently without forgetting older ones. In this work, we leverage meta-learning to encourage the model to learn how to learn continually. Inspired by human concept learning, we develop a generative classifier that efficiently uses data-driven experience to learn new concepts even from few samples while being immune to forgetting. Along with cognitive and theoretical insights, extensive experiments on standard benchmarks demonstrate the effectiveness of the proposed method. The ability to remember all previous concepts, with negligible computational and structural overheads, suggests that generative models provide a natural way for alleviating catastrophic forgetting, which is a major drawback of discriminative models.
Mohammadamin Banayeeanzade, Rasoul Mirzaiezadeh, Hosein Hasani, Mahdieh Soleymani Baghshah
NeurIPS4
2021 DGSAN: Discrete generative self-adversarial network
Ehsan Montahaei, Danial Alihosseini, Mahdieh Soleymani Baghshah
Neurocomputing3
2021 3D Image Segmentation With Sparse Annotation by Self-Training and Internal Registration
abstract
Anatomical image segmentation is one of the foundations for medical planning. Recently, convolutional neural networks (CNN) have achieved much success in segmenting volumetric (3D) images when a large number of fully annotated 3D samples are available. However, rarely a volumetric medical image dataset containing a sufficient number of segmented 3D images is accessible since providing manual segmentation masks is monotonous and time-consuming. Thus, to alleviate the burden of manual annotation, we attempt to effectively train a 3D CNN using a sparse annotation where ground truth on just one 2D slice of the axial axis of each training 3D image is available. To tackle this problem, we propose a self-training framework that alternates between two steps consisting of assigning pseudo annotations to unlabeled voxels and updating the 3D segmentation network by employing both the labeled and pseudo labeled voxels. To produce pseudo labels more accurately, we benefit from both propagation of labels (or pseudo-labels) between adjacent slices and 3D processing of voxels. More precisely, a 2D registration-based method is proposed to gradually propagate labels between consecutive 2D slices and a 3D U-Net is employed to utilize volumetric information. Ablation studies on benchmarks show that cooperation between the 2D registration and the 3D segmentation provides accurate pseudo-labels that enable the segmentation network to be trained effectively when for each training sample only even one segmented slice by an expert is available. Our method is assessed on the CHAOS and Visceral datasets to segment abdominal organs. Results demonstrate that despite utilizing just one segmented slice for each 3D image (that is weaker supervision in comparison with the compared weakly supervised methods) can result in higher performance and also achieve closer results to the fully supervised manner.
Adeleh Bitarafan, Mahdi Nikdan, Mahdieh Soleymani Baghshah
IEEE J. Biomed. Health Informatics3
2021 MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification Challenge
abstract
Detecting various types of cells in and around the tumor matrix holds a special significance in characterizing the tumor micro-environment for cancer prognostication and research. Automating the tasks of detecting, segmenting, and classifying nuclei can free up the pathologists' time for higher value tasks and reduce errors due to fatigue and subjectivity. To encourage the computer vision research community to develop and test algorithms for these tasks, we prepared a large and diverse dataset of nucleus boundary annotations and class labels. The dataset has over 46,000 nuclei from 37 hospitals, 71 patients, four organs, and four nucleus types. We also organized a challenge around this dataset as a satellite event at the International Symposium on Biomedical Imaging (ISBI) in April 2020. The challenge saw a wide participation from across the world, and the top methods were able to match inter-human concordance for the challenge metric. In this paper, we summarize the dataset and the key findings of the challenge, including the commonalities and differences between the methods developed by various participants. We have released the MoNuSAC2020 dataset to the public.
Ruchika Verma, Neeraj Kumar 0002, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E Ahmed Raza, Nasir M. Rajpoot, Xiyi Wu, Huai Chen, Lisheng Wang, Hyun Jung, G. Thomas Brown, Shuolin Liu, Seyed Alireza Fatemi Jahromi, Aliasghar Khani, Ehsan Montahaei, Mahdieh Soleymani Baghshah, Hamid Behroozi, Pavel Semkin, Alexandr Rassadin, Prasad Dutande, Romil Lodaya, Ujjwal Baid, Bhakti Baheti, Sanjay N. Talbar, Amirreza Mahbod, Rupert Ecker, Isabella Ellinger, Bin Dong 0006, Zhengyu Xu, Yuehan Yao, Ming Feng, Kele Xu, Hasib Zunair, A. Ben Hamza, Steven M. Smiley, Tang-Kai Yin, Qi-Rui Fang, Shikhar Srivastava 0001, Dwarikanath Mahapatra, Lubomira Trnavska, Hanyun Zhang, Priya Lakshmi Narayanan, Justin Law, Yinyin Yuan, Abhiroop Tejomay, Aditya Mitkari, Dinesh Koka, Vikas Ramachandra, Lata Kini, Amit Sethi
IEEE Trans. Medical Imaging22
2020 Paraphrase Generation by Learning How to Edit from Samples
abstract
Neural sequence to sequence text generation has been proved to be a viable approach to paraphrase generation. Despite promising results, paraphrases generated by these models mostly suffer from lack of quality and diversity. To address these problems, we propose a novel retrieval-based method for paraphrase generation. Our model first retrieves a paraphrase pair similar to the input sentence from a pre-defined index. With its novel editor module, the model then paraphrases the input sequence by editing it using the extracted relations between the retrieved pair of sentences. In order to have fine-grained control over the editing process, our model uses the newly introduced concept of Micro Edit Vectors. It both extracts and exploits these vectors using the attention mechanism in the Transformer architecture. Experimental results show the superiority of our paraphrase generation method in terms of both automatic metrics, and human evaluation of relevance, grammaticality, and diversity of generated paraphrases.
Amirhossein Kazemnejad, Mohammadreza Salehi, Mahdieh Soleymani Baghshah
ACL3
2020 Conditioning and Processing: Techniques to Improve Information-Theoretic Generalization Bounds
abstract
Obtaining generalization bounds for learning algorithms is one of the main subjects studied in theoretical machine learning. In recent years, information-theoretic bounds on generalization have gained the attention of researchers. This approach provides an insight into learning algorithms by considering the mutual information between the model and the training set. In this paper, a probabilistic graphical representation of this approach is adopted and two general techniques to improve the bounds are introduced, namely conditioning and processing. In conditioning, a random variable in the graph is considered as given, while in processing a random variable is substituted with one of its children. These techniques can be used to improve the bounds by either sharpening them or increasing their applicability. It is demonstrated that the proposed framework provides a simple and unified way to explain a variety of recent tightening results. New improved bounds derived by utilizing these techniques are also proposed.
Hassan Hafez-Kolahi, Zeinab Golgooni, Shohreh Kasaei, Mahdieh Soleymani Baghshah
NeurIPS4
2019 MGCN: semi-supervised classification in multi-layer graphs with graph convolutional networks
abstract
Graph embedding is an important approach for graph analysis tasks such as node classification and link prediction. The goal of graph embedding is to find a low dimensional representation of graph nodes that preserves the graph information. Recent methods like Graph Convolutional Network (GCN) try to consider node attributes (if available) besides node relations and learn node embeddings for unsupervised and semi-supervised tasks on graphs. On the other hand, multi-layer graph analysis has been received attention recently. However, the existing methods for multi-layer graph embedding cannot incorporate all available information (like node attributes). Moreover, most of them consider either type of nodes or type of edges, and they do not treat within and between layer edges differently. In this paper, we propose a method called MGCN that utilizes the GCN for multi-layer graphs. MGCN embeds nodes of multi-layer graphs using both within and between layers relations and nodes attributes. We evaluate our method on the semi-supervised node classification task. Experimental results demonstrate the superiority of the proposed method to other multi-layer and single-layer competitors and also show the positive effect of using cross-layer edges.
Mahsa Ghorbani, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001
ASONAM2
2019 Universal Adversarial Attacks on Text Classifiers
abstract
Despite the vast success neural networks have achieved in different application domains, they have been proven to be vulnerable to adversarial perturbations (small changes in the input), which lead them to produce the wrong output. In this paper, we propose a novel method, based on gradient projection, for generating universal adversarial perturbations for text; namely sequence of words that can be added to any input in order to fool the classifier with high probability. We observed that text classifiers are quite vulnerable to such perturbations: inserting even a single adversarial word to the beginning of every input sequence can drop the accuracy from 93% to 50%.
Melika Behjati, Seyed-Mohsen Moosavi-Dezfooli, Mahdieh Soleymani Baghshah, Pascal Frossard
ICASSP3
2019 Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks
abstract
Numerous neurophysiological studies have revealed that a large number of the primary visual cortex neurons operate in a regime called surround modulation. Surround modulation has a substantial effect on various perceptual tasks, and it also plays a crucial role in the efficient neural coding of the visual cortex. Inspired by the notion of surround modulation, we designed new excitatory-inhibitory connections between a unit and its surrounding units in the convolutional neural network (CNN) to achieve a more biologically plausible network. Our experiments show that this simple mechanism can considerably improve both the performance and training speed of traditional CNNs in visual tasks. We further explore additional outcomes of the proposed structure. We first evaluate the model under several visual challenges, such as the presence of clutter or change in lighting conditions and show its superior generalization capability in handling these challenging situations. We then study possible changes in the statistics of neural activities such as sparsity and decorrelation and provide further insight into the underlying efficiencies of surround modulation. Experimental results show that importing surround modulation into the convolutional layers ensues various effects analogous to those derived by surround modulation in the visual cortex.
Hosein Hasani, Mahdieh Soleymani Baghshah, Hamid K. Aghajan
NeurIPS2
2019 An Efficient Semi-Supervised Multi-label Classifier Capable of Handling Missing Labels
abstract
Multi-label classification has received considerable interest in recent years. Multi-label classifiers usually need to address many issues including: handling large-scale datasets with many instances and a large set of labels, compensating missing label assignments in the training set, considering correlations between labels, as well as exploiting unlabeled data to improve prediction performance. To tackle datasets with a large set of labels, embedding-based methods represent the label assignments in a low-dimensional space. Many state-of-the-art embedding-based methods use a linear dimensionality reduction to map the label assignments to a low-dimensional space. However, by doing so, these methods actually neglect the tail labels - labels that are infrequently assigned to instances. In this paper, we propose an embedding-based method that non-linearly embeds the label vectors using a stochastic approach, thereby predicting the tail labels more accurately. Moreover, the proposed method has excellent mechanisms for handling missing labels, dealing with large-scale datasets, as well as exploiting unlabeled data. Experiments on real-world datasets show that our method outperforms state-of-the-art multi-label classifiers by a large margin, in terms of prediction performance, as well as training time. Our implementation of the proposed method is available online at:https://github.com/Akbarnejad/ESMC_ Implementation.
Amirhossein Akbarnejad, Mahdieh Soleymani Baghshah
IEEE Trans. Knowl. Data Eng.2
2017 NMF-Based Label Space Factorization for Multi-label Classification
abstract
Multi-label classification is a learning task in which each data sample can belong to more than one class. Until now, some methods that are based on reducing the dimensionality of the label space have been proposed. However, these methods have not used specific properties of the label space for this purpose. In this paper, we intend to find a hidden space in which both the input feature vectors and the label vectors are embedded. We propose a modified Non-Negative Matrix Factorization (NMF) method that is suitable for decomposing the label matrix and finding a proper hidden space by a feature-aware approach. We consider that the label matrix is binary and also in this matrix some deserving labels for an instance may not be on (called missing labels). We conduct several experiments and show the superiority of our proposed methods to the state-of-the-art multi- label classification methods.
Mohammad Firouzi, Mahmood Karimian, Mahdieh Soleymani Baghshah
ICMLA3
2017 Joint predictive model and representation learning for visual domain adaptation
Marzieh Gheisari, Mahdieh Soleymani Baghshah
Eng. Appl. Artif. Intell.2
2017 Multi-modal deep distance metric learning
abstract
In many real-world applications, data contain heterogeneous input modalities (e.g., web pages include images, text, etc.). Moreover, data such as images are usually described using different views (i.e. different sets of features). Learning a distance metric or similarity measure that originates fr om all input modalities or views is essential for many tasks such as content-based retrieval ones. In these cases, similar and dissimilar pairs of data can be used to find a better representation of data in which similarity and dissimilarity constraints are better satisfied. In this paper, we incorporate supervision in the form of pairwise similarity and/or dissimilarity constraints into multi-modal deep networks to combine different modalities into a shared latent space. Using properties of multi-modal data, we design multi-modal deep networks and propose a pre-training algorithm for these networks. In fact, the proposed network has the ability of learning intra- and inter-modal high-order statistics from raw features and we control its high flexibility via an efficient multi-stage pre-training phase corresponding to properties of multi-modal data. Experimental results show that the proposed method outperforms recent methods on image retrieval tasks.
Seyed Mahdi Roostaiyan, Ehsan Imani, Mahdieh Soleymani Baghshah
Intell. Data Anal.3
2017 A probabilistic multi-label classifier with missing and noisy labels handling capability
Amirhossein Akbarnejad, Mahdieh Soleymani Baghshah
Pattern Recognit. Lett.2
2016 MDL-CW: A Multimodal Deep Learning Framework with CrossWeights
abstract
Deep learning has received much attention as of the most powerful approaches for multimodal representation learning in recent years. An ideal model for multimodal data can reason about missing modalities using the available ones, and usually provides more information when multiple modalities are being considered. All the previous deep models contain separate modality-specific networks and find a shared representation on top of those networks. Therefore, they only consider high level interactions between modalities to find a joint representation for them. In this paper, we propose a multimodal deep learning framework (MDLCW) that exploits the cross weights between representation of modalities, and try to gradually learn interactions of the modalities in a deep network manner (from low to high level interactions). Moreover, we theoretically show that considering these interactions provide more intra-modality information, and introduce a multi-stage pre-training method that is based on the properties of multi-modal data. In the proposed framework, as opposed to the existing deep methods for multi-modal data, we try to reconstruct the representation of each modality at a given level, with representation of other modalities in the previous layer. Extensive experimental results show that the proposed model outperforms state-of-the-art information retrieval methods for both image and text queries on the PASCAL-sentence and SUN-Attribute databases.
Sarah Rastegar, Mahdieh Soleymani Baghshah, Hamid R. Rabiee 0001, Seyed Mohsen Shojaee
CVPR2
2016 Active Distance-Based Clustering Using K-Medoids
Amin Aghaee, Mehrdad Ghadiri, Mahdieh Soleymani Baghshah
PAKDD (1)3
2016 Incremental Evolving Domain Adaptation
abstract
Almost all of the existing domain adaptation methods assume that all test data belong to a single stationary target distribution. However, in many real world applications, data arrive sequentially and the data distribution is continuously evolving. In this paper, we tackle the problem of adaptation to a continuously evolving target domain that has been recently introduced. We assume that the available data for the source domain are labeled but the examples of the target domain can be unlabeled and arrive sequentially. Moreover, the distribution of the target domain can evolve continuously over time. We propose the Evolving Domain Adaptation (EDA) method that first finds a new feature space in which the source domain and the current target domain are approximately indistinguishable. Therefore, source and target domain data are similarly distributed in the new feature space and we use a semi-supervised classification method to utilize both the unlabeled data of the target domain and the labeled data of the source domain. Since test data arrives sequentially, we propose an incremental approach both for finding the new feature space and for semi-supervised classification. Experiments on several real datasets demonstrate the superiority of our proposed method in comparison to the other recent methods.
Adeleh Bitarafan, Mahdieh Soleymani Baghshah, Marzieh Gheisari
IEEE Trans. Knowl. Data Eng.2
2015 Near Linear-Time Community Detection in Networks with Hardly Detectable Community Structure
abstract
Identifying communities has always been a fundamental task in analysis of complex networks. Many methods have been devised over the last decade for detection of communities. Amongst them, the label propagation algorithm brings great scalability together with high accuracy. However, it has one major flaw; when the community structure in the network is not clear enough, it will assign every node the same label, thus detecting the whole graph as one giant community. We have addressed this issue by setting a capacity for communities, starting from a small value and gradually increasing it over time. Preliminary results show that not only our extension improves the detection capability of the classic label propagation algorithm when communities are not clearly detectable, but also improves the overall quality of the identified clusters in complex networks with a clear community structure.
Aria Rezaei, Saeed Mahlouji Far, Mahdieh Soleymani Baghshah
ASONAM3
2015 Unsupervised domain adaptation via representation learning and adaptive classifier learning
Marzieh Gheisari, Mahdieh Soleymani Baghshah
Neurocomputing2
2014 Scalable semi-supervised clustering by spectral kernel learning
Mahdieh Soleymani Baghshah, Fatemeh Afsari, Saeed Bagheri Shouraki, Esfandiar Eslami
Pattern Recognit. Lett.1
2013 PSSDL: Probabilistic Semi-supervised Dictionary Learning
Behnam Babagholami-Mohamadabadi, Ali Zarghami, Mohammadreza Zolfaghari, Mahdieh Soleymani Baghshah
ECML/PKDD (3)4
2011 Learning low-rank kernel matrices for constrained clustering
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
Neurocomputing1
2010 Efficient Kernel Learning from Constraints and Unlabeled Data
abstract
Recently, distance metric learning has been received an increasing attention and found as a powerful approach for semi-supervised learning tasks. In the last few years, several methods have been proposed for metric learning when must-link and/or cannot-link constraints as supervisory information are available. Although many of these methods learn global Mahalanobis metrics, some recently introduced methods have tried to learn more flexible distance metrics using a kernel-based approach. In this paper, we consider the problem of kernel learning from both pairwise constraints and unlabeled data. We propose a method that adapts a flexible distance metric via learning a nonparametric kernel matrix. We formulate our method as an optimization problem that can be solved efficiently. Experimental evaluations show the effectiveness of our method compared to some recently introduced methods on a variety of data sets.
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
ICPR1
2010 Kernel-based metric learning for semi-supervised clustering
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
Neurocomputing1
2010 Non-linear metric learning using pairwise similarity and dissimilarity constraints and the geometrical structure of data
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
Pattern Recognit.1
2009 Semi-Supervised Metric Learning Using Pairwise Constraints
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
IJCAI1
2009 Metric learning for semi-supervised clustering using pairwise constraints and the geometrical structure of data
abstract
Metric learning is a powerful approach for semi-supervised clustering. In this paper, a metric learning method considering both pairwise constraints and the geometrical structure of data is introduced for semi-supervised clustering. At first, a smoot
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
Intell. Data Anal.1
2008 A fuzzy clustering algorithm for finding arbitrary shaped clusters
abstract
Until now, many algorithms have been introduced for finding arbitrary shaped clusters, but none of these algorithms is able to identify all sorts of cluster shapes and structures that are encountered in practice. Furthermore, the time complexity of the existing algorithms is usually high and applying them on large datasets is time-consuming. In this paper, a novel fast clustering algorithm is proposed. This algorithm distinguishes clusters of different shapes using a two-stage clustering approach. In the first stage, the data points are grouped into a relatively large number of fuzzy ellipsoidal sub-clusters. Then, connections between sub-clusters are established according to the Bhattacharya distances and final clusters are formed from the resulted graph of sub-clusters in the second stage. Experimental results show the ability of the proposed algorithm for finding clusters of different shapes.
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki
AICCSA1
2008 An agent-based clustering algorithm using potential fields
abstract
In this paper, a novel clustering algorithm using an agent-based architecture along with a force-based clustering algorithm is proposed. To this end, a set of simple mobile agents that have limited processing power is used. These agents communicate in a pairwise manner to exchange their position information. As opposed to the bio-inspired clustering algorithms that need a set of local rules to specify the agent movements, in this paper the agent motions are driven from attractive and repulsive potential fields that are created by the data points and the other agents respectively. Each agent moves according to the resulted force from applying the potential fields and announces its next position to the other agents. The movement of agents is continued until all of them reach equilibrium points.
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki, Caro Lucas
AICCSA1
2007 Evolving fuzzy classifiers using a symbiotic approach
abstract
Fuzzy rule-based classifiers are one of the famous forms of the classification systems particularly in the data mining field. Genetic algorithm is a useful technique for discovering this kind of classifiers and it has been used for this purpose in some studies. In this paper, we propose a new symbiotic evolutionary approach to find desired fuzzy rule-based classifiers. For this purpose, a symbiotic combination operator has been designed as an alternative to the recombination operator (crossover) in the genetic algorithms. In the proposed approach, the evolution starts from simple chromosomes and the structure of chromosomes gets complex gradually during the evolutionary process. Experimental results on some standard data sets show the high performance of the proposed approach compared to the other existing approaches.
Mahdieh Soleymani Baghshah, Saeed Bagheri Shouraki, Ramin Halavati, Caro Lucas
IEEE Congress on Evolutionary Computation1