Faruk Ahmed

dblp:56/2471 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 29% Representation and self-supervised learning · 25% Trustworthy machine learning · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 50% Computational science and engineering · 50%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
digital pathology
0.912025
Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital Pathology · ICCV 2025
Computational science and engineering
multi-task learning
0.912025
Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital Pathology · ICCV 2025
Machine learning › Generative modeling
domain translation
0.512021
Integrating Categorical Semantics into Unsupervised Domain Translation · ICLR 2021
Machine learning › Representation and self-supervised learning
invariant representation
0.512021
Systematic generalisation with group invariant predictions · ICLR 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Systematic generalisation with group invariant predictions · ICLR 2021
Machine learning › Representation and self-supervised learning
systematic generalization
0.512021
Systematic generalisation with group invariant predictions · ICLR 2021
Computer vision › Image recognition and object detection
object recognition
0.412020
Detecting Semantic Anomalies · AAAI 2020
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.412020
Detecting Semantic Anomalies · AAAI 2020
Machine learning › Generative modeling
image generation
0.422017
PixelVAE: A Latent Variable Model for Natural Images · ICLR (Poster) 2017
Improved Training of Wasserstein GANs · NIPS 2017
Machine learning › Generative modeling
generative adversarial network
0.312017
Improved Training of Wasserstein GANs · NIPS 2017
Machine learning › Deep learning architectures and training › regularization › gradient regularization
gradient penalty
0.312017
Improved Training of Wasserstein GANs · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.312017
PixelVAE: A Latent Variable Model for Natural Images · ICLR (Poster) 2017
Machine learning › Generative modeling › image generation
natural image synthesis
0.312017
PixelVAE: A Latent Variable Model for Natural Images · ICLR (Poster) 2017
Machine learning › Generative modeling
variational autoencoder
0.312017
PixelVAE: A Latent Variable Model for Natural Images · ICLR (Poster) 2017
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312025
Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital Pathology · ICCV 2025
Computer vision › Segmentation and scene understanding
image segmentation
0.212015
Optimizing Expected Intersection-Over-Union with Candidate-Constrained CRFs · ICCV 2015
Computer vision › Segmentation and scene understanding › image segmentation
intersection-over-union optimization
0.212015
Optimizing Expected Intersection-Over-Union with Candidate-Constrained CRFs · ICCV 2015
Machine learning › Learning paradigms
multi-task learning
0.112020
Detecting Semantic Anomalies · AAAI 2020

Methods — techniques the papers use, named apart from their topics

multimodal learning · 1.7foundation model · 1.7group invariance · 0.5multi-task learning · 0.4weight clipping · 0.3lipschitz constraint · 0.3latent variable modeling · 0.3gradient penalty · 0.3conditional random field · 0.2candidate-constrained inference · 0.2
YearPublicationVenuePosition
2025 Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital Pathology
Vishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed, Maged Goubran, Anne L. Martel
ICCV4
2021 Systematic generalisation with group invariant predictions
Faruk Ahmed, Yoshua Bengio, Harm van Seijen, Aaron C. Courville
ICLR1
2021 Integrating Categorical Semantics into Unsupervised Domain Translation
Samuel Lavoie-Marchildon, Faruk Ahmed, Aaron C. Courville
ICLR2
2020 Detecting Semantic Anomalies
abstract
We critically appraise the recent interest in out-of-distribution (OOD) detection and question the practical relevance of existing benchmarks. While the currently prevalent trend is to consider different datasets as OOD, we argue that out-distributions of practical interest are ones where the distinction is semantic in nature for a specified context, and that evaluative tasks should reflect this more closely. Assuming a context of object recognition, we recommend a set of benchmarks, motivated by practical applications. We make progress on these benchmarks by exploring a multi-task learning based approach, showing that auxiliary objectives for improved semantic awareness result in improved semantic anomaly detection, with accompanying generalization benefits.
Faruk Ahmed, Aaron C. Courville
AAAI1
2020 Assistive System for Navigating Complex Realistic Simulated World Using Reinforcement Learning
abstract
Finding a free path without obstacles or situation that pose minimal risk is critical for safe navigation. People who are sighted and people who are blind or visually impaired require navigation safety while walking on a sidewalk. In this paper we develop assistive navigation on a sidewalk by integrating sensory inputs using reinforcement learning. We train the reinforcement model in a simulated robotic environment which is used to avoid sidewalk obstacles. A conversational agent is built by training with real conversation data. The reinforcement learning model along with a conversational agent improved the obstacle avoidance experience about 2.5% from the base case which is 78.75%.
Faruk Ahmed, Md. Sultan Mahmud, Mohammed Yeasin
IJCNN1
2020 Multivariate Models for Decoding Hearing Impairment using EEG Gamma-Band Power Spectral Density
abstract
Speech-in-noise (SIN) comprehension decreases with age, and these declines have been related to social isolation, depression, and dementia in the elderly. In this work, we build models to distinguish the normal hearing (NH) or mild hearing impairment (HI) using the different genres of machine learning. We compute band wise power spectral density (PSD) of source- derived EEGs as features in building models using support vector machine (SVM), k-nearest neighbors (KNN), and AdaBoost classifiers and compare their performance while listeners perceived clear or noise-degraded sounds. Combining all frequency bands features obtained from the whole-brain, the SVM registered the best performance. The group classification accuracy was found to be 94.90% [area under the curve (AUC) 94.75%; F1-score 95.00%] perceived the clear speech, and for noise- degraded speech perception, accuracy was found to be 92.52% (AUC 91.12%, and F1-score 93.00%). Remarkably, individual frequency band analysis on whole-brain data showed that γ frequency band segregated groups with a best accuracy of 96.78%, AUC 96.79% for clear speech data and noise-degraded speech data yielded slightly less accuracy of 93.62% with AUC 93.17% by using SVM. A separate analysis using the left hemisphere (LH) and right hemisphere (RH) data showed that the LH activity is a better predictor of groups compared to RH. These results are consistent with the dominance of LH in auditory-linguistic processing. Our results demonstrate that spectral features of the γ-band frequency could be used to differentiate NH and HI older adults in terms of their ability to process speech sounds. These findings would be useful to model attentional and listening assistive devices to amplify a more specific pitch than others.
Md. Sultan Mahmud, Faruk Ahmed, Mohammed Yeasin, Claude Alain, Gavin M. Bidelman
IJCNN2
2019 Probability Distillation: A Caveat and Alternatives
Chin-Wei Huang, Faruk Ahmed, Alexandre Lacoste, Aaron C. Courville
UAI2
2017 PixelVAE: A Latent Variable Model for Natural Images
Ishaan Gulrajani, Faruk Ahmed, Adrien Ali Taïga, Francesco Visin, David Vázquez 0001, Aaron C. Courville
ICLR (Poster)3
2017 Optimization and evaluation of deep architectures for ambient awareness on a sidewalk
abstract
We optimize and compare the performance of different deep learning architectures for awareness on a sidewalk using small form factor devices such as Raspberry Pi 3. Our main objective is to find deep learning architecture that is complex enough to accurately classify a set obstacles on the sidewalk. Out selection criteria are: minimum number of parameters, lower power consumption, and robustness against the effect of the diurnal cycle. In particular, we compare the performance of GoogleNet, ResNet, and VGG-16 on a database constructed for AS applications. Empirical evaluation on AS database suggests that the performance of ResNet is superior compared to other architectures e.g., 99.46% and 97.69% on RGB and La ∗ b∗ respectively. To further our objective we optimize the hyperparameters of ResNet to find architecture with a lower number of parameters without losing accuracy. Furthermore, we investigate the efficacy of different color spaces to address problems related to the diurnal cycle and power usage without sacrificing accuracy and generalizability.
Mohammed Yeasin, Faruk Ahmed
IJCNN2
2017 Improved Training of Wasserstein GANs
abstract
Generative Adversarial Networks (GANs) are powerful generative models, but suffer from training instability. The recently proposed Wasserstein GAN (WGAN) makes progress toward stable training of GANs, but sometimes can still generate only poor samples or fail to converge. We find that these problems are often due to the use of weight clipping in WGAN to enforce a Lipschitz constraint on the critic, which can lead to undesired behavior. We propose an alternative to clipping weights: penalize the norm of gradient of the critic with respect to its input. Our proposed method performs better than standard WGAN and enables stable training of a wide variety of GAN architectures with almost no hyperparameter tuning, including 101-layer ResNets and language models with continuous generators. We also achieve high quality generations on CIFAR-10 and LSUN bedrooms.
Ishaan Gulrajani, Faruk Ahmed, Martín Arjovsky, Vincent Dumoulin, Aaron C. Courville
NIPS2
2015 Optimizing Expected Intersection-Over-Union with Candidate-Constrained CRFs
abstract
We study the question of how to make loss-aware predictions in image segmentation settings where the evaluation function is the Intersection-over-Union (IoU) measure that is used widely in evaluating image segmentation systems. Currently, there are two dominant approaches: the first approximates the Expected-IoU (EIoU) score as Expected-Intersection-over-Expected-Union (EIoEU), and the second approach is to compute exact EIoU but only over a small set of high-quality candidate solutions. We begin by asking which approach we should favor for two typical image segmentation tasks. Studying this question leads to two new methods that draw ideas from both existing approaches. Our new methods use the EIoEU approximation paired with high quality candidate solutions. Experimentally we show that our new approaches lead to improved performance on both image segmentation tasks.
Faruk Ahmed, Daniel Tarlow, Dhruv Batra
ICCV1
2014 Removal of High-Density Salt-and-Pepper Noise in Images With an Iterative Adaptive Fuzzy Filter Using Alpha-Trimmed Mean
abstract
Suppression of impulse noise in images is an important problem in image processing. In this paper, we propose a novel adaptive iterative fuzzy filter for denoising images corrupted by impulse noise. It operates in two stages-detection of noisy pixels with an adaptive fuzzy detector followed by denoising using a weighted mean filter on the “good” pixels in the filter window. Experimental results demonstrate the algorithm to be superior to state-of-the-art filters. The filter is also shown to be robust to very high levels of noise, retrieving meaningful detail at noise levels as high as 97%.
Faruk Ahmed, Swagatam Das
IEEE Trans. Fuzzy Syst.1
2008 i-SEGOPubmed: a web interface for semantic enabled browsing of PubMed using Gene Ontology
Mohammed Yeasin, Bhanu Vanteru, Jahangheer S. Shaik, Faruk Ahmed
BMC Bioinform.4