EDBT 2026 Demo / reviewers in the wild / expert
M. Saquib Sarfraz
dblp:12/1561 · also Muhammad Saquib Sarfraz
· DBLP profile ↗
39ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-1271-0005ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 12 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
Kunyu Peng, Di Wen 0006, M. Saquib Sarfraz, Yufan Chen 0001, Junwei Zheng, David Schneider 0006, Kailun Yang 0001, Alina Roitberg, Rainer Stiefelhagen |
Int. J. Comput. Vis. | 3 |
| 2025 | Is Visual in-Context Learning for Compositional Medical Tasks Within Reach?abstractIn this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed. Simon Reiß, Zdravko Marinov, Alexander Jaus, Constantin Seibold, M. Saquib Sarfraz, Erik Rodner, Rainer Stiefelhagen |
ICCV | 5 |
| 2024 | Navigating Open Set Scenarios for Skeleton-Based Action RecognitionabstractIn real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and the distinct sparse structure of body pose sequences. In this paper, we tackle the unexplored Open-Set Skeleton-based Action Recognition (OS-SAR) task and formalize the benchmark on three skeleton-based datasets. We assess the performance of seven established open-set approaches on our task and identify their limits and critical generalization issues when dealing with skeleton information.To address these challenges, we propose a distance-based cross-modality ensemble method that leverages the cross-modal alignment of skeleton joints, bones, and velocities to achieve superior open-set recognition performance. We refer to the key idea as CrossMax - an approach that utilizes a novel cross-modality mean max discrepancy suppression mechanism to align latent spaces during training and a cross-modality distance-based logits refinement method during testing. CrossMax outperforms existing approaches and consistently yields state-of-the-art results across all datasets and backbones. We will release the benchmark, code, and models to the community. Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, David Schneider 0006, Jiaming Zhang 0001, Kailun Yang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
AAAI | 8 |
| 2024 | Improving Single Domain-Generalized Object Detection: A Focus on Diversification and AlignmentabstractIn this work, we tackle the problem of domain generalization for object detection, specifically focusing on the scenario where only a single source domain is available. We propose an effective approach that involves two key steps: diversifying the source domain and aligning detections based on class prediction confidence and localization. Firstly, we demonstrate that by carefully selecting a set of augmentations, a base detector can outperform existing methods for single domain generalization by a good margin. This highlights the importance of domain diversification in improving the performance of object detectors. Secondly, we introduce a method to align detections from multiple views, considering both classification and localization outputs. This alignment procedure leads to better generalized and well-calibrated object detector models, which are crucial for accurate decision-making in safety-critical applications. Our approach is detector-agnostic and can be seamlessly applied to both single-stage and two-stage detectors. To validate the effectiveness of our proposed methods, we conduct extensive experiments and ablations on challenging domain-shift scenarios. The results consistently demonstrate the superiority of our approach compared to existing methods. Our code and models are available at: https://github.com/msohaildanishIDivAlign. Muhammad Sohail Danish, Muhammad Haris Khan, Muhammad Akhtar Munir, M. Saquib Sarfraz, Mohsen Ali |
CVPR | 4 |
| 2024 | Referring Atomic Video Action Recognition
Kunyu Peng, Jia Fu 0001, Kailun Yang 0001, Di Wen 0006, Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
ECCV (19) | 9 |
| 2024 | AltChart: Enhancing VLM-Based Chart Summarization Through Multi-pretext Tasks
Omar Moured, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen |
ICDAR (1) | 3 |
| 2024 | Position: Quo Vadis, Unsupervised Time Series Anomaly Detection?abstractThe current state of machine learning scholarship in Timeseries Anomaly Detection (TAD) is plagued by the persistent use of flawed evaluation metrics, inconsistent benchmarking practices, and a lack of proper justification for the choices made in novel deep learning-based model designs. Our paper presents a critical analysis of the status quo in TAD, revealing the misleading track of current research and highlighting problematic methods, and evaluation practices. Our position advocates for a shift in focus from solely pursuing novel model designs to improving benchmarking practices, creating non-trivial datasets, and critically evaluating the utility of complex methods against simpler baselines. Our findings demonstrate the need for rigorous evaluation protocols, the creation of simple baselines, and the revelation that state-of-the-art deep anomaly detection models effectively learn linear mappings. These findings suggest the need for more exploration and development of simple and interpretable TAD methods. The increment of model complexity in the state-of-the-art deep-learning based models unfortunately offers very little improvement. We offer insights and suggestions for the field to move forward. M. Saquib Sarfraz, Mei-Yen Chen, Lukas Layer, Kunyu Peng, Marios Koulakis |
ICML | 1 |
| 2024 | Fourier Prompt Tuning for Modality-Incomplete Scene SegmentationabstractIntegrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework. However, the modality incompleteness in multi-modal segmentation remains under-explored. In this work, we establish a task called Modality-Incomplete Scene Segmentation (MISS), which encompasses both system-level modality absence and sensor-level modality errors. To avoid the predominant modality reliance in multi-modal fusion, we introduce a Missing-aware Modal Switch (MMS) strategy to proactively manage missing modalities during training. Utilizing bit-level batch-wise sampling enhances the model’s performance in both complete and incomplete testing scenarios. Furthermore, we introduce the Fourier Prompt Tuning (FPT) method to incorporate representative spectral information into a limited number of learnable prompts that maintain robustness against all MISS scenarios. Akin to fine-tuning effects but with fewer tunable parameters (1.1%). Extensive experiments prove the efficacy of our proposed approach, showcasing an improvement of 5.84% mIoU over the prior state-of-the-art parameter-efficient methods in modality missing. The source code is publicly available at https://github.com/RuipingL/MISS. Ruiping Liu 0001, Jiaming Zhang 0001, Kunyu Peng, Yufan Chen 0001, Junwei Zheng, M. Saquib Sarfraz, Kailun Yang 0001, Rainer Stiefelhagen |
IV | 7 |
| 2024 | Towards Video-based Activated Muscle Group Estimation in the WildabstractIn this paper, we tackle the new task of video-based Activated Muscle Group Estimation (AMGE) aiming at identifying active muscle regions during physical activity in the wild.To this intent, we provide the MuscleMap dataset featuring >15𝐾 video clips with 135 different activities and 20 labeled muscle groups.This dataset opens the vistas to multiple video-based applications in sports and rehabilitation medicine under flexible environment constraints.The proposed MuscleMap dataset is constructed with YouTube videos, specifically targeting High-Intensity Interval Training (HIIT) physical exercise in the wild.To make the AMGE model applicable in real-life situations, it is crucial to ensure that the model can generalize well to numerous types of physical activities not present during training and involving new combinations of activated muscles.To achieve this, our benchmark also covers an evaluation setting where the model is exposed to activity types excluded from the training set.Our experiments reveal that the generalizability of existing architectures adapted for the AMGE task remains a challenge.Therefore, we also propose a new approach, TransM 3 E, which employs a multi-modality feature fusion mechanism between both the video transformer model and the skeleton-based graph convolution model with novel cross-modal knowledge distillation executed on multiclassification tokens.The proposed method surpasses all popular video classification models when dealing with both, previously seen and new types of physical activities.The database and code can be found at https://github.com/KPeng9510/MuscleMap. Kunyu Peng, David Schneider 0006, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Chen Deng, M. Saquib Sarfraz, Rainer Stiefelhagen |
ACM Multimedia | 8 |
| 2024 | Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain SchedulerabstractIn Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and accurately quantify category novelty, which is critical for applications in dynamic environments. Recently, meta-learning techniques have demonstrated superior results in OSDG, effectively orchestrating the meta-train and -test tasks by employing varied random categories and predefined domain partition strategies. These approaches prioritize a well-designed training schedule over traditional methods that focus primarily on data augmentation and the enhancement of discriminative feature learning.
The prevailing meta-learning models in OSDG typically utilize a predefined sequential domain scheduler to structure data partitions. However, a crucial aspect that remains inadequately explored is the influence brought by strategies of domain schedulers during training.
In this paper, we observe that an adaptive domain scheduler benefits more in OSDG compared with prefixed sequential and random domain schedulers. We propose the Evidential Bi-Level Hardest Domain Scheduler (EBiL-HaDS) to achieve an adaptive domain scheduler. This method strategically sequences domains by assessing their reliabilities in utilizing a follower network, trained with confidence scores learned in an evidential manner, regularized by max rebiasing discrepancy, and optimized in a bilevel manner. We verify our approach on three OSDG benchmarks, i.e., PACS, DigitsDG, and OfficeHome. The results show that our method substantially improves OSDG performance and achieves more discriminative embeddings for both the seen and unseen categories, underscoring the advantage of a judicious domain scheduler for the generalizability to unseen domains and unseen categories. The source code is publicly available at https://github.com/KPeng9510/EBiL-HaDS. Kunyu Peng, Di Wen 0006, Kailun Yang 0001, Ao Luo, Yufan Chen 0001, Jia Fu 0001, M. Saquib Sarfraz, Alina Roitberg, Rainer Stiefelhagen |
NeurIPS | 7 |
| 2024 | Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle ActivationsabstractExploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of datasets.In this work, we address this issue by establishing Muscles in Time (MinT), a large-scale synthetic muscle activation dataset.For the creation of MinT, we enriched existing motion capture datasets by incorporating muscle activation simulations derived from biomechanical human body models using the OpenSim platform, a common framework used in biomechanics and human motion research.Starting from simple pose sequences, our pipeline enables us to extract detailed information about the timing of muscle activations within the human musculoskeletal system.Muscles in Time contains over nine hours of simulation data covering 227 subjects and 402 simulated muscle strands. We demonstrate the utility of this dataset by presenting results on neural network-based muscle activation estimation from human pose sequences with two different sequence-to-sequence architectures. David Schneider 0006, Simon Reiß, Marco Kugler, Alexander Jaus, Kunyu Peng, Susanne Sutschet, M. Saquib Sarfraz, Sven Matthiesen, Rainer Stiefelhagen |
NeurIPS | 7 |
| 2024 | Recognizing affective states from the expressive behavior of tennis players using convolutional neural networksabstractThis study describes an AI model by leveraging advanced Convolutional Neural Networks (CNNs) to recognize affective states in real-world sports settings, particularly tennis matches. In contrast to prior studies that primarily utilized data acquired from actors and rudimentary statistical methods, the present research emphasizes the analysis of bodily expressions in real-life contexts, aiming for a more naturalistic representation of human emotions. Our CNN-based models demonstrate an accuracy rate of up to 68.9%, outperforming or matching human observers in many instances. Intriguingly, both the machine learning models and human observers exhibited a shared propensity to more effectively identify negative affective states, which may be attributed to the more intense and straightforward expression of these states. These results not only advance the state of the art in affective state recognition but also pave the way for broader applications, including in healthcare and automotive safety sectors, thereby constituting a significant advancement in the development of sophisticated and universally applicable emotional recognition systems. Darko Jekauc, Diana Burkart, Julian Fritsch, Marc Hesenius, Ole Meyer, M. Saquib Sarfraz, Rainer Stiefelhagen |
Knowl. Based Syst. | 6 |
| 2023 | Domain Adaptive Object Detection via Balancing Between Self-Training and Adversarial LearningabstractDeep learning based object detectors struggle generalizing to a new target domain bearing significant variations in object and background. Most current methods align domains by using image or instance-level adversarial feature alignment. This often suffers due to unwanted background and lacks class-specific alignment. A straightforward approach to promote class-level alignment is to use high confidence predictions on unlabeled domain as pseudo-labels. These predictions are often noisy since model is poorly calibrated under domain shift. In this paper, we propose to leverage model's predictive uncertainty to strike the right balance between adversarial feature alignment and class-level alignment. We develop a technique to quantify predictive uncertainty on class assignments and bounding-box predictions. Model predictions with low uncertainty are used to generate pseudo-labels for self-training, whereas the ones with higher uncertainty are used to generate tiles for adversarial feature alignment. This synergy between tiling around uncertain object regions and generating pseudo-labels from highly certain object regions allows capturing both image and instance-level context during the model adaptation. We report thorough ablation study to reveal the impact of different components in our approach. Results on five diverse and challenging adaptation scenarios show that our approach outperforms existing state-of-the-art methods with noticeable margins. Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon S. Kim 0002, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen |
BMVC | 3 |
| 2022 | Hierarchical Nearest Neighbor Graph Embedding for Efficient Dimensionality ReductionabstractDimensionality reduction is crucial both for visualization and preprocessing high dimensional data for machine learning. We introduce a novel method based on a hierarchy built on 1-nearest neighbor graphs in the original space which is used to preserve the grouping properties of the data distribution on multiple levels. The core of the proposal is an optimization-free projection that is competitive with the latest versions of t-SNE and UMAP in performance and visualization quality while being an order of magnitude faster at run-time. Furthermore, its interpretable mechanics, the ability to project new data, and the natural separation of data clusters in visualizations make it a general purpose unsupervised dimension reduction technique. In the paper, we argue about the soundness of the proposed method and evaluate it on a diverse collection of datasets with sizes varying from 1 K to 11M samples and dimensions from 28 to 16K. We perform comparisons with other state-of-the-art methods on multiple metrics and target dimensions high-lighting its efficiency and performance. Code is available at https://github.com/koulakis/h-nne M. Saquib Sarfraz, Marios Koulakis, Constantin Seibold, Rainer Stiefelhagen |
CVPR | 1 |
| 2022 | Breaking with Fixed Set Pathology Recognition Through Report-Guided Contrastive Training
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (5) | 3 |
| 2022 | Towards Improving Calibration in Object Detection Under Domain ShiftabstractWith deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibration have been proposed, they are only limited to visual classification applications and in-domain predictions. Unfortunately, very little to no attention is paid towards addressing calibration of DNN-based visual object detectors, that occupy similar space and importance in many decision making systems as their visual classification counterparts. In this work, we study the calibration of DNN-based object detection models, particularly under domain shift. To this end, we first propose a new, plug-and-play, train-time calibration loss for object detection (coined as TCD). It can be used with various application-specific loss functions as an auxiliary loss function to improve detection calibration. Second, we devise a new implicit technique for improving calibration in self-training based domain adaptive detectors, featuring a new uncertainty quantification mechanism for object detection. We demonstrate TCD is capable of enhancing calibration with notable margins (1) across different DNN-based object detection paradigms both in in-domain and out-of-domain predictions, and (2) in different domain-adaptive detectors across challenging adaptation scenarios. Finally, we empirically show that our implicit calibration technique can be used in tandem with TCD during adaptation to further boost calibration in diverse domain shift scenarios. Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali |
NeurIPS | 3 |
| 2021 | Temporally-Weighted Hierarchical Clustering for Unsupervised Action SegmentationabstractAction segmentation refers to inferring boundaries of semantically consistent visual concepts in videos and is an important requirement for many video understanding tasks. For this and other video understanding tasks, supervised approaches have achieved encouraging performance but require a high volume of detailed frame-level annotations. We present a fully automatic and unsupervised approach for segmenting actions in a video that does not require any training. Our proposal is an effective temporally-weighted hierarchical clustering algorithm that can group semantically consistent frames of the video. Our main finding is that representing a video with a 1-nearest neighbor graph by taking into account the time progression is sufficient to form semantically and temporally consistent clusters of frames where each cluster may represent some action in the video. Additionally, we establish strong unsupervised baselines for action segmentation and show significant performance improvements over published unsupervised methods on five challenging action segmentation datasets. Our code is available.1 M. Saquib Sarfraz, Naila Murray, Vivek Sharma 0001, Ali Diba, Luc Van Gool, Rainer Stiefelhagen |
CVPR | 1 |
| 2021 | Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationabstractIn this paper, we introduce a novel self-supervised visual representation learning method which understands both images and videos in a joint learning fashion. The proposed neural network architecture and objectives are designed to obtain two different Convolutional Neural Networks for solving visual recognition tasks in the domain of videos and images. Our method called Video/Image for Visual Contrastive Learning of Representation(Vi2CLR) uses unlabeled videos to exploit dynamic and static visual cues for self-supervised and instances similarity/dissimilarity learning. Vi2CLR optimization pipeline consists of visual clustering part and representation learning based on groups of similar positive instances within a cluster and negative ones from other clusters and learning visual clusters and their distances. We show how a joint self-supervised visual clustering and instance similarity learning with 2D (image) and 3D (video) CovNet encoders yields such robust and near to supervised learning performance.We extensively evaluate the method on downstream tasks like large scale action recognition, image and object classification on datasets like Kinetics, ImageNet, Pascal VOC’07 and UCF101 and achieve outstanding results compared to state-of-the-art self-supervised methods. Ali Diba, Vivek Sharma 0001, Reza Safdari, Dariush Lotfi, M. Saquib Sarfraz, Rainer Stiefelhagen, Luc Van Gool |
ICCV | 5 |
| 2021 | Deep Learning Based Oil Spill Classification Using Unet Convolutional Neural NetworkabstractOil spills cause a significant threat to marine and coastal ecosystems. It is one of the major causes of water pollution. This research focuses on the use of deep learning for oil spills detection and classification. UNet is a convolutional neural network, originally proposed for biomedical image segmentation and modified for the discrimination of oil spills and look-alikes. The model is trained on a publicly available benchmark oil spill detection dataset of Sentinel-1 synthetic aperture radar (SAR) images. The images have been semantically segmented into multiple regions of interest such as sea surface, oil spills, look-alikes, ships and land. The proposed UNet-based model achieves intersection over union (IoU) value of 95.69% for sea surface, 60.85% for oil spills, 54.90% for look-alikes, 70.27% for ships and 96.79% for land class. The mean intersection over union (mIoU) value for all the classes is 75.70% which consitutes a nearly 10% increase compared to state of the art for this dataset. Abdul Basit 0019, Muhammad Adnan Siddique, M. Saquib Sarfraz |
IGARSS | 3 |
| 2021 | SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object DetectionabstractWe study adapting trained object detectors to unseen domains manifesting significant variations of object appearance, viewpoints and backgrounds. Most current methods align domains by either using image or instance-level feature alignment in an adversarial fashion. This often suffers due to the presence of unwanted background and as such lacks class-specific alignment. A common remedy to promote class-level alignment is to use high confidence predictions on the unlabelled domain as pseudo labels. These high confidence predictions are often fallacious since the model is poorly calibrated under domain shift. In this paper, we propose to leverage model’s predictive uncertainty to strike the right balance between adversarial feature alignment and class-level alignment. Specifically, we measure predictive uncertainty on class assignments and the bounding box predictions. Model predictions with low uncertainty are used to generate pseudo-labels for self-supervision, whereas the ones with higher uncertainty are used to generate tiles for an adversarial feature alignment stage. This synergy between tiling around the uncertain object regions and generating pseudo-labels from highly certain object regions allows us to capture both the image and instance level context during the model adaptation stage. We perform extensive experiments covering various domain shift scenarios. Our approach improves upon existing state-of-the-art methods with visible margins. Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali |
NeurIPS | 3 |
| 2021 | Unsupervised Meta-Domain Adaptation for Fashion RetrievalabstractCross-domain fashion item retrieval naturally arises when unconstrained consumer images are used to query for fashion items in a collection of high-quality photographs provided by retailers. To perform this task, approaches typically leverage both consumer and shop domains from a given dataset to learn a domain invariant representation, allowing these images of different nature to be directly compared. When consumer images are not available beforehand, such training is impossible. In this paper, we focus on this challenging and yet practical scenario, and we propose instead to leverage representations learned for cross-domain retrieval from another source dataset and to adapt them to the target dataset for this particular setting. More precisely, we bypass the lack of consumer images and directly target the more challenging meta-domain gap which occurs between consumer images and shop images, independently of their dataset. Assuming that datasets share some similar fashion items, we cluster their shop images and leverage the clusters to automatically generate pseudo-labels. Those are used to associate consumer and shop images across datasets, which in turn allows to learn meta-domain-invariant representations suitable for cross-domain retrieval in the target dataset. The features and code are available at https://github.com/vivoutlaw/UDMA. Vivek Sharma 0001, Naila Murray, Diane Larlus, M. Saquib Sarfraz, Rainer Stiefelhagen, Gabriela Csurka |
WACV | 4 |
| 2020 | Anchor-free Small-scale Multispectral Pedestrian Detection
Alexander Wolpert, Michael Teutsch, M. Saquib Sarfraz, Rainer Stiefelhagen |
BMVC | 3 |
| 2020 | Clustering based Contrastive Learning for Improving Face RepresentationsabstractA good clustering algorithm can discover natural groupings in data. These groupings, if used wisely, provide a form of weak supervision for learning representations. In this work, we present Clustering-based Contrastive Learning (CCL), a new clustering-based representation learning approach that uses labels obtained from clustering along with video constraints to learn discriminative face features. We demonstrate our method on the challenging task of learning representations for video face clustering. Through several ablation studies, we analyze the impact of creating pair-wise positive and negative labels from different sources. Experiments on three challenging video face clustering datasets: BBT-0101, BF-0502, and ACCIO show that CCL achieves a new state-of-the-art on all datasets. Vivek Sharma 0001, Makarand Tapaswi, M. Saquib Sarfraz, Rainer Stiefelhagen |
FG | 3 |
| 2019 | Content and Colour Distillation for Learning Image Translations with the Spatial Profile Loss
M. Saquib Sarfraz, Constantin Seibold, Haroon Khalid, Rainer Stiefelhagen |
BMVC | 1 |
| 2019 | Efficient Parameter-Free Clustering Using First Neighbor RelationsabstractWe present a new clustering method in the form of a single clustering equation that is able to directly discover groupings in the data. The main proposition is that the first neighbor of each sample is all one needs to discover large chains and finding the groups in the data. In contrast to most existing clustering algorithms our method does not require any hyper-parameters, distance thresholds and/or the need to specify the number of clusters. The proposed algorithm belongs to the family of hierarchical agglomerative methods. The technique has a very low computational overhead, is easily scalable and applicable to large practical problems. Evaluation on well known datasets from different domains ranging between 1077 and 8.1 million samples shows substantial performance gains when compared to the existing clustering techniques. M. Saquib Sarfraz, Vivek Sharma 0001, Rainer Stiefelhagen |
CVPR | 1 |
| 2019 | Self-Supervised Learning of Face Representations for Video Face ClusteringabstractAnalyzing the story behind TV series and movies often requires understanding who the characters are and what they are doing. With improving deep face models, this may seem like a solved problem. However, as face detectors get better, clustering/identification needs to be revisited to address increasing diversity in facial appearance. In this paper, we address video face clustering using unsupervised methods. Our emphasis is on distilling the essential information, identity, from the representations obtained using deep pre-trained face networks. We propose a self-supervised Siamese network that can be trained without the need for video/track based supervision, and thus can also be applied to image collections. We evaluate our proposed method on three video face clustering datasets. The experiments show that our methods outperform current state-of-the-art methods on all datasets. Video face clustering is lacking a common benchmark as current works are often evaluated with different metrics and/or different sets of face tracks. The datasets and code are available at https://github.com/vivoutlaw/SSIAM. Vivek Sharma 0001, Makarand Tapaswi, M. Saquib Sarfraz, Rainer Stiefelhagen |
FG | 3 |
| 2018 | A Pose-Sensitive Embedding for Person Re-Identification With Expanded Cross Neighborhood Re-RankingabstractPerson re-identification is a challenging retrieval task that requires matching a person's acquired image across non-overlapping camera views. In this paper we propose an effective approach that incorporates both the fine and coarse pose information of the person to learn a discriminative embedding. In contrast to the recent direction of explicitly modeling body parts or correcting for misalignment based on these, we show that a rather straightforward inclusion of acquired camera view and/or the detected joint locations into a convolutional neural network helps to learn a very effective representation. To increase retrieval performance, re-ranking techniques based on computed distances have recently gained much attention. We propose a new unsupervised and automatic re-ranking framework that achieves state-of-the-art re-ranking performance. We show that in contrast to the current state-of-the-art re-ranking methods our approach does not require to compute new rank lists for each image pair (e.g., based on reciprocal neighbors) and performs well by using simple direct rank list based comparison or even by just using the already computed euclidean distances between the images. We show that both our learned representation and our re-ranking method achieve state-of-the-art performance on a number of challenging surveillance image and video datasets. Code is available at https://github.com/pse-ecn. M. Saquib Sarfraz, Arne Schumann, Andreas Eberle, Rainer Stiefelhagen |
CVPR | 1 |
| 2017 | Deep View-Sensitive Pedestrian Attribute Inference in an end-to-end Model
M. Saquib Sarfraz, Arne Schumann, Rainer Stiefelhagen |
BMVC | 1 |
| 2017 | Heterogeneous Face Recognition: Recent Advances in Infrared-to-Visible MatchingabstractAn emerging topic in face recognition is matching between facial images acquired from different sensing modalities, referred to as heterogeneous face recognition. Heterogeneous face recognition has the potential to provide key capabilities for the commercial sector as well as for law enforcement, intelligence gathering, and the military, especially in challenging unconstrained settings. However, the difficulty in heterogeneous face recognition is compounded by phenomenology differences between modalities, giving rise to significant facial appearance variations due to the modality gap. In this paper, we focus on a subset of heterogeneous face recognition and present a succinct review of recent work on infrared-to-visible face recognition. Shuowen Hu, Nathan J. Short, Benjamin S. Riggan, Matthew Chasse, M. Saquib Sarfraz |
FG | 5 |
| 2017 | Deep Perceptual Mapping for Cross-Modal Face Recognition
M. Saquib Sarfraz, Rainer Stiefelhagen |
Int. J. Comput. Vis. | 1 |
| 2015 | Deep Perceptual Mapping for Thermal to Visible Face RecogntionabstractCross modal face matching between the thermal and visible spectrum is a much desired capability for night-time surveillance and security applications. Due to a very large modality gap, thermal-to-visible face recognition is one of the most challenging face matching problem. In this paper, we present an approach to bridge this modality gap by a significant margin. Our approach captures the highly non-linear relationship between the two modalities by using a deep neural network. Our model attempts to learn a non-linear mapping from visible to thermal spectrum while preserving the identity information. We show substantive performance improvement on a difficult thermal-visible face dataset. The presented approach improves the state-of-the-art by more than 10% in terms of Rank-1 identification and bridge the drop in performance due to the modality gap by more than 40%. M. Saquib Sarfraz, Rainer Stiefelhagen |
BMVC | 1 |
| 2015 | Exploiting colour information for better scene text detection and recognition
Muhammad Fraz, M. Saquib Sarfraz, Eran A. Edirisinghe |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | Exploiting Color Information for Better Scene Text Recognition
Muhammad Fraz, M. Saquib Sarfraz, Eran A. Edirisinghe |
BMVC | 2 |
| 2014 | Mid-level-Representation Based Lexicon for Vehicle Make and Model RecognitionabstractIn this paper, we present a novel framework for representation of images as a combination of multiple mid-level feature descriptor representation based group of visual words. The mid-level feature representation is computed on discriminative patches of the image to build a lexicon, the visual words of which are used to represent the shape within that image. The proposed image representation method has been applied to the application of vehicles make and model recognition. Each make, model class is represented as an over complete sub-lexicon of mid-level feature representation. The classification of vehicles is performed by comparing the visual words of probe image with the learned lexicon of training data using Euclidean distance. The proposed framework offers the advantage of accurate recognition in the presence of significant background clutter. The experiments have shown that the proposed representation successfully captures the fine-grained inter and intra-class discrimination to recognize the model and make of the vehicle without any strict requirement of precise region of interest segmentation. Another important contribution of the paper is a comprehensive dataset of cars depicting images collected in the wild. Muhammad Fraz, Eran A. Edirisinghe, M. Saquib Sarfraz |
ICPR | 3 |
| 2013 | RPM: Random Points Matching for Pair wise Face-Similarity
M. Saquib Sarfraz, Muhammad Adnan Siddique, Rainer Stiefelhagen |
BMVC | 1 |
| 2012 | Automatic registration of SAR and optical images based on mutual information assisted Monte CarloabstractThe development of Geographical Information Systems applications involving fusion of data from different space-borne imaging sensors inevitably requires a preliminary registration of the images. In case of Synthetic Aperture Radar (SAR) and optical sensors, the registration is particularly challenging due to the vast radiometric differences in the data. In this paper, we present a novel method to register SAR and optical images automatically. It provides an accurate registration despite the radiometric differences in the images. Moreover, this paper introduces a Monte Carlo formulation of the image registration problem. Muhammad Adnan Siddique, M. Saquib Sarfraz, David Bornemann, Olaf Hellwich |
IGARSS | 2 |
| 2010 | Probabilistic learning for fully automatic face recognition across pose
M. Saquib Sarfraz, Olaf Hellwich |
Image Vis. Comput. | 1 |
| 2008 | Statistical appearance models for automatic pose invariant face recognitionabstractRecent pose invariant methods try to model the subject specific appearance change across pose. For this, however, almost all of the existing methods require a perfect alignment between a gallery and a probe image. In this paper we present a pose invariant face recognition method, centered on modeling joint appearance of gallery and probe images across pose, that do not require the facial landmarks to be detected as such. We propose novel extensions by introducing to use a more robust feature description as opposed to pixel-based appearances. Using such features we put forward to synthesize the non-frontal views to frontal. Furthermore, using local kernel density estimation, instead of commonly used normal density assumption, is suggested to derive the prior models. Our method does not require any strict alignment between gallery and probe images which makes it particularly attractive as compared to the existing state of the art methods. Improved recognition across a wide range of poses has been achieved using these extensions. M. Saquib Sarfraz, Olaf Hellwich |
FG | 1 |