VLDB 2026 Research / reviewers in the wild / expert
Majid Mirmehdi
dblp:79/2375
· DBLP profile ↗
110ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-6478-1403ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 9 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 8Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization
Qiushuo Cheng, Catherine Morgan, Alan L. Whone, Majid Mirmehdi |
ICPR (1) | 5 |
| 2026 | Deep in the Jungle: Towards Automating Chimpanzee Population Estimation
Tom Raynes, Otto Brookes, Timm Haucke, Lukas Boesch, Anne-Sophie Crunchant, Hjalmar S. Kühl, Sara Beery, Majid Mirmehdi, Tilo Burghardt |
ICPR (15) | 8 |
| 2026 | GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model - Bringing Motion Generation to the Clinical DomainabstractGait analysis is crucial for the diagnosis and monitoring of movement disorders like Parkinson's Disease. While computer vision models have shown potential for objectively evaluating parkinsonian gait, their effectiveness is limited by scarce clinical datasets and the challenge of collecting large and well-labelled data, impacting model accuracy and risk of bias. To address these gaps, we propose GAITGen, a novel framework that generates realistic gait sequences conditioned on specified pathology severity levels. GAITGen employs a Conditional Residual Vector Quantized Variational Autoencoder to learn disentangled representations of motion dynamics and pathology-specific factors, coupled with Mask and Residual Transformers for conditioned sequence generation. GAITGen generates realistic, diverse gait sequences across severity levels, enriching datasets and enabling large-scale model training in parkinsonian gait analysis. Experiments on our new PD-GaM (real) dataset demonstrate that GAITGen outperforms adapted state-of-the-art models in both reconstruction fidelity and generation quality, accurately capturing critical pathology-specific gait features. A clinical user study confirms the realism and clinical relevance of our generated sequences. Moreover, incorporating GAITGen-generated data into downstream tasks improves parkinsonian gait severity estimation, highlighting its potential for advancing clinical gait analysis. Vida Adeli, Soroush Mehraban, Majid Mirmehdi, Alan L. Whone, Benjamin Filtjens, Amirhossein Dadashzadeh, Alfonso Fasano, Andrea Iaboni, Babak Taati |
WACV | 3 |
| 2026 | Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain AdaptationabstractRecent advances in Source-Free Unsupervised Video Domain Adaptation (SFUVDA) leverage vision-language models to enhance pseudo-label generation. However, challenges such as noisy pseudo-labels and over-confident predictions limit their effectiveness in adapting well across domains. We propose Co-STAR, a novel framework that integrates curriculum learning with collaborative self-training between a source-trained teacher and a contrastive vision-language model (CLIP). Our curriculum learning approach employs a reliability-based weight function that measures bidirectional prediction alignment between the teacher and CLIP, balancing between confident and uncertain predictions. This function preserves uncertainty for difficult samples, while prioritizing reliable pseudo-labels when the predictions from both models closely align. To further improve adaptation, we propose Adaptive Curriculum Regularization, which modifies the learning priority of samples in a probabilistic, adaptive manner based on their confidence scores and prediction stability, mitigating overfitting to noisy and over-confident samples. Extensive experiments across multiple video domain adaptation benchmarks demonstrate that Co-STAR consistently outperforms state-of-the-art SFUVDA methods. Code is available at: https://github.com/Plrbear/Co-Star Amirhossein Dadashzadeh, Parsa Esmati, Majid Mirmehdi |
WACV | 3 |
| 2026 | When Visual Privacy Protection Meets Multimodal Large Language ModelsabstractAbstract The emergence of Multimodal Large Language Models (MLLMs) and the widespread usage of MLLM cloud services such as GPT-4V raised great concerns about privacy leakage in visual data. As these models are typically deployed in cloud services, users are required to submit their images and videos, posing serious privacy risks. However, how to tackle such privacy concerns is an under-explored problem. Thus, in this paper, we aim to conduct a new investigation to protect visual privacy when enjoying the convenience brought by MLLM services. We address the practical case where the MLLM is a “black box”, i.e., we only have access to its input and output without knowing its internal model information. To tackle such a challenging yet demanding problem, we propose a novel framework, in which we carefully design the learning objective with Pareto optimality to seek a better trade-off between visual privacy and MLLM’s performance, and propose critical-history enhanced optimization to effectively optimize the framework with the black-box MLLM. Our experiments show that our method is effective on different benchmarks. Xiaofei Hui, Haoxuan Qu, Majid Mirmehdi, Hossein Rahmani 0001, Jun Liu 0036 |
Int. J. Comput. Vis. | 4 |
| 2026 | Unsupervised Cross-Domain 3D Human Pose Estimation via Pseudo-Label-Guided Global TransformsabstractExisting 3D human pose estimation methods often suffer in performance, when applied to cross-scenario inference, due to domain shifts in characteristics such as camera viewpoint, position, posture, and body size. Among these factors, camera viewpoints and locations have been shown to contribute significantly to the domain gap by influencing the global positions of human poses. To address this, we propose a novel framework that explicitly conducts global transformations between pose positions in the camera coordinate systems of source and target domains. We start with a Pseudo-Label Generation Module that is applied to the 2D poses of the target dataset to generate pseudo-3D poses. Then, a Global Transformation Module leverages a human-centered coordinate system as a novel bridging mechanism to seamlessly align the positional orientations of poses across disparate domains, ensuring consistent spatial referencing. To further enhance generalization, a Pose Augmentor is incorporated to address variations in human posture and body size. This process is iterative, allowing refined pseudo-labels to progressively improve guidance for domain adaptation. Our method is evaluated on various cross-dataset benchmarks, including Human3.6M, MPI-INF-3DHP, and 3DPW. The proposed method outperforms state-of-the-art approaches and even outperforms the target-trained model. Zhiyong Wang 0009, Xinyu Fan 0009, Amirhossein Dadashzadeh, Honghai Liu 0001, Majid Mirmehdi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionabstractComputer vision analysis of camera trap video footage is essential for wildlife conservation, as captured behaviours offer some of the earliest indicators of changes in population health. Recently, several high-impact animal behaviour datasets and methods have been introduced to encourage their use; however, the role of behaviour-correlated background information and its significant effect on out-of-distribution generalisation remain unexplored. In response, we present the PanAf-FGBG dataset, featuring 21 hours of wild chimpanzee behaviours, recorded at 389 individual camera locations. Uniquely, it pairs every video with a chimpanzee (referred to as a foreground video) with a corresponding background video (with no chimpanzee) from the same camera location. We present two views of the dataset: one with overlapping camera locations and one with disjoint locations. This setup enables, for the first time, direct evaluation of in-distribution and out-of-distribution conditions, and for the impact of backgrounds on behaviour recognition models to be quantified. All clips have rich behavioural annotations and metadata, including unique camera IDs. Additionally, we establish several baselines and present a highly effective latent-space normalisation technique that boosts out-of-distribution performance by +5.42% mAP for convolutional and +3.75% mAP for transformer-based models. Finally, we provide an in-depth analysis on the role of backgrounds in out-of-distribution behaviour recognition, including the so far unexplored impact of background durations. (i.e., the count of background frames within foreground videos). The dataset is available at https://obrookes.github.io/panaf-fgbg.github.io/. Otto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens, Paula Dieguez, Thurston C. Hicks, Sorrel Jones, Maureen S. McCarthy, Amelia Meier, Emmanuelle Normand, Erin G. Wessling, Roman M. Wittig, Kevin Langergraber, Klaus Zuberbühler, Lukas Boesch, Thomas Schmid 0003, Mimi Arandjelovic, Hjalmar S. Kühl, Tilo Burghardt |
CVPR | 3 |
| 2025 | Trajectory-guided Motion Perception for Facial Expression Quality Assessment in Neurological DisordersabstractAutomated facial expression quality assessment (FEQA) in neurological disorders is critical for enhancing diagnostic accuracy and improving patient care, yet effectively capturing the subtle motions and nuances of facial muscle movements remains a challenge. We propose to analyse facial landmark trajectories, a compact yet informative representation, that encodes these subtle motions from a high-level structural perspective. Hence, we introduce Trajectory-guided Motion Perception Transformer (TraMP-Former), a novel FEQA framework that fuses landmark trajectory features for fine-grained motion capture with visual semantic cues from RGB frames, ultimately regressing the combined features into a quality score. Extensive experiments demonstrate that TraMP-Former achieves new state-of-the-art performance on bench-mark datasets with neurological disorders, including PFED5 ($\uparrow$ $6.51 \%$) and an augmented Toronto NeuroFace ($\uparrow 7.62 \%$). Our ablation studies further validate the efficiency and effectiveness of landmark trajectories in FEQA. Our code is available at https://github.com/shuchaoduan/TraMP-Former. Shuchao Duan, Amirhossein Dadashzadeh, Alan L. Whone, Majid Mirmehdi |
FG | 4 |
| 2025 | Care-PD: A Multi-Site Anonymized Clinical Dataset for Parkinson's Disease Gait AssessmentabstractObjective gait assessment in Parkinson’s Disease (PD) is limited by the absence of large, diverse, and clinically annotated motion datasets. We introduce Care-PD, the largest publicly available archive of 3D mesh gait data for PD, and the first multi-site collection spanning 9 cohorts from 8 clinical centers. All recordings (RGB video or motion capture) are converted into anonymized SMPL meshes via a harmonized preprocessing pipeline. Care-PD supports two key benchmarks: supervised clinical score prediction (estimating Unified Parkinson’s Disease Rating Scale, UPDRS, gait scores) and unsupervised motion pretext tasks (2D-to-3D keypoint lifting and full-body 3D reconstruction). Clinical prediction is evaluated under four generalization protocols: within-dataset, cross-dataset, leave-one-dataset-out, and multi-dataset in-domain adaptation.To assess clinical relevance, we compare state-of-the-art motion encoders with a traditional gait-feature baseline, finding that encoders consistently outperform handcrafted features. Pretraining on Care-PD reduces MPJPE (from 60.8mm to 7.5mm) and boosts PD severity macro-F1 by 17\%, underscoring the value of clinically curated, diverse training data. Care-PD and all benchmark code are released for non-commercial research (Code, Data). Vida Adeli, Ivan Klabucar, Javad Rajabi, Benjamin Filtjens, Soroush Mehraban, Diwei Wang, Trung-Hieu Hoang, Minh N. Do, Hyewon Seo, Candice Müller, Daniel Boari Coelho, Claudia de Oliveira, Pieter Ginis, Moran Gilat, Alice Nieuwboer, Joke Spildooren, J. Lucas McKay, Hyeokhyen Kwon, Gari D. Clifford, Christine D. Esper, Stewart A. Factor, Imari Genias, Amirhossein Dadashzadeh, Leia C. Shum, Alan L. Whone, Majid Mirmehdi, Andrea Iaboni, Babak Taati |
NeurIPS | 26 |
| 2025 | Your turn: At home turning angle estimation for Parkinson's disease severity assessmentabstractPeople with Parkinson's Disease (PD) often experience progressively worsening gait, including changes in how they turn around, as the disease progresses. Existing clinical rating tools are not capable of capturing hour-by-hour variations of PD symptoms, as they are confined to brief assessments within clinic settings, leaving gait performance outside these controlled environments unaccounted for. Measuring turning angles continuously and passively is a component step towards using gait characteristics as sensitive indicators of disease progression in PD. This paper presents a deep learning-based approach to automatically quantify turning angles by extracting 3D skeletons from videos and calculating the rotation of hip and knee joints. We utilise advanced human pose estimation models, Fastpose and Strided Transformer, on a total of 1386 turning video clips from 24 subjects (12 people with PD and 12 healthy control volunteers), trimmed from a PD dataset of unscripted free-living videos in a home-like setting (Turn-REMAP). We also curate a turning video dataset, Turn-H3.6M, from the public Human3.6M human pose benchmark with 3D groundtruth, to further validate our method. Previous gait research has primarily taken place in clinics or laboratories evaluating scripted gait outcomes, but this work focuses on free-living home settings where complexities exist, such as baggy clothing and poor lighting. Due to difficulties in obtaining accurate groundtruth data in a free-living setting, we quantise the angle into the nearest bin 45° based on the manual labelling of expert clinicians. Our method achieves a turning calculation accuracy of 41.6%, a Mean Absolute Error (MAE) of 34.7°, and a weighted precision (WPrec) of 68.3% for Turn-REMAP. On Turn-H3.6M, it achieves an accuracy of 73.5%, an MAE of 18.5°, and a WPrec of 86.2%. This is the first work to explore the use of single monocular camera data to quantify turns by PD patients in a home setting. All data and models are publicly available, providing a baseline for turning parameter measurement to promote future PD gait research. Qiushuo Cheng, Catherine Morgan, Arindam Sikdar, Alessandro Masullo, Alan L. Whone, Majid Mirmehdi |
Artif. Intell. Medicine | 6 |
| 2025 | Deep dive into KABR: a dataset for understanding ungulate behavior from in-situ drone video
Maksim Kholiavchenko, Jenna Kline, Maksim Kukushkin, Otto Brookes, Samuel Stevens 0001, Isla Duporge, Alec Sheets, Reshma Ramesh Babu, Namrata Banerji, Elizabeth G. Campolongo, Matthew J. Thompson, Nina Van Tiel, Jackson Miliko, Eduardo Bessa, Majid Mirmehdi, Thomas Schmid 0003, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Tilo Burghardt, Charles V. Stewart |
Multim. Tools Appl. | 15 |
| 2024 | PECoP: Parameter Efficient Continual Pretraining for Action Quality AssessmentabstractThe limited availability of labelled data in Action Quality Assessment (AQA), has forced previous works to fine-tune their models pretrained on large-scale domain-general datasets. This common approach results in weak generalisation, particularly when there is a significant domain shift. We propose a novel, parameter efficient, continual pretraining framework, PECoP, to reduce such domain shift via an additional pretraining stage. In PECoP, we introduce 3D-Adapters, inserted into the pretrained model, to learn spatiotemporal, in-domain information via self-supervised learning where only the adapter modules’ parameters are updated. We demonstrate PECoP’s ability to enhance the performance of recent state-of-the-art methods (MUSDL, CoRe, and TSA) applied to AQA, leading to considerable improvements on benchmark datasets, JIGSAWS (↑ 6.0%), MTL-AQA (↑ 0.99%), and FineDiving (↑ 2.54%). We also present a new Parkinson’s Disease dataset, PD4T, of real patients performing four various actions, where we surpass (↑ 3.56%) the state-of-the-art in comparison. Our code, pretrained models, and the PD4T dataset are available at https://github.com/Plrbear/PECoP. Amirhossein Dadashzadeh, Shuchao Duan, Alan L. Whone, Majid Mirmehdi |
WACV | 4 |
| 2024 | PanAf20K: A Large Video Dataset for Wild Ape Detection and Behaviour RecognitionabstractAbstract We present the PanAf20K dataset, the largest and most diverse open-access annotated video dataset of great apes in their natural environment. It comprises more than 7 million frames across $$\sim $$ ∼ 20,000 camera trap videos of chimpanzees and gorillas collected at 18 field sites in tropical Africa as part of the Pan African Programme: The Cultured Chimpanzee. The footage is accompanied by a rich set of annotations and benchmarks making it suitable for training and testing a variety of challenging and ecologically important computer vision tasks including ape detection and behaviour recognition. Furthering AI analysis of camera trap information is critical given the International Union for Conservation of Nature now lists all species in the great ape family as either Endangered or Critically Endangered. We hope the dataset can form a solid basis for engagement of the AI community to improve performance, efficiency, and result interpretation in order to support assessments of great ape presence, abundance, distribution, and behaviour and thereby aid conservation efforts. The dataset and code are available from the project website: PanAf20K Otto Brookes, Majid Mirmehdi, Colleen Stephens, Samuel Angedakin, Katherine Corogenes, Dervla Dowd, Paula Dieguez, Thurston C. Hicks, Sorrel Jones, Vera Leinert, Juan Lapuente, Maureen S. McCarthy, Amelia Meier, Mizuki Murai, Emmanuelle Normand, Virginie Vergnes, Erin G. Wessling, Roman M. Wittig, Kevin Langergraber, Nuria Maldonado, Klaus Zuberbühler, Christophe Boesch, Mimi Arandjelovic, Hjalmar S. Kühl, Tilo Burghardt |
Int. J. Comput. Vis. | 2 |
| 2023 | Use Your Head: Improving Long-Tail Video RecognitionabstractThis paper presents an investigation into long-tail video recognition. We demonstrate that, unlike naturally-collected video datasets and existing long-tail image benchmarks, current video benchmarks fall short on multiple long-tailed properties. Most critically, they lack few-shot classes in their tails. In response, we propose new video benchmarks that better assess long-tail recognition, by sampling subsets from two datasets: SSv2 and VideoLT. We then propose a method, Long-Tail Mixed Reconstruction (LMR), which reduces overfitting to instances from few-shot classes by reconstructing them as weighted combinations of samples from head classes. LMR then employs label mixing to learn robust decision boundaries. It achieves state-of-the-art average class accuracy on EPIC-KITCHENS and the proposed SSv2-LT and VideoLT-LT. Benchmarks and code at: github.com/tobyperrett/lmr Toby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi, Dima Damen |
CVPR | 4 |
| 2023 | Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance SegmentationabstractThis paper presents a deep learning framework for medical video segmentation. Convolution neural network (CNN) and transformer-based methods have achieved great milestones in medical image segmentation tasks due to their incredible semantic feature encoding and global information comprehension abilities. However, most existing approaches ignore a salient aspect of medical video data - the temporal dimension. Our proposed framework explicitly extracts features from neighbouring frames across the temporal dimension and incorporates them with a temporal feature blender, which then tokenises the high-level spatio-temporal feature to form a strong global feature encoded via a Swin Transformer. The final segmentation results are produced via a UNet-like encoder-decoder architecture. Our model outperforms other approaches by a significant margin and improves the segmentation benchmarks on the VFSS2022 dataset, achieving a dice coefficient of 0.8986 and 0.8186 for the two datasets tested. Our studies also show the efficacy of the temporal feature blending scheme and cross-dataset transferability of learned capabilities. Code and models are available at https://github.com/SimonZeng7108/Video-SwinUNet. Chengxi Zeng, Xinyu Yang 0003, David Smithard, Majid Mirmehdi, Alberto M. Gambaruto, Tilo Burghardt |
ICIP | 4 |
| 2023 | Dynamic Curriculum Learning for Great Ape Detection in the WildabstractAbstract We propose a novel end-to-end curriculum learning approach for sparsely labelled animal datasets leveraging large volumes of unlabelled data to improve supervised species detectors. We exemplify the method in detail on the task of finding great apes in camera trap footage taken in challenging real-world jungle environments. In contrast to previous semi-supervised methods, our approach adjusts learning parameters dynamically over time and gradually improves detection quality by steering training towards virtuous self-reinforcement. To achieve this, we propose integrating pseudo-labelling with curriculum learning policies and show how learning collapse can be avoided. We discuss theoretical arguments, ablations, and significant performance improvements against various state-of-the-art systems when evaluating on the Extended PanAfrican Dataset holding approx. 1.8M frames. We also demonstrate our method can outperform supervised baselines with significant margins on sparse label versions of other animal datasets such as Bees and Snapshot Serengeti. We note that performance advantages are strongest for smaller labelled ratios common in ecological applications. Finally, we show that our approach achieves competitive benchmarks for generic object detection in MS-COCO and PASCAL-VOC indicating wider applicability of the dynamic learning concepts introduced. We publish all relevant source code, network weights, and data access details for full reproducibility. Tilo Burghardt, Majid Mirmehdi |
Int. J. Comput. Vis. | 3 |
| 2022 | Refining Action Boundaries for One-stage DetectionabstractCurrent one-stage action detection methods, which simultaneously predict action boundaries and the corresponding class, do not estimate or use a measure of confidence in their boundary predictions, which can lead to inaccurate boundaries. We incorporate the estimation of boundary confidence into one-stage anchor-free detection, through an additional prediction head that predicts the refined boundaries with higher confidence. We obtain state-of-the-art performance on the challenging EPICKITCHENS-100 action detection as well as the standard THUMOS14 action detection benchmarks, and achieve improvement on the ActivityNet-1.3 benchmark. Hanyuan Wang, Majid Mirmehdi, Dima Damen, Toby Perrett |
AVSS | 2 |
| 2022 | Video-TransUNet: temporally blended vision transformer for CT VFSS instance segmentationabstractWe propose Video-TransUNet, a deep architecture for instance segmentation in medical CT videos constructed by integrating temporal feature blending into the TransUNet deep learning framework. In particular, our approach amalgamates strong frame representation via a ResNet CNN backbone, multi-frame feature blending via a Temporal Context Module (TCM), non-local attention via a Vision Transformer, and reconstructive capabilities for multiple targets via a UNet-based convolutional-deconvolutional architecture with multiple heads. We show that this new network design can significantly outperform other state-of-the-art systems when tested on the segmentation of bolus and pharynx/larynx in Videofluoroscopic Swallowing Study (VFSS) CT sequences. On our VFSS2022 dataset it achieves a dice coefficient of 0.8796 and an average surface distance of 1.0379 pixels. Note that tracking the pharyngeal bolus accurately is a particularly important application in clinical practice since it constitutes the primary method for diagnostics of swallowing impairment. Our findings suggest that the proposed model can indeed enhance the TransUNet architecture via exploiting temporal information and improving segmentation performance by a significant margin. We publish key source code, network weights, and ground truth annotations for simplified performance reproduction. Chengxi Zeng, Xinyu Yang 0003, Majid Mirmehdi, Alberto M. Gambaruto, Tilo Burghardt |
ICMV | 3 |
| 2021 | Unsupervised View-Invariant Human Posture Representation
Faegheh Sardari, Björn Ommer, Majid Mirmehdi |
BMVC | 3 |
| 2021 | Back to the Future: Cycle Encoding Prediction for Self-supervised Video Representation Learning
Majid Mirmehdi, Tilo Burghardt |
BMVC | 2 |
| 2021 | Temporal-Relational CrossTransformers for Few-Shot Action RecognitionabstractWe propose a novel approach to few-shot action recognition, finding temporally-corresponding frame tuples between the query and videos in the support set. Distinct from previous few-shot works, we construct class prototypes using the CrossTransformer attention mechanism to observe relevant sub-sequences of all support videos, rather than using class averages or single best matches. Video representations are formed from ordered tuples of varying numbers of frames, which allows sub-sequences of actions at different speeds and temporal offsets to be compared.1Our proposed Temporal-Relational CrossTransformers (TRX) achieve state-of-the-art results on few-shot splits of Kinetics, Something-Something V2 (SSv2), HMDB51 and UCF101. Importantly, our method outperforms prior work on SSv2 by a wide margin (12%) due to the its ability to model temporal relations. A detailed ablation showcases the importance of matching to multiple support set videos and learning higher-order relational CrossTransformers. Toby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi, Dima Damen |
CVPR | 4 |
| 2021 | Exploring Motion Boundaries in an End-to-End Network for Vision-based Parkinson's Severity AssessmentabstractEvaluating neurological disorders such as Parkinson's disease (PD) is a challenging task that requires the assessment of several motor and non-motor functions. In this paper, we present an end-to-end deep learning framework to measure PD severity in two important components, hand movement and gait, of the Unified Parkinson's Disease Rating Scale (UPDRS). Our method leverages on an Inflated 3D CNN trained by a temporal segment framework to learn spatial and long temporal structure in video data. We also deploy a temporal attention mechanism to boost the performance of our model. Further, motion boundaries are explored as an extra input modality to assist in obfuscating the effects of camera motion for better movement assessment. We ablate the effects of different data modalities on the accuracy of the proposed network and compare with other popular architectures. We evaluate our proposed method on a dataset of 25 PD patients, obtaining 72.3% and 77.1% top-1 accuracy on hand movement and gait tasks respectively. Amirhossein Dadashzadeh, Alan L. Whone, Michal Rolinski, Majid Mirmehdi |
ICPRAM | 4 |
| 2020 | Meta-Learning with Context-Agnostic Initialisations
Toby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi, Dima Damen |
ACCV (4) | 4 |
| 2019 | HGR-Net: a fusion network for hand gesture segmentation and recognitionabstractWe propose a two‐stage convolutional neural network (CNN) architecture for robust recognition of hand gestures, called HGR‐Net, where the first stage performs accurate semantic segmentation to determine hand regions, and the second stage identifies the gesture. The segmentation stage architecture is based on the combination of fully convolutional residual network and atrous spatial pyramid pooling. Although the segmentation sub‐network is trained without depth information, it is particularly robust against challenges such as illumination variations and complex backgrounds. The recognition stage deploys a two‐stream CNN, which fuses the information from the red–green–blue and segmented images by combining their deep representations in a fully connected layer before classification. Extensive experiments on public datasets show that our architecture achieves almost as good as state‐of‐the‐art performance in segmentation and recognition of static hand gestures, at a fraction of training time, run time, and model size. Our method can operate at an average of 23 ms per frame. Amirhossein Dadashzadeh, Alireza Tavakoli Targhi, Maryam Tahmasbi, Majid Mirmehdi |
IET Comput. Vis. | 4 |
| 2018 | Action Completion: A Temporal Model for Moment Detection
Farnoosh Heidarivincheh, Majid Mirmehdi, Dima Damen |
BMVC | 2 |
| 2018 | CaloriNet: From silhouettes to calorie estimation in private environments
Alessandro Masullo, Tilo Burghardt, Dima Damen, Sion L. Hannuna, Víctor Ponce-López, Majid Mirmehdi |
BMVC | 6 |
| 2018 | Markerless Active Trunk Shape Modelling for Motion Tolerant Remote Respiratory AssessmentabstractWe present a vision-based trunk-motion tolerant approach which estimates lung volume-time data remotely in forced vital capacity (FVC) and slow vital capacity (SVC) spirometry tests. After temporal modelling of trunk shape, generated using two opposing Kinects in a sequence, the chest-surface respiratory pattern is computed by performing principal component analysis on temporal geometrical features extracted from the chest and posterior shapes. We evaluate our method on a publicly available dataset of 35 subjects (300 sequences) and compare against the state-of-the-art. By filtering complex trunk motions, our proposed method calibrates the entire volume-time data using only the tidal volume scaling factor which reduces the state-of-the-art average normalised L2 error from 0.136 to 0.05. Vahid Soleimani, Majid Mirmehdi, Dima Damen, James Dodd |
ICIP | 2 |
| 2018 | Energy expenditure estimation using visual and inertial sensorsabstractDeriving a person's energy expenditure accurately forms the foundation for tracking physical activity levels across many health and lifestyle monitoring tasks. In this study, the authors present a method for estimating calorific expenditure from combined visual and accelerometer sensors by way of an RGB‐Depth camera and a wearable inertial sensor. The proposed individual‐independent framework fuses information from both modalities which leads to improved estimates beyond the accuracy of single modality and manual metabolic equivalents of task (MET) lookup table based methods. For evaluation, the authors introduce a new dataset called SPHERE_RGBD + Inertial_calorie , for which visual and inertial data are simultaneously obtained with indirect calorimetry ground truth measurements based on gas exchange. Experiments show that the fusion of visual and inertial data reduces the estimation error by 8 and 18% compared with the use of visual only and inertial sensor only, respectively, and by 33% compared with a MET‐based approach. The authors conclude from their results that the proposed approach is suitable for home monitoring in a controlled environment. Lili Tao, Tilo Burghardt, Majid Mirmehdi, Dima Damen, Ashley Cooper, Massimo Camplani, Sion L. Hannuna, Adeline Paiement, Ian Craddock |
IET Comput. Vis. | 3 |
| 2017 | Detection of valuable left-behind items in vehicle cabinsabstractWe propose a method for detecting valuable left-behind items in vehicle cabins which uses a single overhead camera. An additional sub-network is incorporated into the Faster R-CNN framework in order to allow it to estimate item value based on visual properties, as well as to perform detection. A loss function which contains a user-specified minimum-value threshold is introduced, which enables warnings to be given if a detected item is above this threshold. As a significant amount of real data is time consuming to collect on the scale necessary for (deep) learning-based methods, an ImageNet model is first retrained on synthetic data to adapt it to our environment, before training on some real data. The effectiveness of this detection and validation approach is demonstrated by integrating additional valuation subnetworks into two convolutional neural network detection architectures. Toby Perrett, Majid Mirmehdi, Eduardo Dias 0002 |
Intelligent Vehicles Symposium | 2 |
| 2017 | Multiple human tracking in RGB-depth data: a surveyabstractMultiple human tracking (MHT) is a fundamental task in many computer vision applications. Appearance‐based approaches, primarily formulated on RGB data, are constrained and affected by problems arising from occlusions and/or illumination variations. In recent years, the arrival of cheap RGB‐depth devices has led to many new approaches to MHT, and many of these integrate colour and depth cues to improve each and every stage of the process. In this survey, the authors present the common processing pipeline of these methods and review their methodology based (a) on how they implement this pipeline and (b) on what role depth plays within each stage of it. They identify and introduce existing, publicly available, benchmark datasets and software resources that fuse colour and depth data for MHT. Finally, they present a brief comparative evaluation of the performance of those works that have applied their methods to these datasets. Massimo Camplani, Adeline Paiement, Majid Mirmehdi, Dima Damen, Sion L. Hannuna, Tilo Burghardt, Lili Tao |
IET Comput. Vis. | 3 |
| 2017 | Visual Monitoring of Driver and Passenger Control Panel InteractionsabstractAdvances in vehicular technology have resulted in more controls being incorporated into cabin designs. We present a system to determine which vehicle occupant is interacting with a control on the center console when it is activated, enabling the full use of dual-view touchscreens and the removal of duplicate controls. The proposed method relies on a background subtraction algorithm incorporating information from a superpixel segmentation stage. A manifold generated via the diffusion maps process handles the large variation in hand shapes, along with determining which part of the hand interacts with controls for a given gesture. We demonstrate superior results compared with other approaches on a challenging dataset. Toby Perrett, Majid Mirmehdi, Eduardo Dias 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | 3D Data Acquisition and Registration Using Two Opposing KinectsabstractWe present an automatic, open source data acquisition and calibration approach using two opposing RGBD sensors (Kinect V2) and demonstrate its efficacy for dynamic object reconstruction in the context of monitoring for remote lung function assessment. First, the relative pose of the two RGBD sensors is estimated through a calibration stage and rigid transformation parameters are computed. These are then used to align and register point clouds obtained from the sensors at frame level. We validated the proposed system by performing experiments on known-size box objects with the results demonstrating accurate measurements. We also report on dynamic object reconstruction by way of human subjects undergoing respiratory functional assessment. Vahid Soleimani, Majid Mirmehdi, Dima Damen, Sion L. Hannuna, Massimo Camplani |
3DV | 2 |
| 2016 | Beyond Action Recognition: Action Completion in RGB-D Data
Farnoosh Heidarivincheh, Majid Mirmehdi, Dima Damen |
BMVC | 2 |
| 2016 | A comparative study of pose representation and dynamics modelling for online motion quality assessmentabstractQuantitative assessment of the quality of motion is increasingly in demand by clinicians in healthcare and rehabilitation monitoring of patients. We study and compare the performances of different pose representations and HMM models of dynamics of movement for online quality assessment of human motion. In a general sense, our assessment framework builds a model of normal human motion from skeleton-based samples of healthy individuals. It encapsulates the dynamics of human body pose using robust manifold representation and a first-order Markovian assumption. We then assess deviations from it via a continuous online measure. We compare different feature representations, reduced dimensionality spaces, and HMM models on motions typically tested in clinical settings, such as gait on stairs and flat surfaces, and transitions between sitting and standing. Our dataset is manually labelled by a qualified physiotherapist. The continuous-state HMM, combined with pose representation based on body-joints’ location, outperforms standard discrete-state HMM approaches and other skeleton-based features in detecting gait abnormalities, as well as assessing deviations from the motion model on a frame-by-frame basis. Lili Tao, Adeline Paiement, Dima Damen, Majid Mirmehdi, Sion L. Hannuna, Massimo Camplani, Tilo Burghardt, Ian Craddock |
Comput. Vis. Image Underst. | 4 |
| 2016 | Registration and Modeling From Spaced and Misaligned Image VolumesabstractWe address the problem of object modeling from 3D and 3D+T data made up of images, which contain different parts of an object of interest, are separated by large spaces, and are misaligned with respect to each other. These images have only a limited number of intersections, hence making their registration particularly challenging. Furthermore, such data may result from various medical imaging modalities and can, therefore, present very diverse spatial configurations. Previous methods perform registration and object modeling (segmentation and interpolation) sequentially. However, sequential registration is ill-suited for the case of images with few intersections. We propose a new methodology, which, regardless of the spatial configuration of the data, performs the three stages of registration, segmentation, and shape interpolation from spaced and misaligned images simultaneously. We integrate these three processes in a level set framework, in order to benefit from their synergistic interactions. We also propose a new registration method that exploits segmentation information rather than pixel intensities, and that accounts for the global shape of the object of interest, for increased robustness and accuracy. The accuracy of registration is compared against traditional mutual information based methods, and the total modeling framework is assessed against traditional sequential processing and validated on artificial, CT, and MRI data. Adeline Paiement, Majid Mirmehdi, Xianghua Xie, Mark C. K. Hamilton |
IEEE Trans. Image Process. | 2 |
| 2015 | Real-time RGB-D Tracking with Depth Scaling Kernelised Correlation Filters and Occlusion HandlingabstractWe present a real-time RGB-D object tracker which manages occlusions and scale changes in a wide variety of scenarios. Its accuracy matches, and in many cases outperforms, state-of-the-art algorithms for precision and it far exceeds most in speed. We build our algorithm on the existing colour-only KCF tracker which uses the `kernel trick' to extend correlation filters for fast tracking. We fuse colour and depth cues as the tracker's features, and furthermore, exploit the depth data to both adjust a given target's scale, and detect and manage occlusions in such a way as to maintain real-time performance, exceeding on average 40~fps. We benchmark our approach using 2 publicly available datasets and make our easy-to-extend modularised code available to other researchers. Massimo Camplani, Sion L. Hannuna, Majid Mirmehdi, Dima Damen, Adeline Paiement, Lili Tao, Tilo Burghardt |
BMVC | 3 |
| 2015 | A comparative home activity monitoring study using visual and inertial sensorsabstractMonitoring actions at home can provide essential information for rehabilitation management. This paper presents a comparative study and a dataset for the fully automated, sample-accurate recognition of common home actions in the living room environment using commercial-grade, inexpensive inertial and visual sensors. We investigate the practical home-use of body-worn mobile phone inertial sensors together with an Asus Xmotion RGB-Depth camera to achieve monitoring of daily living scenarios. To test this setup against realistic data, we introduce the challenging SPHERE-H130 action dataset containing 130 sequences of 13 household actions recorded in a home environment. We report automatic recognition results at maximal temporal resolution, which indicate that a vision-based approach outperforms accelerometer provided by two phone-based inertial sensors by an average of 14.85% accuracy for home actions. Further, we report improved accuracy of a vision-based approach over accelerometry on particularly challenging actions as well as when generalising across subjects. Lili Tao, Tilo Burghardt, Sion L. Hannuna, Massimo Camplani, Adeline Paiement, Dima Damen, Majid Mirmehdi, Ian Craddock |
HealthCom | 7 |
| 2015 | Detection and Recognition of Painted Road Surface Markings
Jack Greenhalgh, Majid Mirmehdi |
ICPRAM (1) | 2 |
| 2015 | Correspondence, Matching and Recognition
Tilo Burghardt, Dima Damen, Walterio W. Mayol-Cuevas, Majid Mirmehdi |
Int. J. Comput. Vis. | 4 |
| 2015 | Recognizing Text-Based Traffic SignsabstractWe propose a novel system for the automatic detection and recognition of text in traffic signs. Scene structure is used to define search regions within the image, in which traffic sign candidates are then found. Maximally stable extremal regions (MSERs) and hue, saturation, and value color thresholding are used to locate a large number of candidates, which are then reduced by applying constraints based on temporal and structural information. A recognition stage interprets the text contained within detected candidate regions. Individual text characters are detected as MSERs and are grouped into lines, before being interpreted using optical character recognition (OCR). Recognition accuracy is vastly improved through the temporal fusion of text results across consecutive frames. The method is comparatively evaluated and achieves an overall $F_{\rm measure}$ of 0.87. Jack Greenhalgh, Majid Mirmehdi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Online quality assessment of human motion from skeleton data
Adeline Paiement, Lili Tao, Massimo Camplani, Sion L. Hannuna, Dima Damen, Majid Mirmehdi |
BMVC | 6 |
| 2014 | Text Line AggregationabstractWe present a new approach to text line aggregation that can work as both a line formation stage for a myriad of text segmentation methods (over all orientations) and as an extra level of filtering to remove false text candidates. The proposed method is centred on the processing of candidate text components based on local and global measures. We use orientation histograms to build an understanding of paragraphs, and filter noise and construct lines based on the discovery of prominent orientations. Paragraphs are then reduced to seed components and lines are reconstructed around these components. We demonstrate results for text aggregation on the ICDAR 2003 Robust Reading Competition data, and also present results on our own more complex data set. Copyright © 2014 SCITEPRESS. Chris Beck, Alan Broun, Majid Mirmehdi, Anthony G. Pipe, Chris Melhuish |
ICPRAM | 3 |
| 2014 | Segmentation of the Right Ventricle Using Diffusion Maps and Markov Random Fields
Oliver Moolan-Feroze, Majid Mirmehdi, Mark C. K. Hamilton, Chiara Bucciarelli-Ducci |
MICCAI (1) | 2 |
| 2014 | Weakly supervised learning of semantic colour termsabstractRecognition of visual attributes in images allows an image's information content to be expressed textually. This has benefits for web search and image archiving, especially since visual attributes transcend language barriers. Classifiers are traditionally trained using manually segmented images, which are expensive and time consuming to produce. The authors propose a method which uses raw, noisy and unsegmented results of web image searches, to learn semantic colour terms. They use probabilistic graphical models on continuous domain, both for weakly supervised learning, and for segmentation of novel images. Experiments show that the authors methods give better results than the current state of the art in colour naming using noisy, weakly labelled training data. ∗Note: Colour figures are available in the online version of this paper. David Hanwell, Majid Mirmehdi |
IET Comput. Vis. | 2 |
| 2014 | Real-time text tracking in natural scenesabstractThe authors present a system that automatically detects, recognises and tracks text in natural scenes in real‐time. The focus of the author's method is on large text found in outdoor environments, such as shop signs, street names, billboards and so on. Built on top of their previously developed techniques for scene text detection and orientation estimation, the main contribution of this work is to present a complete end‐to‐end scene text reading system based on text tracking. They propose to use a set of unscented Kalman filters to maintain each text region's identity and to continuously track the homography transformation of the text into a fronto‐parallel view, thereby being resilient to erratic camera motion and wide baseline changes in orientation. The system is designed for continuous, unsupervised operation in a handheld or wearable system over long periods of time. It is completely automatic and features quick failure recovery and interactive text reading. It is also highly parallelised to maximise usage of available processing power and achieve real‐time operation. They demonstrate the performance of the system on sequences recorded in outdoor scenarios. Carlos Merino-Gracia, Majid Mirmehdi |
IET Comput. Vis. | 2 |
| 2014 | QUAC: Quick unsupervised anisotropic clustering
David Hanwell, Majid Mirmehdi |
Pattern Recognit. | 2 |
| 2014 | Integrated Segmentation and Interpolation of Sparse DataabstractWe address the two inherently related problems of segmentation and interpolation of 3D and 4D sparse data and propose a new method to integrate these stages in a level set framework. The interpolation process uses segmentation information rather than pixel intensities for increased robustness and accuracy. The method supports any spatial configurations of sets of 2D slices having arbitrary positions and orientations. We achieve this by introducing a new level set scheme based on the interpolation of the level set function by radial basis functions. The proposed method is validated quantitatively and/or subjectively on artificial data and MRI and CT scans and is compared against the traditional sequential approach, which interpolates the images first, using a state-of-the-art image interpolation method, and then segments the interpolated volume in 3D or 4D. In our experiments, the proposed framework yielded similar segmentation results to the sequential approach but provided a more robust and accurate interpolation. In particular, the interpolation was more satisfactory in cases of large gaps, due to the method taking into account the global shape of the object, and it recovered better topologies at the extremities of the shapes where the objects disappear from the image slices. As a result, the complete integrated framework provided more satisfactory shape reconstructions than the sequential approach. Adeline Paiement, Majid Mirmehdi, Xianghua Xie, Mark C. K. Hamilton |
IEEE Trans. Image Process. | 2 |
| 2013 | Shape and appearance priors for level set-based left ventricle segmentationabstractThe authors propose a novel spatiotemporal constraint based on shape and appearance and combine it with a level‐set deformable model for left ventricle (LV) segmentation in four‐dimensional gated cardiac SPECT, particularly in the presence of perfusion defects. The model incorporates appearance and shape information into a ‘soft‐to‐hard’ probabilistic constraint, and utilises spatiotemporal regularisation via a maximum a posteriori framework. This constraint force allows more flexibility than the rigid forces of shape constraint‐only schemes, as well as other state of the art joint shape and appearance constraints. The combined model can hypothesise defective LV borders based on prior knowledge. The authors present comparative results to illustrate the improvement gain. A brief defect detection example is finally presented as an application of the proposed method. Ronghua Yang, Majid Mirmehdi, Xianghua Xie |
IET Comput. Vis. | 2 |
| 2013 | Fast perspective recovery of text in natural scenes
Carlos Merino-Gracia, Majid Mirmehdi, José F. Sigut, José L. González-Mora |
Image Vis. Comput. | 2 |
| 2012 | Detection of Lane Departure on High-speed Roads
David Hanwell, Majid Mirmehdi |
ICPRAM (2) | 2 |
| 2012 | Line Histogram - A Fast Method for Rotated Rectangular Area Histogramming
Sion L. Hannuna, Xianghua Xie, Majid Mirmehdi |
ICPRAM (2) | 4 |
| 2012 | A Diffusion Model for Detecting and Classifying Vesicle Fusion and Undocking Events
Lorenz Berger, Majid Mirmehdi, Sam Reed, Jeremy Tavare |
MICCAI (3) | 2 |
| 2012 | Automatic Bootstrapping and Tracking of Object ContoursabstractA new fully automatic object tracking and segmentation framework is proposed. The framework consists of a motion-based bootstrapping algorithm concurrent to a shape-based active contour. The shape-based active contour uses finite shape memory that is automatically and continuously built from both the bootstrap process and the active-contour object tracker. A scheme is proposed to ensure that the finite shape memory is continuously updated but forgets unnecessary information. Two new ways of automatically extracting shape information from image data given a region of interest are also proposed. Results demonstrate that the bootstrapping stage provides important motion and shape information to the object tracker. This information is found to be essential for good (fully automatic) initialization of the active contour. Further results also demonstrate convergence properties of the content of the finite shape memory and similar object tracking performance in comparison with an object tracker with unlimited shape memory. Tests with an active contour using a fixed-shape prior also demonstrate superior performance for the proposed bootstrapped finite-shape-memory framework and similar performance when compared with a recently proposed active contour that uses an alternative online learning model. John Chiverton, Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Image Process. | 3 |
| 2012 | Archive Film Defect Detection and Removal: An Automatic Restoration FrameworkabstractIn this paper, we present an automatic restoration system targeting on dirt and blotches in digitized archive films. The system is composed of mainly two modules: defect detection and defect removal. In defect detection, we locate the defects by combing temporal and spatial information across a number of frames. An HMM is trained for normal observation sequences and then applied within a framework to detect defective pixels. The resulting defect maps are refined in a two-stage false alarm elimination process and then passed over to the defect removal procedure. A labelled (degraded) pixels is restored in a multiscale framework by first searching the optimal replacement in its dynamically generated, random walk based region of candidate pixel-exemplars and then updating all its features (intensity, motion and texture). Finally, the proposed system is compared against state-of-the-art methods to demonstrate improved accuracy in both detection and restoration using synthetic and real degraded image sequences. Xiaosong Wang 0001, Majid Mirmehdi |
IEEE Trans. Image Process. | 2 |
| 2012 | Real-Time Detection and Recognition of Road Traffic SignsabstractThis paper proposes a novel system for the automatic detection and recognition of traffic signs. The proposed system detects candidate regions as maximally stable extremal regions (MSERs), which offers robustness to variations in lighting conditions. Recognition is based on a cascade of support vector machine (SVM) classifiers that were trained using histogram of oriented gradient (HOG) features. The training data are generated from synthetic template images that are freely available from an online database; thus, real footage road signs are not required as training data. The proposed system is accurate at high vehicle speeds, operates under a range of weather conditions, runs at an average speed of 20 frames per second, and recognizes all classes of ideogram-based (nontext) traffic symbols from an online road sign database. Comprehensive comparative results to illustrate the performance of the system are presented. Jack Greenhalgh, Majid Mirmehdi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2011 | Radial basis function based level set interpolation and evolution for deformable modelling
Xianghua Xie, Majid Mirmehdi |
Image Vis. Comput. | 2 |
| 2010 | Fast Dynamic Texture Detection
V. Javier Traver, Majid Mirmehdi, Xianghua Xie, Raúl Montoliu |
ECCV (4) | 2 |
| 2010 | Archive Film Restoration Based on Spatiotemporal Random Walks
Xiaosong Wang 0001, Majid Mirmehdi |
ECCV (5) | 2 |
| 2010 | Initialisation-Free Active Contour SegmentationabstractWe present a region based active contour model which does not require any initialisation and is capable of modelling multi-modal image regions. Its external force is based on statistically learning and grouping of image primitives in multiscale, and its numerical solution is carried out using radial basis function interpolation and time dependent expansion coefficient updating. The initialisation-free property makes it attractive to applications such as detecting unkown number of objects with unkown topologies. Xianghua Xie, Majid Mirmehdi |
ICPR | 2 |
| 2009 | On-line Learning of Shape Information for Object Segmentation and TrackingabstractWe present segmentation and tracking of deformable objects using non-linear on-line learning of high-level shape information in the form of a level set function. The emphasis is for successful tracking of objects that undergo smooth arbitrary deformations, but without the a priori learning of shape constraints. The high-level shape information is learnt on-line by defining a memory of object samples in a high-dimensional shape space. These shape samples are then used as weights via a locally defined shape space kernel function to define a template against which potential future shapes of the tracked object can be compared. Results for the successful tracking of a range of deformable motions are presented. 1 John Chiverton, Majid Mirmehdi, Xianghua Xie |
BMVC | 2 |
| 2009 | HMM based Archive Film Defect Detection with Spatial and Temporal ConstraintsabstractWe propose a novel probabilistic approach to detect defects in digitized archive film, by combining temporal and spatial information across a number of frames. An HMM is trained for normal observation sequences and then applied within a framework to detect defective pixels by examining each new observation sequence and its subformations via a leave-one-out process. A two-stage false alarm elimination process is then applied on the resulting defect maps, comprising MRF modelling and localised feature tracking, which impose spatial and temporal constraints respectively. The proposed method is compared against state-of-the-art and industry-standard methods to demonstrate its superior detection rate. Xiaosong Wang 0001, Majid Mirmehdi |
BMVC | 2 |
| 2008 | Tracking with Active Contours Using Dynamically Updated Shape InformationabstractAn active contour based tracking framework is described that generates and integrates dynamic shape information without having to learn a priori shape constraints. This dynamic shape information is combined with fixative pho-tometric foreground model matching and background mismatching. Bound-ary based optical flow is also used to estimate the location of the object in each new video frame, incorporating Procrustes based shape alignment. Promising results under complex deformations of shape, varied levels of noise, and close-to-complete occlusion in the presence of complex textured backgrounds are presented. 1 John Chiverton, Xianghua Xie, Majid Mirmehdi |
BMVC | 3 |
| 2008 | Variational Maximum A Posteriori model similarity and dissimilarity matchingabstractA new variational Maximum A Posteriori (MAP) contextual modeling approach is presented that minimizes the product of two ratios: (a) the ratio of the model distribution to the distribution of currently estimated foreground pixels; (b) the ratio of the background distribution to the model distribution for all estimated background pixels. This approach provides robust discrimination to identify the division between foreground and background pixels, which is useful for applications such as object tracking. John Chiverton, Majid Mirmehdi, Xianghua Xie |
ICPR | 2 |
| 2008 | MAC: Magnetostatic Active Contour ModelabstractWe propose an active contour model using an external force field that is based on magnetostatics and hypothesized magnetic interactions between the active contour and object boundaries. The major contribution of the method is that the interaction of its forces can greatly improve the active contour in capturing complex geometries and dealing with difficult initializations, weak edges and broken boundaries. The proposed method is shown to achieve significant improvements when compared against six well-known and state-of-the-art shape recovery methods, including the geodesic snake, the generalized version of GVF snake, the combined geodesic and GVF snake, and the charged particle model. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Correction to "MAC: Magnetostatic Active Contour Model"abstractIn the above titled paper (ibid., vol. 30, no. 4, pp. 632-646, Apr 08), there was an error in a definition. The correct definition is presented here. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Implicit Active Model using Radial Basis Function Interpolated Level SetsabstractBuilding on recent work by others that introduced RBFs into level sets for structural topology optimisation, we introduce the concept into active models and present a new level set formulation able to handle more complex topological changes, in particular perturbation away from the evolving front. This allows the initial contour or surface to be placed arbitrarily in the image. The proposed level set updating scheme is efficient and does not suffer from self-flattening while evolving which will cause large numerical error. Unlike conventional level set based active models, periodic re-initialisation is also no longer necessary and the computational grid can be much coarser, thus, it has great potential in modelling in high dimensional space. We show results on synthetic and real data for active modelling in 2D and 3D. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 2 |
| 2007 | TEXEMS: Texture Exemplars for Defect Detection on Random Textured SurfacesabstractWe present an approach to detecting and localizing defects in random color textures which requires only a few defect free samples for unsupervised training. It is assumed that each image is generated by a superposition of various-size image patches with added variations at each pixel position. These image patches and their corresponding variances are referred to here as textural exemplars or texems. Mixture models are applied to obtain the texems using multiscale analysis to reduce the computational costs. Novelty detection on color texture surfaces is performed by examining the same-source similarity based on the data likelihood in multiscale, followed by logical processes to combine the defect candidates to localize defects. The proposed method is compared against a Gabor filter bank-based novelty detection method. Also, we compare different texem generalization schemes for defect detection in terms of accuracy and efficiency. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | A Charged Active Contour Based on Electrostatics
Ronghua Yang, Majid Mirmehdi, Xianghua Xie |
ACIVS | 2 |
| 2006 | Magnetostatic Field for the Active Contour Model: A Study in ConvergenceabstractA new external velocity field for active contours is proposed. The velocity field is based on magnetostatics and hypothesised magnetic interactions be-tween the active contour and image gradients. In this paper, we introduce the method and study its convergence capability for the recovery of shapes with complex topology and geometry, including deep, narrow concavities. The proposed active contour can be arbitrarily initialised. Level sets are used to achieve topological freedom. The proposed method is compared against shape recovery methods based on distance vector flow, constant flow, gener-alised version of GVF, geodesic GGVF, and curvature vector flow. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 2 |
| 2006 | Colour tonality inspection using eigenspace features
Xianghua Xie, Majid Mirmehdi, Barry T. Thomas |
Mach. Vis. Appl. | 2 |
| 2005 | Recognizing Animals Using Motion PartsabstractWe describe a method for automatically recognizing animals in image sequences based on their distinctive locomotive movement patterns. The 2-D motion field associated with the animal is represented using a `configuration of motion parts' model, the characteristics of which are learned from training data. We adopt an unsupervised approach to learning model parameters, based on minimal a priori knowledge of the physical or locomotive characteristics of the animals concerned. Results are presented demonstrating excellent classification performance, with accuracy exceeding 98% on a test set consisting of over 100 sequences of 7 different species. Changming Kong, Andrew Calway, Majid Mirmehdi |
BMVC | 3 |
| 2005 | Localising surface defects in random colour textures using multiscale texem analysis in image eigenchannelsabstractA novel method is presented to detect defects in random colour textures which requires only a very few normal samples for unsupervised training. We decorrelate the colour image by generating three eigenchannels in each of which the surface texture image is divided into overlapping patches of various sizes. Then, a mixture model and EM is applied to reduce groupings of patches to a small number of textural exemplars, or texems. Localised defect detection is achieved by comparing the learned texems to patches in the unseen image eigenchannels. Xianghua Xie, Majid Mirmehdi |
ICIP (3) | 2 |
| 2005 | Special issue on camera-based text and document recognition
Majid Mirmehdi |
Int. J. Document Anal. Recognit. | 1 |
| 2005 | Computer vision elastography: speckle adaptive motion estimation for elastography using ultrasound sequencesabstractWe present the development and validation of an image based speckle tracking methodology, for determining temporal two-dimensional (2-D) axial and lateral displacement and strain fields from ultrasound video streams. We refine a multiple scale region matching approach incorporating novel solutions to known speckle tracking problems. Key contributions include automatic similarity measure selection to adapt to varying speckle density, quantifying trajectory fields, and spatiotemporal elastograms. Results are validated using tissue mimicking phantoms and in vitro data, before applying them to in vivo musculoskeletal ultrasound sequences. The method presented has the potential to improve clinical knowledge of tendon pathology from carpel tunnel syndrome, inflammation from implants, sport injuries, and many others. James Revell, Majid Mirmehdi, Donal McNally |
IEEE Trans. Medical Imaging | 2 |
| 2004 | Restructured Eigenfilter Matching for Novelty Detection in Random TexturesabstractA new eigenfilter-based novelty detection approach to find abnormalities in random textures is presented. The proposed algorithm reconstructs a given texture twice using a subset of its own eigenfilter bank and a subset of a reference (template) eigenfilter bank, and measures the reconstruction error as the level of novelty. We then present an improved reconstruction generated by structurally matched eigenfilters through rotation, negation, and mirroring. We apply the method to the detection of defects in textured ceramic tiles. The method is over $90\%$ accurate, and is fast and amenable to implementation on a production line. S. Amirhassan Monadjemi, Majid Mirmehdi, Barry T. Thomas |
BMVC | 2 |
| 2004 | RAGS: region-aided geometric snakeabstractAn enhanced, region-aided, geometric active contour that is more tolerant toward weak edges and noise in images is introduced. The proposed method integrates gradient flow forces with region constraints, composed of image region vector flow forces obtained through the diffusion of the region segmentation map. We refer to this as the Region-aided Geometric Snake or RAGS. The diffused region forces can be generated from any reliable region segmentation technique, greylevel or color. This extra region force gives the snake a global complementary view of the boundary information within the image which, along with the local gradient flow, helps detect fuzzy boundaries and overcome noisy regions. The partial differential equation (PDE) resulting from this integration of image gradient flow and diffused region flow is implemented using a level set approach. We present various examples and also evaluate and compare the performance of RAGS on weak boundaries and noisy images. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Image Process. | 2 |
| 2003 | Video Indexing using Motion EstimationabstractSummarising video data is essential to enable content-based video indexing and retrieval. A novel graph theoretic approach is presented to extract representative key frames corresponding to the shortest path of the graph for each shot. We distinguish further amongst paths of similar weight by examining the standard deviation of their constituent edge weights which improves the distribution of the selected key frames. The perceived camera motions contained within each shot are also annotated to introduce a further level of indexing and searching video content. 1 Sarah V. Porter, Majid Mirmehdi, Barry T. Thomas |
BMVC | 2 |
| 2003 | Strain Quantification In Ultrasound SequencesabstractA novel methodology to quantify displacements and strain in musculoskeletal ultrasound sequences is presented. We extend the principles of 2D interframe displacements produced by our earlier work using hierarchical variable block size matching, to quantify displacement \emph{trajectories}. We provide novel solutions for probe motion, quantification of objects moving in the 3D volume traversing the 2D plane, and improving the temporal coherence of displacements for longer image sequences than the frame pairs traditionally applied in ultrasound. We also present trajectory strain that yields a novel strain history for musculoskeletal tendon tissue samples. James Revell, Majid Mirmehdi, Donal McNally |
BMVC | 2 |
| 2003 | Geodesic Colour Active Contour Resistent to Weak Edges and NoiseabstractThe standard geometric or geodesic active contour is a powerful segmentation method, yet it is susceptible to weak edges and image noise. We propose a new region-aided, geometric, colour active contour that integrates gradient flow forces with region con-straints. These constraints are composed of image region vector flow forces obtained through the diffusion of the region segmentation map. The extra region force gives the snake a global view of the boundary information within the image which, along with the local gradient flow, helps detect fuzzy boundaries and overcome noisy re-gions. The partial differential equation (PDE) resulting from this integration of image gradient flow and diffused region flow is implemented using the level set approach. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 2 |
| 2003 | Text Selection by Structured Light Marking for Hand-held CamerasabstractWe describe a method to enable the selection of specific text regions with a hand-held camera by means of projecting a structured light pointer on the document. The user indicates the text required by dragging the laser pointer over it while a sequence of images is captured. By tracking the motion of the camera over the document and matching the trajectory to the actual text, the system is able to precisely determine the text portion the user intended to capture. By using this selection method along with a text-processing pipeline and OCR, a general purpose hand-held device (such as a PDA or mobile phone) with a camera could be used as effectively as single-purpose pen scanning devices. We present our results showing successful capture and extraction of text. Eve Bertucci, Maurizio Pilu, Majid Mirmehdi |
ICDAR | 3 |
| 2003 | Level-set based geometric colour snake with region supportabstractA novel method is introduced to force a geometric-based snake be more tolerant towards weak edges and noise in images. The method integrates gradient flow forces with region constraints obtained from diffused region segmentation forces. The diffusion is obtained from the region map vector flow field. This extra region force gives the snake a global view of the boundary information within the image. We present results on both graylevel and colour images. Xianghua Xie, Majid Mirmehdi |
ICIP (2) | 2 |
| 2003 | Temporal video segmentation and classification of edit effects
Sarah V. Porter, Majid Mirmehdi, Barry T. Thomas |
Image Vis. Comput. | 2 |
| 2003 | A non-contact method of capturing low-resolution text for OCR
Majid Mirmehdi, Paul Clark, J. Lam |
Pattern Anal. Appl. | 1 |
| 2003 | Rectifying perspective views of text in 3D scenes using vanishing points
Paul Clark, Majid Mirmehdi |
Pattern Recognit. | 2 |
| 2002 | Speed v. Accuracy for High Resolution Colour Texture ClassificationabstractMethods for extracting features and classifying textures in high resolution colour im-ages are presented. The proposed features are directional texture features obtained from the convolution of the Walsh-Hadamard transform with different orientations of texture patches from high resolution images, as well as simple chromatic features that correspond to hue and saturation in the HLS colour space. We compare the perfor-mance of these new features against Gabor transform features combined with HLS and Lab colour space features. Multiple classifiers are employed to combine both textural and chromatic features for better classification performance. We demonstrate a considerable reduction in computational costs, whilst maintaining close accuracy. 1 S. Amirhassan Monadjemi, Barry T. Thomas, Majid Mirmehdi |
BMVC | 3 |
| 2002 | Classification and Localisation of Diabetic-Related Eye Disease
Alireza Osareh, Majid Mirmehdi, Barry T. Thomas, Richard Markham |
ECCV (4) | 2 |
| 2002 | Comparative Exudate Classification Using Support Vector Machines and Neural Networks
Alireza Osareh, Majid Mirmehdi, Barry T. Thomas, Richard Markham |
MICCAI (2) | 2 |
| 2002 | Recognising text in real scenes
Paul Clark, Majid Mirmehdi |
Int. J. Document Anal. Recognit. | 2 |
| 2002 | Perceptual primitives from an extended 4D Hough transform
María J. Carreira, Majid Mirmehdi, Barry T. Thomas, Marta Penas |
Image Vis. Comput. | 2 |
| 2002 | British Machine Vision Conference 2000
Majid Mirmehdi |
Image Vis. Comput. | 1 |
| 2002 | Perceptual Image Indexing and Retrieval
Majid Mirmehdi, Radhakrishnan Periasamy |
J. Vis. Commun. Image Represent. | 1 |
| 2001 | Estimating the Orientation and Recovery of Text Planes in a Single ImageabstractA method for the fronto-parallel recovery of paragraphs of text under full perspective transformation is presented. The horizontal vanishing point of the text plane is found using an extension of 2D projection profiles. This allows the accurate segmentation of the lines of text. Analysis of the lines will then reveal the style of justification of the paragraph, and provide an estimate of the vertical vanishing point of the plane. The text is finally recovered to a fronto-parallel view suitable for OCR or other higher-level recognition. Paul Clark, Majid Mirmehdi |
BMVC | 2 |
| 2001 | CBIR with Perceptual Region FeaturesabstractA perceptual approach to generating features for use in indexing and retrieving images is described. Salient regions that immediately attract the eye are colour (textured) regions that usually dominate an image. Features derived from these will allow search for images that are similar perceptually. We compute colour features and Gabor colour texture features on regions identified from a coarse representation of the image, generated by a multi-band smoothing algorithm based on human psychophysical measurements of colour appearance. Images are retrieved, using a multi-feedback retrieval and ranking mechanism. We examine the performance of the features and the feedback mechanism. Majid Mirmehdi, Radhakrishnan Periasamy |
BMVC | 1 |
| 2001 | Detection and Classification of Shot TransitionsabstractThe process of shot break detection is a fundamental component in automatic video indexing, editing and archiving. This paper introduces a novel approach to the detection and classification of shot transitions in video sequences including cuts, fades and dissolves. It uses the average interframe correlation coefficient and block-based motion estimation to track image blocks through the video sequence and to distinguish changes caused by shot transitions from those caused by camera and object motion. We achieve better results compared with two established techniques. 1 Sarah V. Porter, Majid Mirmehdi, Barry T. Thomas |
BMVC | 2 |
| 2000 | Finding Text Regions using Localised Statistical Measures
Paul Clark, Majid Mirmehdi |
BMVC | 2 |
| 2000 | Grouping of Directional Features Using an Extended Hough TransformabstractDirectional features extracted from Gabor responses are used as primitives for perceptual grouping. In previous work, we extracted Gabor features in 8 directions and then applied two self-organising maps, thus classifying each pixel in the image within a neuron-map, each corner of which represents one of four main directions. In this work we group pixels with similar directional features to detect salient structures within an image. Results obtained from application to forward-looking infrared (FLIR) images are very promising. María J. Carreira, Majid Mirmehdi, Barry T. Thomas, John F. Haddon |
ICPR | 2 |
| 2000 | Combining Statistical Measures to Find Image Text RegionsabstractWe present a method based on statistical properties of local image pixels for focussing attention on regions of text in arbitrary scenes where the text plane is not necessarily fronto-parallel to the camera. This is particularly useful for desktop or wearable computing applications. The statistical measures are chosen to reveal characteristic properties of text. We combine a number of localised measures using a neural network to classify each pixel as text or non-text. We demonstrate our results on typical images. Paul Clark, Majid Mirmehdi |
ICPR | 2 |
| 2000 | Video Cut Detection using Frequency Domain CorrelationabstractA common video indexing technique is to segment a video sequence into shots and then select representative key-frames. The process of shot break detection is a fundamental component in automatic video indexing, editing, and archiving. This paper introduces a novel video cut detection technique which performs in the frequency domain. It is computationally tractable and robust with respect to sudden changes in mean intensity within a shot. The method uses the average interframe correlation coefficients to determine whether an abrupt shot change has occurred. We compare our method against three established techniques and present our results using different video sequences. Sarah V. Porter, Majid Mirmehdi, Barry T. Thomas |
ICPR | 2 |
| 2000 | Segmentation of Color TexturesabstractThis paper describes an approach to perceptual segmentation of color image textures. A multiscale representation of the texture image, generated by a multiband smoothing algorithm based on human psychophysical measurements of color appearance is used as the input. Initial segmentation is achieved by applying a clustering algorithm to the image at the coarsest level of smoothing. The segmented clusters are then restructured in order to isolate core clusters, i.e., patches in which the pixels are definitely associated with the same region. The image pixels representing the core clusters are used to form 3D color histograms which are then used for probabilistic assignment of all other pixels to the core clusters to form larger clusters and categorise the rest of the image. The process of setting up color histograms and probabilistic reassignment of the pixels to the clusters is then propagated through finer levels of smoothing until a full segmentation is achieved at the highest level of resolution. Majid Mirmehdi, Maria Petrou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Road Recognition Using Fuzzy ClassifiersabstractCurrent learning approaches to computer vision have mainly focussed on low-level image processing and object recognition, while tending to ignore higher level processing for understanding. We propose an approach to scene analysis that facilitates the transition from recognition to understanding. It begins by segmenting the image into regions using standard approaches, which are then classified using a discovered fuzzy Cartesian granule feature classifier. Understanding is made possible through the transparent and succinct nature of the discovered models. The recognition of roads in images is taken as an illustrative problem. The discovered fuzzy models while providing high levels of accuracy (97%), also provide understanding of the problem domain through the transparency of the learnt models. The learning step in the proposed approach is compared with other techniques such as decision trees, naive Bayes and neural networks using a variety of performance criteria such as accuracy, understandability, and efficiency. James G. Shanahan, Barry T. Thomas, Majid Mirmehdi, Trevor P. Martin, Jim F. Baldwin |
BMVC | 3 |
| 1999 | A Prototype Hotel Browsing System Using Java3DabstractJava3D is an application-centred approach to building 3D worlds. We use Java3D and VRML to design a prototype WWW-based 3D Hotel Browsing system. A Java3D scene graph viewer was implemented to interactively explore objects in a virtual universe using models generated by a commercial computer graphics suite and imported using a VRML file loader. A special collision prevention mechanism is also devised. This case study is reported on by reviewing the current aspects of the prototype system. D. Ball, Majid Mirmehdi |
IV | 2 |
| 1999 | Feedback control strategies for object recognitionabstractWe present a paradigm for feedback strategies that find instances of a generic class of objects by improving on established single-pass hypothesis generation and verification approaches. We improve upon the mechanisms of the traditional or classical image processing systems by introducing control strategies at low, intermediate, and high levels of analysis. We produce optimal sets of low-level features to reduce the number of hypotheses generated. The feedback further enables updated sets of features to be extracted so that the target object may be located even in very, noisy data. The use of an interest operator in the feedback directs the search through the hypotheses in an optimal manner, so minimizing the amount of feedback to false alarms. Furthermore, we aim to obtain detailed information about a complex object and not just its location. Thus, following top-down recognition of the object our feedback control directs the search for missing information. The system can extract complex objects in a scale and rotation independent manner where the objects may be partially occluded. The method is illustrated using box shaped objects and noisy IR images of a number of bridges. Majid Mirmehdi, Phil L. Palmer, Josef Kittler, Homam Dabis |
IEEE Trans. Image Process. | 1 |
| 1998 | Optimising the Complete Image Feature Extraction Chain
Majid Mirmehdi, Phil L. Palmer, Josef Kittler |
ACCV (2) | 1 |
| 1998 | Detection and Tracking of Very Small Low Contrast ObjectsabstractWe present a Kalman tracking algorithm that can track a number of very small, low contrast objects through an image sequence taken from a static camera. The issues that we have addressed to achieve this are twofold. Firstly, the detection of small objects comprising a few pixels only, mov-ing slowly in the image, and secondly, tracking of multiple small targets even though they may be lost either through occlusion or in noisy signal. The ap-proach uses a combination of wavelet filtering for detection with an interest operator for testing multiple target hypotheses based within the framework of a Kalman tracker. We demonstrate the robustness of the approach to oc-clusion and for multiple targets. 1 D. Davies, Phil L. Palmer, Majid Mirmehdi |
BMVC | 3 |
| 1998 | Perceptual Smoothing and Segmentation of Colour Textures
Maria Petrou, Majid Mirmehdi, M. Coors |
ECCV (1) | 2 |
| 1997 | Multi-Level Probabilistic Relaxation
Maria Petrou, Majid Mirmehdi, M. Coors |
BMVC | 2 |
| 1997 | Genetic optimisation of the image feature extraction process
Majid Mirmehdi, Phil L. Palmer, Josef Kittler |
Pattern Recognit. Lett. | 1 |
| 1996 | Complex Feedback Strategies for Hypothesis Generation and VerificationabstractWe extend the idea of the single-pass feedback framework by employing complex feedback strategies for both more robust hypothesis generation and hypothesis verification. These strategies are developed at every level of our object recognition application; from low-level parameter optimisation, through the low level processing chain, to higher level recognition stages. The strategies are independent of the techniques used at each level. We introduce various control mechanisms to achieve such complex feedback strategies. Within our implementation, we minimise the amount of feedback to false alarms by using an interest operator which directs the search through the hypotheses in an optimal manner. Furthermore, we obtain detailed information about a complex object and not just its location. Thus, following top-down recognition of the object our feedback control directs the search for missing information. We illustrate our approach using noisy infra-red images of bridges in real outdoor scene... Majid Mirmehdi, Phil L. Palmer, Josef Kittler, Homam Dabis |
BMVC | 1 |
| 1993 | Parallel approach to tracking edge segments in dynamic scenes
Majid Mirmehdi, T. J. Ellis |
Image Vis. Comput. | 1 |
| 1991 | Product label inspection using transputersabstractAbstract A product label inspection system implemented on a transputer network is described. Two data routing algorithms are considered, one for a general M × N array and one for a customized setup of transputers. Data transfer results for various image sizes are provided. Reliable and fast label inspection algorithms are described for rectangular, acute‐angled and oval‐shaped product labels, and their implementation on both the general and the customized networks are discussed. Simulated test results are provided to illustrate the present capability of the system for inspecting large numbers of labels per second. Majid Mirmehdi |
Concurr. Pract. Exp. | 1 |