EDBT 2026 Demo / reviewers in the wild / expert
Hossein Rahmani 0001
dblp:40/7475
· DBLP profile ↗
66ranked-venue papers
11as first author
50since 2021 · last 2026
0000-0003-1920-0371ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 8 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 30 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Visual Privacy Protection Meets Multimodal Large Language ModelsabstractAbstract The emergence of Multimodal Large Language Models (MLLMs) and the widespread usage of MLLM cloud services such as GPT-4V raised great concerns about privacy leakage in visual data. As these models are typically deployed in cloud services, users are required to submit their images and videos, posing serious privacy risks. However, how to tackle such privacy concerns is an under-explored problem. Thus, in this paper, we aim to conduct a new investigation to protect visual privacy when enjoying the convenience brought by MLLM services. We address the practical case where the MLLM is a “black box”, i.e., we only have access to its input and output without knowing its internal model information. To tackle such a challenging yet demanding problem, we propose a novel framework, in which we carefully design the learning objective with Pareto optimality to seek a better trade-off between visual privacy and MLLM’s performance, and propose critical-history enhanced optimization to effectively optimize the framework with the black-box MLLM. Our experiments show that our method is effective on different benchmarks. Xiaofei Hui, Haoxuan Qu, Majid Mirmehdi, Hossein Rahmani 0001, Jun Liu 0036 |
Int. J. Comput. Vis. | 5 |
| 2026 | Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional SportsabstractAbstract Reasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval. However, this task has not been explored due to the lack of relevant datasets and the challenging nature it presents. Most datasets for video question answering (VideoQA) focus mainly on general and coarse-grained understanding of daily-life videos, which is not applicable to sports scenarios requiring professional action understanding and fine-grained motion analysis. In this paper, we introduce the first dataset, named Sports-QA, specifically designed for the sports VideoQA task. The Sports-QA dataset includes various types of questions, such as descriptions, chronologies, causalities, and counterfactual conditions, covering multiple sports. Furthermore, to address the characteristics of the sports VideoQA task, we propose a new Auto-Focus Transformer (AFT) capable of automatically focusing on particular scales of temporal information for question answering. We conduct extensive experiments on Sports-QA, including baseline studies and the evaluation of different methods. The results demonstrate that our AFT achieves state-of-the-art performance. Haopeng Li 0001, Andong Deng, Jun Liu 0036, Hossein Rahmani 0001, Yulan Guo, Bernt Schiele, Mohammed Bennamoun, Qiuhong Ke |
Int. J. Comput. Vis. | 4 |
| 2026 | Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Hossein Rahmani 0001, Nicu Sebe, Ajmal Mian |
Int. J. Comput. Vis. | 8 |
| 2026 | DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action RecognitionabstractZero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantics. This coarse-grained alignment fails to bridge the domain shift between seen, unseen classes, thereby impeding the effective transfer of fine-grained visual knowledge. To address these limitations, we introduce DynaPURLS, a unified framework that establishes robust, multi-scale visual-semantic correspondences, dynamically refines them at inference time to enhance generalization. Our framework leverages a large language model to generate hierarchical textual descriptions that encompass both global movements, local body-part dynamics. Concurrently, an adaptive partitioning module produces fine-grained visual representations by semantically grouping skeleton joints. To fortify this fine-grained alignment against the train-test domain shift, DynaPURLS incorporates a dynamic refinement module. During inference, this module adapts textual features to the incoming visual stream via a lightweight learnable projection. This refinement process is stabilized by a confidence-aware, class-balanced memory bank, which mitigates error propagation from noisy pseudo-labels. Extensive experiments on three large-scale benchmark datasets, including NTU RGB+D 60/120, PKU-MMD, demonstrate that DynaPURLS significantly outperforms prior art, setting new state-of-the-art records. Jingmin Zhu, James Bailey 0001, Jun Liu 0036, Hossein Rahmani 0001, Mohammed Bennamoun, Farid Boussaïd, Qiuhong Ke |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Prompt-guided selective frequency network for real-world scene text image super-Resolutionabstract• We introduce PGSFNet for real-world scene text image super-resolution. • Adaptive Frequency Modulator is proposed to extract informative frequency components. • Text Information Enhancement module is designed to incorporate text priors. • We develop a Sobel loss to guide optimization towards sharper text details. • PGSFNet is shown to achieve superior performance on public text image datasets. Real-world scene text image super-resolution is challenging due to complex writing strokes, random text distribution, and diverse scene degradations. Existing text super-resolution methods focus on pure text images or fixed-size single-line text, which limits their practical utility. To address that, we propose a Prompt-Guided Selective Frequency super-resolution Network (PGSFNet). Our unique bicephalous neural model comprises a super-resolution branch and a prompt guidance branch. The latter specifically helps in leveraging text content-aware information priors. To that end, we propose a Text Information Enhancement module. To exploit selective frequency information present in the image, PGSFNet employs a proposed Adaptive Frequency Modulator fused with multi-attention structures. Considering the criticality of text edges in our task, we also propose a tailored text edge perception loss. Extensive experiments on the standard open real-world scene text image datasets demonstrate remarkable performance of our method, achieving up to 8.75% PNSR gain for × 2 and 2.28% SSIM gain for × 4 super-resolution on the Real-CE dataset. Our code will be made public at https://github.com/holastq/PGSFNet . Tianqi Shan, Hanlin Qin, Naveed Akhtar, Hossein Rahmani 0001, Ajmal Mian |
Pattern Recognit. | 6 |
| 2026 | Scalable Unseen Objects 6-DoF Absolute Pose Estimation With Robotic Integration
Jian Liu 0014, Wei Sun 0028, Hossein Rahmani 0001, Ajmal Mian, Lin Wang 0025 |
IEEE Trans. Robotics | 6 |
| 2025 | An Image-like Diffusion Method for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection often faces high levels of ambiguity and indeterminacy, as the same interaction can appear vastly different across different human-object pairs. Additionally, the indeterminacy can be further exacerbated by issues such as occlusions and cluttered backgrounds. To handle such a challenging task, in this work, we begin with a key observation: the output of HOI detection for each human-object pair can be recast as an image. Thus, inspired by the strong image generation capabilities of image diffusion models, we propose a new framework, HOI-IDiff. In HOI-IDiff, we tackle HOI detection from a novel perspective, using an Image-like Diffusion process to generate HOI detection outputs as images. Furthermore, recognizing that our recast images differ in certain properties from natural images, we enhance our framework with a customized HOI diffusion process and a slice patchification model architecture, which are specifically tailored to generate our recast “HOI images”. Extensive experiments demonstrate the efficacy of our framework. Xiaofei Hui, Haoxuan Qu, Hossein Rahmani 0001, Jun Liu 0036 |
CVPR | 3 |
| 2025 | LongDiff: Training-Free Long Video Generation in One GoabstractVideo diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components – Position Mapping (PM) and Informative Frame Selection (IFS) – to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method. Zhuoling Li, Hossein Rahmani 0001, Qiuhong Ke, Jun Liu 0036 |
CVPR | 2 |
| 2025 | Boundary Probing for Input Privacy Protection when Using LMM Services
Xiaofei Hui, Haoxuan Qu, Ping Hu 0001, Hossein Rahmani 0001, Jun Liu 0036 |
ICCV | 4 |
| 2025 | DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion Models
Zhuoling Li, Haoxuan Qu, Jason Kuen, Jiuxiang Gu, Qiuhong Ke, Jun Liu 0036, Hossein Rahmani 0001 |
ICCV | 7 |
| 2025 | Performing Defocus Deblurring by Modeling its Formation Process
Zhengbo Zhang, Lin Geng Foo, Hossein Rahmani 0001, Jun Liu 0036, De Wen Soh |
ICCV | 3 |
| 2025 | GaussianBlock: Building Part-Aware Compositional and Editable 3D Scene by Primitives and GaussiansabstractRecently, with the development of Neural Radiance Fields and Gaussian Splatting, 3D reconstruction techniques have achieved remarkably high fidelity. However, the latent representations learnt by these methods are highly entangled and lack interpretability. In this paper, we propose a novel part-aware compositional reconstruction method, called GaussianBlock, that enables semantically coherent and disentangled representations, allowing for precise and physical editing akin to building blocks, while simultaneously maintaining high fidelity.
Our GaussianBlock introduces a hybrid representation that leverages the advantages of both primitives, known for their flexible actionability and editability, and 3D Gaussians, which excel in reconstruction quality. Specifically, we achieve semantically coherent primitives through a novel attention-guided centering loss derived from 2D semantic priors, complemented by a dynamic splitting and fusion strategy.
Furthermore, we utilize 3D Gaussians that hybridize with primitives to refine structural details and enhance fidelity.
Additionally, a binding inheritance strategy is employed to strengthen and maintain the connection between the two.
Our reconstructed scenes are evidenced to be disentangled, compositional, and compact across diverse benchmarks, enabling seamless, direct and precise editing while maintaining high quality. Shuyi Jiang, Qihao Zhao, Hossein Rahmani 0001, De Wen Soh, Jun Liu 0036, Na Zhao 0004 |
ICLR | 3 |
| 2025 | TSTMotion: Training-free Scene-aware Text-to-motion GenerationabstractText-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted exploration into scene-aware text-to-motion generation methods. Yet, existing scene-aware methods often rely on large-scale ground-truth motion sequences in diverse 3D scenes, which poses practical challenges due to the expensive cost. To mitigate this challenge, we are the first to propose a Training-free Scene-aware Text-to-Motion framework, dubbed as TSTMotion, that efficiently empowers pre-trained blank-background motion generators with the scene-aware capability. Specifically, conditioned on the given 3D scene and text description, we adopt foundation models together to reason, predict and validate a scene-aware motion guidance. Then, the motion guidance is incorporated into the blank-background motion generators with two modifications, resulting in scene-aware text-driven motion sequences. Extensive experiments demonstrate the efficacy and generalizability of our proposed framework. We release our code in Project Page. Ziyan Guo, Haoxuan Qu, Hossein Rahmani 0001, De Wen Soh, Ping Hu 0001, Qiuhong Ke, Jun Liu 0036 |
ICME | 3 |
| 2025 | MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion ModelabstractObject pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects. We propose a diffusion-based monocular category-level 9D object pose generation method, MonoDiff9D. Our motivation is to leverage the probabilistic nature of diffusion models to alleviate the need for shape priors, CAD models, or depth sensors for intra-class unknown object pose estimation. We first estimate coarse depth via DINOv2 from the monocular image in a zero-shot manner and convert it into a point cloud. We then fuse the global features of the point cloud with the input image and use the fused features along with the encoded time step to condition MonoDiff9D. Finally, we design a transformer-based denoiser to recover the object pose from Gaussian noise. Extensive experiments on two popular benchmark datasets show that MonoDiff9D achieves state-of-the-art monocular category-level 9D object pose estimation accuracy without the need for shape priors or CAD models at any stage. Our code will be made public at https://github.com/CNJianLiu/MonoDiff9D. Jian Liu 0014, Wei Sun 0028, Zichen Geng, Hossein Rahmani 0001, Ajmal Mian |
ICRA | 6 |
| 2025 | Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time AdaptationabstractWe introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-parametric cache that stores structured skeleton representations, combining both global and fine-grained local descriptors. To guide the fusion of descriptor-wise predictions, we leverage the semantic reasoning capabilities of large language models (LLMs) to assign class-specific importance weights. By integrating these structured descriptors with LLM-guided semantic priors, Skeleton-Cache dynamically adapts to unseen actions without any additional training or access to training data. Extensive experiments on NTU RGB+D 60/120 and PKU-MMD II demonstrate that Skeleton-Cache consistently boosts the performance of various SZAR backbones under both zero-shot and generalized zero-shot settings. The code is publicly available at https://github.com/Alchemist0754/Skeleton-Cache. Jingmin Zhu, Hossein Rahmani 0001, Jun Liu 0036, Mohammed Bennamoun, Qiuhong Ke |
NeurIPS | 3 |
| 2025 | Recent Advances of Continual Learning in Computer Vision: An OverviewabstractABSTRACT In contrast to batch learning where all training data is available at once, continual learning represents a family of methods that accumulate knowledge and learn continuously with data available in sequential order. Similar to the human learning process with the ability of learning, fusing and accumulating new knowledge acquired at different time steps, continual learning is considered to have high practical significance. Hence, continual learning has been studied in various artificial intelligence tasks. In this paper, we present a comprehensive review of the recent progress of continual learning in computer vision. In particular, the works are grouped by their representative techniques, including regularisation, knowledge distillation, memory, generative replay, parameter isolation and a combination of the above techniques. For each category of these techniques, both its characteristics and applications in computer vision are presented. At the end of this overview, several subareas, where continuous knowledge accumulation is potentially helpful while continual learning has not been well studied, are discussed. Haoxuan Qu, Hossein Rahmani 0001, Bryan M. Williams 0001, Jun Liu 0036 |
IET Comput. Vis. | 2 |
| 2025 | Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose EstimationabstractNine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization to intra-class unknown objects. However, these methods require manual collection and labeling of large-scale real-world training data. To address this problem, we introduce a diffusion-based paradigm for domain-generalized category-level 9-DoF object pose estimation. Our motivation is to leverage the latent generalization ability of the diffusion model to address the domain generalization challenge in object pose estimation. This entails training the model exclusively on rendered synthetic data to achieve generalization to real-world scenes. We propose an effective diffusion model to redefine 9-DoF object pose estimation from a generative perspective. Our model does not require any 3D shape priors during training or inference. By employing the Denoising Diffusion Implicit Model, we demonstrate that the reverse diffusion process can be executed in as few as 3 steps, achieving near real-time performance. Finally, we design a robotic grasping system comprising both hardware and software components. Through comprehensive experiments on two benchmark datasets and the real-world robotic system, we show that our method achieves state-of-the-art domain generalization performance. Jian Liu 0014, Wei Sun 0028, Pengchao Deng, Chongpei Liu, Nicu Sebe, Hossein Rahmani 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | 3D Points Splatting for real-time dynamic Hand ReconstructionabstractWe present 3D Points Splatting Hand Reconstruction (3D-PSHR), a real-time and photo-realistic hand reconstruction approach. We propose a self-adaptive canonical points upsampling strategy to achieve high-resolution hand geometry representation. This is followed by a self-adaptive deformation that deforms the hand from the canonical space to the target pose, adapting to the dynamic changing of canonical points which, in contrast to the common practice of subdividing the MANO model, offers greater flexibility and results in improved geometry fitting. To model texture, we disentangle the appearance color into the intrinsic albedo and pose-aware shading, which are learned through a Context-Attention module. Moreover, our approach allows the geometric and the appearance models to be trained simultaneously in an end-to-end manner. We demonstrate that our method is capable of producing animatable, photorealistic and relightable hand reconstructions using multiple datasets, including monocular videos captured with handheld smartphones and large-scale multi-view videos featuring various hand poses. We also demonstrate that our approach achieves real-time rendering speeds while simultaneously maintaining superior performance compared to existing state-of-the-art methods. • We propose 3D-PSHR, a real-time, photo-realistic hand reconstruction via point clouds. • Our method creates animatable, photorealistic, relightable hands from various datasets. • Our approach shows real-time rendering with superior performance over state-of-the-art. Zheheng Jiang, Hossein Rahmani 0001, Sue Black 0002, Bryan M. Williams 0001 |
Pattern Recognit. | 2 |
| 2025 | Unpaired 3D Shape-to-Shape Translation via Gradient-Guided Triplane DiffusionabstractUnpaired shape-to-shape translation refers to the task of transforming the geometry and semantics of an input shape into a new shape domain without paired training data. Previous methods utilize GAN-based architectures to perform shape translation, employing adversarial training to transform the source shape encoding into the target domain in the low-dimensional latent feature space. However, these methods encounter difficulties in generating diverse and high-quality results, as they often suffer from issues such as "mode collapse". This leads to limited generation diversity and makes it challenging to find an accurate latent code that adequately represents the input shape. In this article, we achieve unpaired shape-to-shape translation via a triplane diffusion model, in which we factorize 3D objects into triplane representations and conduct a diffusion process on these representations to accomplish shape domain transformation. We observe that by adding an appropriate amount of noise to an input object during the forward diffusion process, domain-specific shape structures are smoothed out while the overall structure is still preserved. Subsequently, we progressively remove the noise via an unconditional diffusion model trained on the target shape domain in the reverse diffusion process. This allows us to obtain a denoised output that retains the structural similarities of the source input while aligning with the distribution of the target shape domain. During this process, we propose two gradient-based guidance mechanisms to guide the translation process to guarantee more faithful results during the denoising process. We conduct extensive experiments on different shape domains, and the experimental results demonstrate that our method achieves superior shape fidelity with high quality compared to current state-of-the-art baselines. Hossein Rahmani 0001, Jun Liu 0036 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Action Detection via an Image Diffusion ProcessabstractAction detection aims to localize the starting and ending points of action instances in untrimmed videos, and predict the classes of those instances. In this paper, we make the observation that the outputs of the action detection task can be formulated as images. Thus, from a novel perspective, we tackle action detection via a three-image generation process to generate starting point, ending point and action-class predictions as images via our proposed Action Detection Image Diffusion (ADI-Diff) framework. Furthermore, since our images differ from natural images and exhibit special properties, we further explore a Discrete Action-Detection Diffusion Process and a Row-Column Transformer design to better handle their processing. Our ADI-Diff framework achieves state-of-the-art results on two widely-used datasets. Lin Geng Foo, Hossein Rahmani 0001, Jun Liu 0036 |
CVPR | 3 |
| 2024 | LLMs are Good Sign Language TranslatorsabstractSign Language Translation (SLT) is a challenging task that aims to translate sign videos into spoken language. Inspired by the strong translation capabilities of large language models (LLMs) that are trained on extensive multilingual text corpora, we aim to harness off-the-shelf LLMs to handle SLT. In this paper, we regularize the sign videos to embody linguistic characteristics of spoken language, and propose a novel SignLLM framework to transform sign videos into a language-like representation for improved readability by off-the-shelf LLMs. SignLLM comprises two key modules: (1) The Vector-Quantized Visual Sign module converts sign videos into a sequence of discrete character-level sign tokens, and (2) the Codebook Reconstruction and Alignment module converts these character-level tokens into word-level sign representations using an optimal transport formulation. A sign-text alignment loss further bridges the gap between sign and text tokens, enhancing semantic compatibility. We achieve state-of-the-art gloss-free results on two widely-used SLT benchmarks. Jia Gong, Lin Geng Foo, Hossein Rahmani 0001, Jun Liu 0036 |
CVPR | 4 |
| 2024 | Class-Agnostic Object Counting with Text-to-Image Diffusion Model
Xiaofei Hui, Hossein Rahmani 0001, Jun Liu 0036 |
ECCV (69) | 3 |
| 2024 | Weakly Supervised Co-training with Swapping Assignments for Semantic Segmentation
Hossein Rahmani 0001, Sue Black 0002, Bryan M. Williams 0001 |
ECCV (56) | 2 |
| 2024 | Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
Zhengbo Zhang, Duo Peng, Hossein Rahmani 0001, Jun Liu 0036 |
ECCV (28) | 4 |
| 2024 | Towards Adaptive Pseudo-Label Learning for Semi-Supervised Temporal Action Localization
Feixiang Zhou, Bryan M. Williams 0001, Hossein Rahmani 0001 |
ECCV (62) | 3 |
| 2024 | Reverse2Complete: Unpaired Multimodal Point Cloud Completion via Guided DiffusionabstractUnpaired point cloud completion involves filling in missing parts of a point cloud without requiring partial-complete correspondence. Meanwhile, since point cloud completion is an ill-posed problem, there are multiple ways to generate the missing parts. Existing unpaired completion methods usually leverage generative adversarial training by transforming partial shape encoding into a complete one in the low-dimensional latent feature space. However, "mode collapse" often occurs, where only a subset of the shapes is represented in the low-dimensional space, reducing the diversity of the generated shapes. In this paper, we propose a novel unpaired multimodal shape completion approach that directly operates on point coordinate space. We achieve unpaired completion via a single diffusion model trained on complete data by "hijacking" the generative process. We further augment the diffusion model by introducing two guidance mechanisms to facilitate mapping the partial point cloud to the complete one while preserving its original structure. We conduct extensive evaluations of our approach, which show that our method generates shapes that are more diverse and better preserve the original structures compared to alternative methods. Hossein Rahmani 0001, Xun Yang 0001, Jun Liu 0036 |
ACM Multimedia | 2 |
| 2024 | DisC-GS: Discontinuity-aware Gaussian SplattingabstractRecently, Gaussian Splatting, a method that represents a 3D scene as a collection of Gaussian distributions, has gained significant attention in addressing the task of novel view synthesis. In this paper, we highlight a fundamental limitation of Gaussian Splatting: its inability to accurately render discontinuities and boundaries in images due to the continuous nature of Gaussian distributions. To address this issue, we propose a novel framework enabling Gaussian Splatting to perform discontinuity-aware image rendering. Additionally, we introduce a B\'ezier-boundary gradient approximation strategy within our framework to keep the ``differentiability'' of the proposed discontinuity-aware rendering process. Extensive experiments demonstrate the efficacy of our framework. Haoxuan Qu, Zhuoling Li, Hossein Rahmani 0001, Yujun Cai, Jun Liu 0036 |
NeurIPS | 3 |
| 2024 | DL-Reg: A deep learning regularization technique using linear regressionabstractRegularization is an essential aspect in the context of deep learning as it mitigates the risk of overfitting in deep neural networks. This study presents a novel deep learning regularization method, referred as DL-Reg, which effectively reduces the nonlinearity of deep networks by enforcing linearity to a certain extent. The proposed method is based on the incorporation of a linear constraint in the objective function of deep neural networks, which is defined as the error of a linear mapping from the inputs to the outputs of the model. Specifically, DL-Reg imposes a linear constraint on the network, which is further adjusted by a regularization factor, thereby preventing the network from overfitting. The effectiveness of DL-Reg is evaluated by training state-of-the-art deep network models on several benchmark datasets. The results of the experiments demonstrate that the proposed regularization method provides significant improvements over existing regularization techniques and enhances the performance of deep neural networks, particularly when dealing with small-sized training datasets. The main code of DL-Reg written in PyTorch is also available here: https://github.com/m2dgithub/DL-Reg.git (alternative link: shorturl.at/afmA4). Maryam Dialameh, Ali Hamzeh, Hossein Rahmani 0001, Safoura Dialameh, Hyock Ju Kwon |
Expert Syst. Appl. | 3 |
| 2024 | Deep orientated distance-transform network for geometric-aware centerline detection
Zheheng Jiang, Hossein Rahmani 0001, Plamen Angelov 0001, Ritesh Vyas, Huiyu Zhou 0001, Sue Black 0002, Bryan M. Williams 0001 |
Pattern Recognit. | 2 |
| 2024 | Progressive Channel-Shrinking NetworkabstractCurrently, salience-based channel pruning makes continuous breakthroughs in network compression. In the realization, the salience mechanism is used as a metric of channel salience to guide pruning. Therefore, salience-based channel pruning can dynamically adjust the channel width at run-time, which provides a flexible pruning scheme. However, there are two problems emerging: a gating function is often needed to truncate the specific salience entries to zero, which destabilizes the forward propagation; dynamic architecture brings more cost for indexing in inference which bottlenecks the inference speed. In this article, we propose a Progressive Channel-Shrinking (PCS) method to compress the selected salience entries at run-time instead of roughly approximating them to zero. We also propose a Running Shrinking Policy to provide a testing-static pruning scheme that can reduce the memory access cost for filter indexing. We evaluate our method on ImageNet and CIFAR10 datasets over two prevalent networks: ResNet and VGG, and demonstrate that our PCS outperforms all baselines and achieves state-of-the-art in terms of compression-performance tradeoff. Moreover, we observe a significant and practical acceleration of inference. The code is available athttps://github.com/JianhongPan-VLG/Progressive.Channel-Shrinking.Network. Jianhong Pan, Siyuan Yang 0001, Lin Geng Foo, Qiuhong Ke, Hossein Rahmani 0001, Zhipeng Fan 0001, Jun Liu 0036 |
IEEE Trans. Multim. | 5 |
| 2023 | Unified Pose Sequence ModelingabstractWe propose a Unified Pose Sequence Modeling approach to unify heterogeneous human behavior understanding tasks based on pose data, e.g., action recognition, 3D pose estimation and 3D early action prediction. A major obstacle is that different pose-based tasks require different output data formats. Specifically, the action recognition and prediction tasks require class predictions as outputs, while 3D pose estimation requires a human pose output, which limits existing methods to leverage task-specific network architectures for each task. Hence, in this paper, we propose a novel Unified Pose Sequence (UPS) model to unify heterogeneous output formats for the aforementioned tasks by considering text-based action labels and coordinate-based human poses as language sequences. Then, by optimizing a single auto-regressive transformer, we can obtain a unified output sequence that can handle all the aforementioned tasks. Moreover, to avoid the interference brought by the heterogeneity between different tasks, a dynamic routing mechanism is also proposed to empower our UPS with the ability to learn which subsets of parameters should be shared among different tasks. To evaluate the efficacy of the proposed UPS, extensive experiments are conducted on four different tasks with four popular behavior understanding benchmarks. Lin Geng Foo, Hossein Rahmani 0001, Qiuhong Ke, Jun Liu 0036 |
CVPR | 3 |
| 2023 | DiffPose: Toward More Reliable 3D Pose EstimationabstractMonocular 3D human pose estimation is quite challenging due to the inherent ambiguity and occlusion, which often lead to high uncertainty and indeterminacy. On the other hand, diffusion models have recently emerged as an effective tool for generating high-quality images from noise. In-spired by their capability, we explore a novel pose estimation framework (DiffPose) that formulates 3D pose estimation as a reverse diffusion process. We incorporate novel designs into our DiffPose to facilitate the diffusion process for 3D pose estimation: a pose-specific initialization of pose uncertainty distributions, a Gaussian Mixture Model-based forward diffusion process, and a context-conditioned re-verse diffusion process. Our proposed DiffPose significantly outperforms existing methods on the widely used pose estimation benchmarks Human3.6M and MPI-INF-3DHP. Project page: https://gongjia0208.github.io/Diffpose/. Jia Gong, Lin Geng Foo, Zhipeng Fan 0001, Qiuhong Ke, Hossein Rahmani 0001, Jun Liu 0036 |
CVPR | 5 |
| 2023 | A Probabilistic Attention Model with Occlusion-aware Texture Regression for 3D Hand Reconstruction from a Single RGB ImageabstractRecently, deep learning based approaches have shown promising results in 3D hand reconstruction from a single RGB image. These approaches can be roughly divided into model-based approaches, which are heavily dependent on the model's parameter space, and model-free approaches, which require large numbers of 3D ground truths to reduce depth ambiguity and struggle in weakly-supervised scenarios. To overcome these issues, we propose a novel probabilistic model to achieve the robustness of model-based approaches and reduced dependence on the model's parameter space of model-free approaches. The proposed probabilistic model incorporates a model-based network as a prior-net to estimate the prior probability distribution of joints and vertices. An Attention-based Mesh Vertices Uncertainty Regression (AMVUR) model is proposed to capture dependencies among vertices and the correlation between joints and mesh vertices to improve their feature representation. We further propose a learning based occlusion-aware Hand Texture Regression model to achieve high-fidelity texture reconstruction. We demonstrate the flexibility of the proposed probabilistic model to be trained in both supervised and weakly-supervised scenarios. The experimental results demonstrate our probabilistic model's state-of-the-art accuracy in 3D hand and texture reconstruction from a single image in both training schemes, including in the presence of severe occlusions. Zheheng Jiang, Hossein Rahmani 0001, Sue Black 0002, Bryan M. Williams 0001 |
CVPR | 2 |
| 2023 | Token Boosting for Robust Self-Supervised Visual Transformer Pre-trainingabstractLearning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially when we pre-train via the masked autoencoding approach, where both the inputs and masked “ground truth” targets can potentially be unreliable in this case. To address this limitation, we introduce the Token Boosting Module (TBM) as a plug-and-play component for VTs that effectively allows the VT to learn to extract clean and robust features during masked autoencoding pre-training. We provide theoretical analysis to show how TBM improves model pre-training with more robust and generalizable representations, thus benefiting down stream tasks. We conduct extensive experiments to analyze TBM's effectiveness, and results on four corrupted datasets demonstrate that TBM consistently improves performance on downstream tasks. Lin Geng Foo, Ping Hu 0001, Xindi Shang, Hossein Rahmani 0001, Zehuan Yuan, Jun Liu 0036 |
CVPR | 5 |
| 2023 | Distribution-Aligned Diffusion for Human Mesh RecoveryabstractRecovering a 3D human mesh from a single RGB image is a challenging task due to depth ambiguity and self-occlusion, resulting in a high degree of uncertainty. Meanwhile, diffusion models have recently seen much success in generating high-quality outputs by progressively denoising noisy inputs. Inspired by their capability, we explore a diffusion-based approach for human mesh recovery, and propose a Human Mesh Diffusion (HMDiff) framework which frames mesh recovery as a reverse diffusion process. We also propose a Distribution Alignment Technique (DAT) that injects input-specific distribution information into the diffusion process, and provides useful prior knowledge to simplify the mesh recovery task. Our method achieves state-of-the-art performance on three widely used datasets. Project page: https://gongjia0208.github.io/HMDiff/. Lin Geng Foo, Jia Gong, Hossein Rahmani 0001, Jun Liu 0036 |
ICCV | 3 |
| 2023 | Reinforced Learning for Label-Efficient 3D Face Reconstructionabstract3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face images preferably captured along different viewpoints of the subject. However, collecting the required large amounts of 3D annotated face data is expensive and time-consuming. To address the high annotation cost and due to the importance of training on a useful set, we propose an Active Learning (AL) framework that actively selects the most informative and representative samples to be labeled. To the best of our knowledge, this paper is the first work on tackling active learning for 3D face reconstruction to enable a label-efficient training strategy. In particular, we propose a Reinforcement Active Learning approach in conjunction with a clustering-based pooling strategy to select informative view-points of the subjects. Experimental results on 300W-LP and AFLW2000 datasets demonstrate that our proposed method is able to 1) efficiently select the most influencing view-points for labeling and outperforms several baseline AL techniques and 2) further improve the performance of a 3D Face Reconstruction network trained on the full dataset. Hoda Mohaghegh, Hossein Rahmani 0001, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
ICRA | 2 |
| 2023 | Robust monocular 3D face reconstruction under challenging viewing conditions
Hoda Mohaghegh, Farid Boussaïd, Hamid Laga, Hossein Rahmani 0001, Mohammed Bennamoun |
Neurocomputing | 4 |
| 2023 | GradMDM: Adversarial Attack on Dynamic NetworksabstractDynamic neural networks can greatly reduce computation redundancy without compromising accuracy by adapting their structures based on the input. In this paper, we explore the robustness of dynamic neural networks against energy-oriented attacks targeted at reducing their efficiency. Specifically, we attack dynamic models with our novel algorithm GradMDM. GradMDM is a technique that adjusts the direction and the magnitude of the gradients to effectively find a small perturbation for each input, that will activate more computational units of dynamic models during inference. We evaluate GradMDM on multiple datasets and dynamic models, where it outperforms previous energy-oriented attack techniques, significantly increasing computation complexity while reducing the perceptibility of the perturbations https://github.com/lingengfoo/GradMDM. Jianhong Pan, Lin Geng Foo, Qichen Zheng, Zhipeng Fan 0001, Hossein Rahmani 0001, Qiuhong Ke, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Human Action Recognition From Various Data Modalities: A ReviewabstractHuman Action Recognition (HAR) aims to understand human behavior and assign a label to each action. It has a wide range of applications, and therefore has been attracting increasing attention in the field of computer vision. Human actions can be represented using various data modalities, such as RGB, skeleton, depth, infrared, point cloud, event stream, audio, acceleration, radar, and WiFi signal, which encode different sources of useful yet distinct information and have various advantages depending on the application scenarios. Consequently, lots of existing works have attempted to investigate different types of approaches for HAR using various modalities. In this article, we present a comprehensive survey of recent progress in deep learning methods for HAR based on the type of input data modality. Specifically, we review the current mainstream deep learning methods for single data modalities and multiple data modalities, including the fusion-based and the co-learning-based frameworks. We also present comparative results on several benchmark datasets for HAR, together with insightful observations and inspiring future research directions. Zehua Sun, Qiuhong Ke, Hossein Rahmani 0001, Mohammed Bennamoun, Gang Wang 0012, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in VideosabstractExisting approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose variations in videos, which enables effective learning of 2D pose estimation in sparsely annotated videos. Specifically, we introduce a Motion Transformer (MT) module to perform cross frame reconstruction, aiming to learn motion dynamic knowledge in videos. Besides, a novel reinforcement learning-based Frame Selection Agent (FSA) is designed within our framework, which is able to harness informative frame pairs on the fly to enhance the pose estimator under our cross reconstruction mechanism. We conduct extensive experiments that show the efficacy of our proposed REMOTE framework. Xianzheng Ma, Hossein Rahmani 0001, Zhipeng Fan 0001, Bin Yang 0026, Jun Chen 0001, Jun Liu 0036 |
AAAI | 2 |
| 2022 | Meta Agent Teaming Active Learning for Pose EstimationabstractThe existing pose estimation approaches often require a large number of annotated images to attain good estimation performance, which are laborious to acquire. To reduce the human efforts on pose annotations, we propose a novel Meta Agent Teaming Active Learning (MATAL) framework to actively select and label informative images for effective learning. Our MATAL formulates the image selection procedure as a Markov Decision Process and learns an optimal sampling policy that directly maximizes the performance of the pose estimator based on the reward. Our framework consists of a novel state-action representation as well as a multi-agent team to enable batch sampling in the active learning procedure. The framework could be effectively optimized via Meta-Optimization to accelerate the adaptation to the gradually expanded labeled data during deployment. Finally, we show experimental results on both human hand and body pose estimation benchmark datasets and demonstrate that our method significantly outperforms all baselines continuously under the same amount of annotation budget. Moreover, to obtain similar pose estimation accuracy, our MATAL framework can save around 40% labeling efforts on average compared to state-of-the-art active learning frameworks. Jia Gong, Zhipeng Fan 0001, Qiuhong Ke, Hossein Rahmani 0001, Jun Liu 0036 |
CVPR | 4 |
| 2022 | Graph-context Attention Networks for Size-varied Deep Graph MatchingabstractDeep learning for graph matching has received growing interest and developed rapidly in the past decade. Although recent deep graph matching methods have shown excellent performance on matching between graphs of equal size in the computer vision area, the size-varied graph matching problem, where the number of keypoints in the images of the same category may vary due to occlusion, is still an open and challenging problem. To tackle this, we firstly propose to formulate the combinatorial problem of graph matching as an Integer Linear Programming (ILP) problem, which is more flexible and efficient to facilitate comparing graphs of varied sizes. A novel Graph-context Attention Network (GCAN), which jointly capture intrinsic graph structure and cross-graph information for improving the discrimination of node features, is then proposed and trained to resolve this ILP problem with node correspondence supervision. We further show that the proposed GCAN model is efficient to resolve the graph-level matching problem and is able to automatically learn node-to-node similarity via graph-level matching. The proposed approach is evaluated on three public keypoint-matching datasets and one graph-matching dataset for blood vessel patterns, with experimental results showing its superior performance over existing state-of-the-art algorithms for keypoint and graph-level matching. Zheheng Jiang, Hossein Rahmani 0001, Plamen Angelov 0001, Sue Black 0002, Bryan M. Williams 0001 |
CVPR | 2 |
| 2022 | ERA: Expert Retrieval and Assembly for Early Action Prediction
Lin Geng Foo, Hossein Rahmani 0001, Qiuhong Ke, Jun Liu 0036 |
ECCV (34) | 3 |
| 2022 | Dynamic Spatio-Temporal Specialization Learning for Fine-Grained Action Recognition
Lin Geng Foo, Qiuhong Ke, Hossein Rahmani 0001, Anran Wang 0001, Jun Liu 0036 |
ECCV (4) | 4 |
| 2022 | GradAuto: Energy-Oriented Attack on Dynamic Neural Networks
Jianhong Pan, Qichen Zheng, Zhipeng Fan 0001, Hossein Rahmani 0001, Qiuhong Ke, Jun Liu 0036 |
ECCV (4) | 4 |
| 2022 | IGFormer: Interaction Graph Transformer for Skeleton-Based Human Interaction Recognition
Yunsheng Pang, Qiuhong Ke, Hossein Rahmani 0001, James Bailey 0001, Jun Liu 0036 |
ECCV (25) | 3 |
| 2022 | Multi-Branch with Attention Network for Hand-Based Person RecognitionabstractIn this paper, we propose a novel hand-based person recognition method for the purpose of criminal investigations since the hand image is often the only available information in cases of serious crime such as sexual abuse. Our proposed method, Multi-Branch with Attention Network (MBA-Net), incorporates both channel and spatial attention modules in branches in addition to a global (without attention) branch to capture global structural information for discriminative feature learning. The attention modules focus on the relevant features of the hand image while suppressing the irrelevant backgrounds. In order to overcome the weakness of the attention mechanisms, equivariant to pixel shuffling, we integrate relative positional encodings into the spatial attention module to capture the spatial positions of pixels. Extensive evaluations on two large multi-ethnic and publicly available hand datasets demonstrate that our proposed method achieves state-of-the-art performance, surpassing the existing hand-based identification methods. The source code is available at https://github.com/nathanlem1/MBA-Net. Nathanael L. Baisa, Bryan M. Williams 0001, Hossein Rahmani 0001, Plamen Angelov 0001, Sue Black 0002 |
ICPR | 3 |
| 2022 | Self-Supervised Learning With Adaptive Distillation for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is an important topic in the community of remote sensing, which has a wide range of applications in geoscience. Recently, deep learning-based methods have been widely used in HSI classification. However, due to the scarcity of labeled samples in HSI, the potential of deep learning-based methods has not been fully exploited. To solve this problem, a self-supervised learning (SSL) method with adaptive distillation is proposed to train the deep neural network with extensive unlabeled samples. The proposed method consists of two modules: adaptive knowledge distillation with spatial–spectral similarity and 3-D transformation on HSI cubes. The SSL with adaptive knowledge distillation uses the self-supervised information to train the network by knowledge distillation, where self-supervised knowledge is the adaptive soft label generated by spatial–spectral similarity measurement. The SSL with adaptive knowledge distillation mainly includes the following three steps. First, the similarity between unlabeled samples and object classes in HSI is generated based on the spatial–spectral joint distance (SSJD) between unlabeled samples and labeled samples. Second, the adaptive soft label of each unlabeled sample is generated to measure the probability that the unlabeled sample belongs to each object class. Third, a progressive convolutional network (PCN) is trained by minimizing the cross-entropy between the adaptive soft labels and the probabilities generated by the forward propagation of the PCN. The SSL with 3-D transformation rotates the HSI cube in both the spectral domain and the spatial domain to fully exploit the labeled samples. Experiments on three public HSI data sets have demonstrated that the proposed method can achieve better performance than existing state-of-the-art methods. Jun Yue 0004, Leyuan Fang, Hossein Rahmani 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Robust End-to-End Hand Identification via Holistic Multi-Unit Knuckle RecognitionabstractIn many cases of serious crime, images of a hand can be the only evidence available for the forensic identification of the offender. As well as placing them at the scene, such images and video evidence offer proof of the offender committing the crime. The knuckle creases of the human hand have emerged as an effective biometric trait and been used to identify the perpetrators of child abuse in forensic investigations. However, manual utilization of knuckle creases for identification is highly time consuming and can be subjective, requiring the expertise of experienced forensic anthropologists whose availability is very limited. Hence, there arises a need for an automated approach for localization and comparison of knuckle patterns. In this paper, we present a fully automatic end-to-end approach which localizes the minor, major and base knuckles in images of the hand, and effectively uses them for identification achieving state-of-the-art results. This work improves on existing approaches and allows us to strengthen cases further by objectively combining multiple knuckles and knuckle types to obtain a holistic matching result for comparing two hands. This yields a stronger and more robust multi-unit biometric and facilitates the large-scale examination of the potential of knuckle-based identification. Evaluated on two large landmark datasets, the proposed framework achieves equal error rates (EER) of 1.0-1.9%, rank-1 accuracies of 99.3-100% and decidability indices of 5.04-5.83. We make the full results available via a novel online GUI to raise awareness with the general public and forensic investigators about the identifiability of various knuckle regions. These strong results demonstrate the value of our holistic approach to hand identification from knuckle patterns and their utility in forensic investigations. Ritesh Vyas, Hossein Rahmani 0001, Ricki Boswell-Challand, Plamen Angelov 0001, Sue Black 0002, Bryan M. Williams 0001 |
IJCB | 2 |
| 2021 | Else-Net: Elastic Semantic Network for Continual Action Recognition from Skeleton DataabstractMost of the state-of-the-art action recognition methods focus on offline learning, where the samples of all types of actions need to be provided at once. Here, we address continual learning of action recognition, where various types of new actions are continuously learned over time. This task is quite challenging, owing to the catastrophic forgetting problem stemming from the discrepancies between the previously learned actions and current new actions to be learned. Therefore, we propose Else-Net, a novel Elastic Semantic Network with multiple learning blocks to learn diversified human actions over time. Specifically, our Else-Net is able to automatically search and update the most relevant learning blocks w.r.t. the current new action, or explore new blocks to store new knowledge, preserving the unmatched ones to retain the knowledge of previously learned actions and alleviates forgetting when learning new actions. Moreover, even though different human actions may vary to a large extent as a whole, their local body parts can still share many homogeneous features. Inspired by this, our proposed Else-Net mines the shared knowledge of the decomposed human body parts from different actions, which benefits continual learning of actions. Experiments show that the proposed approach enables effective continual action recognition and achieves promising performance on two large-scale action recognition datasets. Qiuhong Ke, Hossein Rahmani 0001, Rui En Ho, Henghui Ding, Jun Liu 0036 |
ICCV | 3 |
| 2020 | Learning Latent Global Network for Skeleton-Based Action PredictionabstractHuman actions represented with 3D skeleton sequences are robust to clustered backgrounds and illumination changes. In this paper, we investigate skeleton-based action prediction, which aims to recognize an action from a partial skeleton sequence that contains incomplete action information. We propose a new Latent Global Network based on adversarial learning for action prediction. We demonstrate that the proposed network provides latent long-term global information that is complementary to the local action information of the partial sequences and helps improve action prediction. We show that action prediction can be improved by combining the latent global information with the local action information. We test the proposed method on three challenging skeleton datasets and report state-of-the-art performance. Qiuhong Ke, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Image Process. | 3 |
| 2019 | Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
Jian Liu 0014, Hossein Rahmani 0001, Naveed Akhtar, Ajmal Mian |
Int. J. Comput. Vis. | 2 |
| 2019 | Single image dehazing using deep neural networks
Cameron Hodges, Mohammed Bennamoun, Hossein Rahmani 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Global Regularizer and Temporal-Aware Cross-Entropy for Skeleton-Based Early Action Recognition
Qiuhong Ke, Jun Liu 0036, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
ACCV (4) | 4 |
| 2018 | Learning a Deep Model for Human Action Recognition from Novel ViewpointsabstractRecognizing human actions from unknown and unseen (novel) views is a challenging problem. We propose a Robust Non-Linear Knowledge Transfer Model (R-NKTM) for human action recognition from novel views. The proposed R-NKTM is a deep fully-connected neural network that transfers knowledge of human actions from any unknown view to a shared high-level virtual view by finding a set of non-linear transformations that connects the views. The R-NKTM is learned from 2D projections of dense trajectories of synthetic 3D human models fitted to real motion capture data and generalizes to real videos of human actions. The strength of our technique is that we learn a single R-NKTM for all actions and all viewpoints for knowledge transfer of any real human action video without the need for re-training or fine-tuning the model. Thus, R-NKTM can efficiently scale to incorporate new action classes. R-NKTM is learned with dummy labels and does not require knowledge of the camera viewpoint at any stage. Experiments on three benchmark cross-view human action datasets show that our method outperforms existing state-of-the-art. Hossein Rahmani 0001, Ajmal Mian, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Learning Action Recognition Model from Depth and Skeleton VideosabstractDepth sensors open up possibilities of dealing with the human action recognition problem by providing 3D human skeleton data and depth images of the scene. Analysis of human actions based on 3D skeleton data has become popular recently, due to its robustness and view-invariant representation. However, the skeleton alone is insufficient to distinguish actions which involve human-object interactions. In this paper, we propose a deep model which efficiently models human-object interactions and intra-class variations under viewpoint changes. First, a human body-part model is introduced to transfer the depth appearances of body-parts to a shared view-invariant space. Second, an end-to-end learning framework is proposed which is able to effectively combine the view-invariant body-part representation from skeletal and depth images, and learn the relations between the human body-parts and the environmental objects, the interactions between different human body-parts, and the temporal structure of human actions. We have evaluated the performance of our proposed model against 15 existing techniques on two large benchmark human action recognition datasets including NTU RGB+D and UWA3DII. The Experimental results show that our technique provides a significant improvement over state-of-the-art methods. Hossein Rahmani 0001, Mohammed Bennamoun |
ICCV | 1 |
| 2016 | 3D Action Recognition from Novel ViewpointsabstractWe propose a human pose representation model that transfers human poses acquired from different unknown views to a view-invariant high-level space. The model is a deep convolutional neural network and requires a large corpus of multiview training data which is very expensive to acquire. Therefore, we propose a method to generate this data by fitting synthetic 3D human models to real motion capture data and rendering the human poses from numerous viewpoints. While learning the CNN model, we do not use action labels but only the pose labels after clustering all training poses into k clusters. The proposed model is able to generalize to real depth images of unseen poses without the need for re-training or fine-tuning. Real depth videos are passed through the model frame-wise to extract view-invariant features. For spatio-temporal representation, we propose group sparse Fourier Temporal Pyramid which robustly encodes the action specific most discriminative output features of the proposed human pose model. Experiments on two multiview and three single-view benchmark datasets show that the proposed method dramatically outperforms existing state-of-the-art in action recognition. Hossein Rahmani 0001, Ajmal Mian |
CVPR | 1 |
| 2016 | Histogram of Oriented Principal Components for Cross-View Action RecognitionabstractExisting techniques for 3D action recognition are sensitive to viewpoint variations because they extract features from depth images which are viewpoint dependent. In contrast, we directly process pointclouds for cross-view action recognition from unknown and unseen views. We propose the histogram of oriented principal components (HOPC) descriptor that is robust to noise, viewpoint, scale and action speed variations. At a 3D point, HOPC is computed by projecting the three scaled eigenvectors of the pointcloud within its local spatio-temporal support volume onto the vertices of a regular dodecahedron. HOPC is also used for the detection of spatio-temporal keypoints (STK) in 3D pointcloud sequences so that view-invariant STK descriptors (or Local HOPC descriptors) at these key locations only are used for action recognition. We also propose a global descriptor computed from the normalized spatio-temporal distribution of STKs in 4-D, which we refer to as STK-D. We have evaluated the performance of our proposed descriptors against nine existing techniques on two cross-view and three single-view human action recognition datasets. The experimental results show that our techniques provide significant improvement over state-of-the-art methods. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Discriminative human action classification using locality-constrained linear coding
Hossein Rahmani 0001, Du Q. Huynh, Arif Mahmood, Ajmal Mian |
Pattern Recognit. Lett. | 1 |
| 2015 | Learning a non-linear knowledge transfer model for cross-view action recognitionabstractThis paper concerns action recognition from unseen and unknown views. We propose unsupervised learning of a non-linear model that transfers knowledge from multiple views to a canonical view. The proposed Non-linear Knowledge Transfer Model (NKTM) is a deep network, with weight decay and sparsity constraints, which finds a shared high-level virtual path from videos captured from different unknown viewpoints to the same canonical view. The strength of our technique is that we learn a single NKTM for all actions and all camera viewing directions. Thus, NKTM does not require action labels during learning and knowledge of the camera viewpoints during training or testing. NKTM is learned once only from dense trajectories of synthetic points fitted to mocap data and then applied to real video data. Trajectories are coded with a general codebook learned from the same mocap data. NKTM is scalable to new action classes and training data as it does not require re-learning. Experiments on the IXMAS and N-UCLA datasets show that NKTM outperforms existing state-of-the-art methods for cross-view action recognition. Hossein Rahmani 0001, Ajmal Mian |
CVPR | 1 |
| 2014 | HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ECCV (2) | 1 |
| 2014 | Action Classification with Locality-Constrained Linear CodingabstractWe propose an action classification algorithm which uses Locality-constrained Linear Coding (LLC) to capture discriminative information of human body variations in each spatio-temporal subsequence of a video sequence. Our proposed method divides the input video into equally spaced overlapping spatio-temporal sub sequences, each of which is decomposed into blocks and then cells. We use the Histogram of Oriented Gradient (HOG3D) feature to encode the information in each cell. We justify the use of LLC for encoding the block descriptor by demonstrating its superiority over Sparse Coding (SC). Our sequence descriptor is obtained via a logistic regression classifier with L2 regularization. We evaluate and compare our algorithm with ten state-of-the-art algorithms on five benchmark datasets. Experimental results show that, on average, our algorithm gives better accuracy than these ten algorithms. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ICPR | 1 |
| 2014 | Real time action recognition using histograms of depth gradients and random decision forestsabstractWe propose an algorithm which combines the discriminative information from depth images as well as from 3D joint positions to achieve high action recognition accuracy. To avoid the suppression of subtle discriminative information and also to handle local occlusions, we compute a vector of many independent local features. Each feature encodes spatiotemporal variations of depth and depth gradients at a specific space-time location in the action volume. Moreover, we encode the dominant skeleton movements by computing a local 3D joint position difference histogram. For each joint, we compute a 3D space-time motion volume which we use as an importance indicator and incorporate in the feature vector for improved action discrimination. To retain only the discriminant features, we train a random decision forest (RDF). The proposed algorithm is evaluated on three standard datasets and compared with nine state-of-the-art algorithms. Experimental results show that, on the average, the proposed algorithm outperform all other algorithms in accuracy and have a processing speed of over 112 frames/second. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
WACV | 1 |
| 2011 | Hardware design of a new genetic based disk scheduling method
Hossein Rahmani 0001, Mohammad Reza Bonyadi, Amir Momeni, Mohsen Ebrahimi Moghaddam, Maghsoud Abbaspour |
Real Time Syst. | 1 |
| 2010 | A genetic based disk scheduling method to decrease makespan and missed tasks
Mohammad Reza Bonyadi, Hossein Rahmani 0001, Mohsen Ebrahimi Moghaddam |
Inf. Syst. | 2 |
| 2010 | A new real time disk-scheduling method based on GSR algorithm
Hossein Rahmani 0001, Mohammad Mehdi Faghih, Mohsen Ebrahimi Moghaddam |
J. Syst. Softw. | 1 |