Kumie Gedamu

dblp:295/8875 · also Kumie Alemu Gedamu · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-6458-1882ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vision-Language Collaborative Representation Learning for Action Quality Assessment
abstract
Action Quality Assessment (AQA) has gained significant attention due to its potential real-world applications, which require a fine-grained understanding of action sequences. Recent works have attempted to utilize multimodal video features and address some existing challenges. However, these approaches primarily focus on leveraging textual information from language models only, leading to instability and suboptimal performance due to directional bias in a vision-language joint embedding space. To tackle these issues, we propose a Vision-Language Collaboration Representation Learning approach (VLC-Net) to understand fine-grained action sequences and create a unified feature representation along with their temporal dependencies for accurate AQA score prediction. Specifically, we design a bidirectional knowledge distillation operation to perform collaboration learning between vision-language pre-trained knowledge and visual action knowledge for fine-grained action feature learning. Furthermore, we design vision-language alignment guidance to explicitly align action features with the same action semantics across modalities, thereby unifying their joint representation. Leveraging these aligned features, we propose multimodal contrastive learning to relate modalities and align subactions with textual descriptions, ensuring accurate action representation. We conduct experiments on the FineDiving, MTL-AQA, FineFS, and Fis-V datasets, demonstrating the effectiveness of our approach, which outperforms state-of-the-art methods.
Kumie Gedamu, Yanli Ji, Wangmeng Zuo, Jamal Bentahar, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen
IEEE Trans. Image Process.1
2026 DUCore: Dual Uncertainty-Guided Consistency and Regional Contrastive Learning for Semi-Supervised Medical Image Segmentation
abstract
Uncertainty-aware consistency learning is one of the reliable approaches in semi-supervised medical image segmentation, enforcing robust model predictions under various perturbations. However, existing methods often rely on multiple stochastic predictions or dual-network/decoder discrepancies to estimate uncertainty, which increases computational cost and discards uncertain regions, potentially missing complex structures such as ambiguous lesion boundaries. To address these challenges, we introduce a Dual Uncertainty-Guided Consistency and Regional Contrastive Learning (DUCore) framework. DUCore improves segmentation robustness by integrating two complementary loss functions within consistency learning. The dual uncertainty-guided consistency loss (DuCL) adaptively calibrates the prediction alignment by prioritizing uncertain regions. DuCL uses deterministic single-pass uncertainty estimation, employing entropy-based calibration for aleatoric uncertainty and Proxy Dirichlet calibration for epistemic uncertainty. These uncertainty measures are computed directly from network output, and moderately uncertain regions are weighted instead of being discarded, which preserves valuable learning signals. The Regional Contrastive Loss (ReCL) further refines feature separability using boundary- and gradient-based hard negative mining in the encoded representation space. By explicitly targeting structural ambiguities, ReCL distinguishes lesion and organ edges from visually similar boundary-adjacent regions and mitigates intensity overlaps in gradient-rich transitions. As a result, DUCore is able to delineate fine structures and complex boundaries with higher precision. Extensive experiments on various medical segmentation benchmarks reveal that DUCore outperforms existing consistency methods.
Maregu Assefa, Muzammal Naseer, Kumie Gedamu, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Mohamed L. Seghier, Ernesto Damiani, Naoufel Werghi
IEEE J. Biomed. Health Informatics3
2025 Decoupling Representations with Quantized Vectors for Semi-Supervised Action Quality Assessment
abstract
Semi-supervised Action Quality Assessment (AQA) aims to predict action execution scores by utilizing limited labeled and massive unlabeled samples. However, existing approaches often oversimplify the modeling process of scene-invariant and fine-grained action sequences for score prediction, which fails to fully leverage rich information available in unlabeled data. Thus, we propose a Vector Quantized Decoupling representation Network (VQD-Net), which decouples sub-action categories and action execution quality to enable a fine-grained understanding of actions for semi-supervised AQA. The proposed VQD-Net effectively captures and learns discriminative features through common semantic representations between labeled and unlabeled samples within a shared embedding space, enabling accurate AQA score prediction. By leveraging the differences between features and embeddings, we achieve more accurate confidence estimates for unlabeled samples and enhance the model performance by selecting reliable pseudo-labels. Experiments on three public AQA datasets, including MTL-AQA, RG, and FineFS, demonstrate that the proposed VQD-Net achieves state-of-the-art performance. The source code is available at https://github.com/Pix0611/VQD-Net.
Lingfeng Ye, Kumie Gedamu, Jie Shao 0001
ICME2
2025 Visual-Semantic Alignment Temporal Parsing for Action Quality Assessment
abstract
Action Quality Assessment (AQA) is a challenging task involving analyzing fine-grained technical subactions, aligning high-level visual-semantic representations, and exploring internal temporal structures that capture the overall meaning of given action sequences. To address these challenges, we propose a Visual-semantic Alignment Temporal Parsing Network (VATP-Net) to understand the high-level visual semantics of subaction sequences and internal temporal structures without explicit supervision for action quality assessment. The proposed approach designs a self-supervised temporal parsing module to generate subaction sequences from the given video by aligning the visual and semantic action features. It captures high-level semantics and the internal temporal dynamics of subaction sequences. Furthermore, a multimodal interaction module is proposed to capture the interaction between different modalities of action features, enabling a comprehensive assessment of fine-grained and scene-invariant action details. The proposed module captures the intricate relationships and encourages interactions between different modalities within an action sequence, enhancing the overall understanding of action assessment. We exhaustively evaluate our proposed approach on the MTL-AQA, Rhythmic Gymnastics (RG), FineFS, and Fis-V datasets. Extensive experimental results demonstrate the effectiveness and feasibility of our proposed approach, which outperforms state-of-the-art methods by a significant margin.
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.1
2024 Self-Supervised Sub-Action Parsing Network for Semi-Supervised Action Quality Assessment
abstract
Semi-supervised Action Quality Assessment (AQA) using limited labeled and massive unlabeled samples to achieve high-quality assessment is an attractive but challenging task. The main challenge relies on how to exploit solid and consistent representations of action sequences for building a bridge between labeled and unlabeled samples in the semi-supervised AQA. To address the issue, we propose a Self-supervised sub-Action Parsing Network (SAP-Net) that employs a teacher-student network structure to learn consistent semantic representations between labeled and unlabeled samples for semi-supervised AQA. We perform actor-centric region detection and generate high-quality pseudo-labels in the teacher branch and assists the student branch in learning discriminative action features. We further design a self-supervised sub-action parsing solution to locate and parse fine-grained sub-action sequences. Then, we present the group contrastive learning with pseudo-labels to capture consistent motion-oriented action features in the two branches. We evaluate our proposed SAP-Net on four public datasets: the MTL-AQA, FineDiving, Rhythmic Gymnastics, and FineFS datasets. The experiment results show that our approach outperforms state-of-the-art semi-supervised methods by a significant margin.
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen
IEEE Trans. Image Process.1
2024 Audio-Visual Contrastive and Consistency Learning for Semi-Supervised Action Recognition
abstract
Semi-supervised video learning is an increasingly popular approach for improving video understanding tasks by utilizing large-scale unlabeled videos along with a few labels. Recent studies have shown that multimodal contrastive learning and consistency regularization are effective techniques for generating high-quality pseudo-labels for semi-supervised action recognition. However, existing pseudo-labeling approaches are solely based on the model's class predictions and can suffer from confirmation biases due to the accumulation of false predictions. To address this issue, we propose exploiting audio-visual feature correlations to achieve high-quality pseudo-labels instead of relying on model confidence. To achieve this goal, we introduce Audio-visual Contrastive and Consistency Learning (AvCLR) for semi-supervised action recognition. AvCLR generates reliable pseudo-labels from audio-visual feature correlations using deep embedded clustering to mitigate confirmation biases. Additionally, AvCLR introduces two contrastive modules: intra-modal contrastive learning (ImCL) and cross-modal contrastive learning (XmCL) to discover complementary information from audio-visual alignments. The ImCL module learns informative representations within audio and video independently, while the XmCL module aims to leverage global high-level features of audio-visual information. Furthermore, the XmCL is constrained by introducing intra-instance negatives from one modality to the other. We jointly optimize the model with ImCL, XmCL, and consistency regularization in an end-to-end semi-supervised manner. Experimental results have demonstrated that the proposed AvCLR framework is effective in reducing confirmation biases and outperforms existing confidence-based semi-supervised action recognition methods.
Maregu Assefa, Wei Jiang 0016, Jinyu Zhan, Kumie Gedamu, Getinet Yilma, Melese Ayalew, Deepak Adhikari
IEEE Trans. Multim.4
2023 Relation-mining self-attention network for skeleton-based human action recognition
Kumie Gedamu, Yanli Ji, Lingling Gao, Yang Yang 0002, Heng Tao Shen
Pattern Recognit.1
2023 Actor-Aware Self-Supervised Learning for Semi-Supervised Video Representation Learning
abstract
Self-supervised contrastive learning has shown a significant improvement in performance for action recognition tasks by discovering useful signals from unlabeled videos. Nevertheless, the unique features of existing video benchmark datasets have led the learned video representations to be contextually biased toward dominant backgrounds and scene correlations. Thus, ultimately leading to poor generalizations on scene-invariant action recognition. Therefore, we propose Actor-aware Self-supervised Learning for Semi-supervised Video Representation Learning (ActorSL). We aligned localized actors and their corresponding scene information to encourage the model to learn discriminative regions and mitigate the model’s dependency on the video background during contrastive training. Furthermore, we present an inter-video Background Mixing (iBM) augmentation strategy to introduce scene consistency into the model. We patch inter-video crops of four randomly selected frames for iBM to create a unique frame for each video. The patched frame is blended with the target video frames to generate a spatially augmented sample. Then, the actor-scene aligned features and features of iBM-augmented videos are utilized to optimize contrastive loss and consistency regularization jointly in a semi-supervised way. Moreover, iBM combines the one-hot-encoded labels of patches with the label of the target video as a label smoothing regularizer to soften the decision boundaries of the semi-supervised model. Our experimental results reveal that, ActorSL notably improved current state-of-the-art semi-supervised methods on the Kinetics-400, UCF101, and HMDB51 datasets under a low-label regime. Code released athttps://github.com/Endarzboy/ActorSL.
Maregu Assefa, Wei Jiang 0016, Kumie Gedamu, Getinet Yilma, Deepak Adhikari, Melese Ayalew, Aiman Erbad
IEEE Trans. Circuits Syst. Video Technol.3
2023 Fine-Grained Spatio-Temporal Parsing Network for Action Quality Assessment
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Jie Shao 0001, Heng Tao Shen
IEEE Trans. Image Process.1
2023 Self-Supervised Scene-Debiasing for Video Representation Learning via Background Patching
abstract
Self-supervised learning has considerably improved video representation learning by discovering supervisory signals automatically from unlabeled videos. However, due to the scene-biased nature of existing video datasets, the current methods are biased to the dominant scene context during action inference. Hence, this paper proposes Background Patching (BP), a scene-debiasing augmentation strategy to alleviate the model reliance on the video background in a self-supervised contrastive manner. The BP reduces the negative influence of the video background by mixing a randomly patched frame to the video background. BP randomly crops four frames from four different videos and patches them to construct a new frame for each video separately. The patched frame is mixed with all frames of the target video to produce a spatially distorted video sample. Then, we use existing self-supervised contrastive frameworks to pull representations of the distorted and original videos closer together. Moreover, BP mixes the semantic labels of patches with the target video's label, resulting in the regularization of the contrastive model to soften the decision boundaries in the embedding space. Therefore, the model is explicitly constrained to suppress the background influence by emphasizing more on the motion changes. The extensive experimental results show that our BP significantly improved the performance of various video understanding downstream tasks including action recognition, action detection, and video retrieval.
Maregu Assefa, Wei Jiang 0016, Kumie Gedamu, Getinet Yilma, Bulbula Kumeda, Melese Ayalew
IEEE Trans. Multim.3
2022 Actor-Aware Contrastive Learning for Semi-Supervised Action Recognition
abstract
The unique features of existing video datasets have led self-supervised contrastive learning to scene correlations and background biases, resulting in poor generalization in scene-invariant action recognition. Therefore, we propose Actor-aware Contrastive Learning for semi-supervised action recognition (ActorCLR). We employ localized actors to encourage the model to learn discriminative regions and mitigate the model's reliance on the video background during contrastive training. Furthermore, we introduce Inter-video Background Mixing (iBM) augmentation strategy to inject scene-invariance into the model. For iBM, we patch inter-video crops of four randomly selected frames to create a distinct frame for each video individually. The patched frame is mixed with the target video frames to produce a spatially distorted sample. Then, we jointly optimize contrastive loss and consistency regularization with localized actors and corresponding iBM-augmented videos in a semi-supervised manner. iBM also mixes the one-hot-encoded labels of patches with the target video's label, which softens the decision boundaries of the semi-supervised model. Our experimental results show that ActorCLR significantly improved action recognition on Kinetics-400, UCF101, and HMDB51 datasets under a low-label regime.
Maregu Assefa, Wei Jiang 0016, Kumie Gedamu, Getinet Yilma, Melese Ayalew, Mohammed Seid
ICTAI3
2022 View-Invariant Human Action Recognition Via View Transformation Network (VTN)
abstract
Since the human body is non-rigid, actions captured in different views always involve action occlusion and information loss. Recently, view-variation-related human action recognition is still a challenging problem. To address the problem, we propose a View Transformation Network (VTN) that realizes the view normalization by transforming arbitrary-view action samples to a base view to seek for a view-invariant representation. an attention learning module is designed to learn a co-attention for action samples of different views, that contributes to output a similar feature representation to erase the view diversity in different views. Extensive and fair evaluations are performed on the UESTC varying-view RGB-D dataset, the NTU RGB-D 60 dataset, and the NTU RGB-D 120 dataset, where three evaluation types,i.e.X-subject, X-view, and A-view recognition, are performed. Experiments illustrate that our VTN model achieves outstanding performance.
Lingling Gao, Yanli Ji, Kumie Gedamu, Xiaofeng Zhu 0001, Xing Xu 0001, Heng Tao Shen
IEEE Trans. Multim.3
2021 Arbitrary-view human action recognition via novel-view action generation
Kumie Gedamu, Yanli Ji, Yang Yang 0002, Lingling Gao, Heng Tao Shen
Pattern Recognit.1