Tonmoy Hossain

dblp:287/0350 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-5808-6391ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 50% Image recognition and object detection · 50%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d shape representation
1.012026
Learning to Transform: Unifying Latent Geometric Shape and Appearance Representations in Healthcare Imaging · AAAI 2026
Computer vision › Image recognition and object detection
medical image analysis
1.012026
Learning to Transform: Unifying Latent Geometric Shape and Appearance Representations in Healthcare Imaging · AAAI 2026

Methods — techniques the papers use, named apart from their topics

neural network · 1.0deep learning · 1.0
YearPublicationVenuePosition
2026 Learning to Transform: Unifying Latent Geometric Shape and Appearance Representations in Healthcare Imaging
abstract
Recent advances in deep neural networks have highlighted the importance of geometric shape in various image analysis and computer vision tasks. However, most current approaches rely on coarse or simplified shape representations, such as binary masks, meshes, or point clouds, that are primarily designed to capture global structures of objects presented in images. While effective for general image and visual understanding, these methods often fail to learn fine-grained geometric information that is critical for accurately modeling complex shapes and subtle anatomical variations. This limitation is particularly consequential in healthcare applications, where understanding fine-grained anatomical shapes and their changes is crucial for accurate disease detection and diagnosis. My research focuses on developing a set of advanced deep learning frameworks that learn robust and complex shape representations from dense image data and integrate them into the current paradigm of image appearance and texture learning.
Tonmoy Hossain
AAAI1
2026 Learning Group Actions In Disentangled Latent Image Representations
abstract
Modeling group actions on latent representations enables controllable transformations of high-dimensional image data. Prior works applying group-theoretic priors or modeling transformations typically operate in the high-dimensional data space, where group actions apply uniformly across the entire input, making it difficult to disentangle the subspace that varies under transformations. While latent-space methods offer greater flexibility, they still require manual partitioning of latent variables into equivariant and invariant subspaces, limiting the ability to robustly learn and operate group actions within the representation space. To address this, we introduce a novel end-to-end framework that for the first time learns group actions on latent image manifolds, automatically discovering transformation-relevant structures without manual intervention. Our method uses learnable binary masks with straight-through estimation to dynamically partition latent representations into transformation-sensitive and invariant components. We formulate this within a unified optimization framework that jointly learns latent disentanglement and group transformation mappings. The framework can be seamlessly integrated with any standard encoder-decoder architecture. We validate our approach on five 2D/3D image datasets, demonstrating its ability to automatically learn disentangled latent factors for group actions in diverse data, while downstream classification tasks confirm the effectiveness of the learned representations. Our code is publicly available at GitHub.
Farhana Hossain Swarnali, Miaomiao Zhang 0002, Tonmoy Hossain
WACV3
2025 Invariant Shape Representation Learning for Image Classification
abstract
Geometric shape features have been widely used as strong predictors for image classification. Nevertheless, most existing classifiers such as deep neural networks (DNNs) directly leverage the statistical correlations between these shape features and target variables. However, these correlations can often be spurious and unstable across different environments (e.g., in different age groups, certain types of brain changes have unstable relations with neurodegenerative disease); hence leading to biased or inaccurate predictions. In this paper, we introduce a novel framework that for the first time develops invariant shape representation learning (ISRL) to further strengthen the robustness of image classifiers. In contrast to existing approaches that mainly derive features in the image space, our model ISRL is designed to jointly capture invariant features in latent shape spaces parameterized by deformable transformations. To achieve this goal, we develop a new learning paradigm based on invariant risk minimization (IRM) to learn invariant representations of image and shape features across multiple training distributions/environments. By embedding the features that are invariant with regard to target variables in different environments, our model consistently offers more accurate predictions. We validate our method by performing classification tasks on both simulated 2D images, real 3D brain and cine cardiovascular magnetic resonance images (MRIs). Our code is publicly available at https://github.com/tonmoy-hossain/ISRL.
Tonmoy Hossain, Jing Ma 0002, Jundong Li, Miaomiao Zhang 0002
WACV1