Chaitanya Kaul

dblp:236/4424 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-4893-6222ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 HandSolo: A Mid-Air Hand Pose Interaction Method Based on Disentangled Degrees-of-Hand-Freedom
abstract
This study aims to utilise mid-air hand-pose movements to implement various interactive controls, e.g. dial and slider controlling, through independent low-dimensional embeddings. Towards this, we develop a novel adjustable hand-pose space disentanglement approach for a learnable VAE-based high-to-low dimensional embedding model (HandSolo). It disentangles the latent embeddings intomultiple independent one- or two-dimensional embedding spaces, enabling independent control. HandSolo allows multi-dimensional settings and multi-DOF combinations, providing a new paradigm for flexible and extensible hand-pose interaction systems. Additionally, to exploit model potential and make user interaction comfortable, we propose a visual interaction evaluation strategy (VIEs) to help system designers understand model capability and user habits. Finally, we provide an example virtual interaction system that integrates various virtual interaction objects, showing how our innovations improve their interaction capabilities. Experimental user studies demonstrate the effectiveness of our embedding-disentanglement designs, including discovery experiment (n=4) for VIEs, inspiration experiment (n=4) for approach extensibility, and exploration experiment (n=8) for the virtual interaction system.
Songpei Xu, Xuri Ge, Chaitanya Kaul, Roderick Murray-Smith
ACM Multimedia3
2024 HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive Systems
abstract
We present a novel Hand-pose Embedding Interactive System (HpEIS) as a virtual sensor, which maps users’ flexible hand poses to a two-dimensional visual space using a Variational Autoencoder (VAE) trained on a variety of hand poses. HpEIS enables visually interpretable and guidable support for user explorations in multimedia collections, using only a camera as an external hand pose acquisition device. We identify general usability issues associated with system stability and smoothing requirements through pilot experiments with expert and inexperienced users. We then design stability and smoothing improvements, including hand-pose data augmentation, an anti-jitter regularisation term added to loss function, stabilising post-processing for movement turning points and smoothing post-processing based on One Euro Filters. In target selection experiments (n=12), we evaluate HpEIS by measures of task completion time and the final distance to target points, with and without the gesture guidance window condition. Experimental responses indicate that HpEIS provides users with a learnable, flexible, stable and smooth mid-air hand movement interaction experience.
Songpei Xu, Xuri Ge, Chaitanya Kaul, Roderick Murray-Smith
ICME3
2024 Is One GPU Enough? Pushing Image Generation at Higher-Resolutions with Foundation Models
abstract
In this work, we introduce Pixelsmith, a zero-shot text-to-image generative framework to sample images at higher resolutions with a single GPU. We are the first to show that it is possible to scale the output of a pre-trained diffusion model by a factor of 1000, opening the road to gigapixel image generation at no extra cost. Our cascading method uses the image generated at the lowest resolution as baseline to sample at higher resolutions. For the guidance, we introduce the Slider, a mechanism that fuses the overall structure contained in the first-generated image with enhanced fine details. At each inference step, we denoise patches rather than the entire latent space, minimizing memory demands so that a single GPU can handle the process, regardless of the image's resolution. Our experimental results show that this method not only achieves higher quality and diversity compared to existing techniques but also reduces sampling time and ablation artifacts.
Athanasios Tragakis, Marco Aversa, Chaitanya Kaul, Roderick Murray-Smith, Daniele Faccio
NeurIPS3
2023 Optimizing Vision Transformers for Medical Image Segmentation
abstract
For medical image semantic segmentation (MISS), Vision Transformers have emerged as strong alternatives to convolutional neural networks thanks to their inherent ability to capture long-range correlations. However, existing research uses off-the-shelf vision Transformer blocks based on linear projections and feature processing which lack spatial and local context to refine organ boundaries. Furthermore, Transformers do not generalize well on small medical imaging datasets and rely on large-scale pre-training due to limited inductive biases. To address these problems, we demonstrate the design of a compact and accurate Transformer network for MISS, CS-Unet, which introduces convolutions in a multi-stage design for hierarchically enhancing spatial and local modeling ability of Transformers. This is mainly achieved by our well-designed Convolutional Swin Transformer (CST) block which merges convolutions with Multi-Head Self-Attention and Feed-Forward Networks for providing inherent localized spatial context and inductive biases. Experiments demonstrate CS-Unet without pre-training out- performs other counterparts by large margins on multi-organ and cardiac datasets with fewer parameters and achieves state-of-the-art performance. Our code is available at Github1.
Qianying Liu, Chaitanya Kaul, Jun Wang 0121, Christos Anagnostopoulos 0001, Roderick Murray-Smith, Fani Deligianni
ICASSP2
2023 mmSense: Detecting Concealed Weapons with a Miniature Radar Sensor
abstract
For widespread adoption, public security and surveillance systems must be accurate, portable, compact, and real-time, without impeding the privacy of the individuals being observed. Current systems broadly fall into two categories – image-based which are accurate, but lack privacy, and RF signal-based, which preserve privacy but lack portability, compactness and accuracy. Our paper proposes mmSense, an end-to-end portable miniaturised real-time system that can accurately detect the presence of concealed metallic objects on persons in a discrete, privacy-preserving modality. mm-Sense features millimeter wave radar technology, provided by Google’s Soli sensor for its data acquisition, and TransDope, our real-time neural network, capable of processing a single radar data frame in 19 ms. mmSense achieves high recognition rates on a diverse set of challenging scenes while running on standard laptop hardware, demonstrating a significant advancement towards creating portable, cost-effective real-time radar based surveillance systems.
Kevin J. Mitchell, Khaled Kassem, Chaitanya Kaul, Valentin Kapitany, Philip Binner, Andrew Ramsay, Daniele Faccio, Roderick Murray-Smith
ICASSP3
2023 Continuous Interaction with A Smart Speaker via Low-Dimensional Embeddings of Dynamic Hand Pose
abstract
This paper presents a new continuous interaction strategy with visual feedback of hand pose and mid-air gesture recognition and control for a smart music speaker, which utilizes only 2 video frames to recognize gestures. Frame-based hand pose features from MediaPipe Hands, containing 21 landmarks, are embedded into a 2 dimensional pose space by an autoencoder. The corresponding space for interaction with the music content is created by embedding high-dimensional music track profiles to a compatible two-dimensional embedding. A PointNet-based model is then applied to classify gestures which are used to control the device interaction or explore music spaces. By jointly optimising the autoencoder with the classifier, we manage to learn a more useful embedding space for discriminating gestures. We demonstrate the functionality of the system with experienced users selecting different musical moods by varying their hand pose.
Songpei Xu, Chaitanya Kaul, Xuri Ge, Roderick Murray-Smith
ICASSP2
2023 The Fully Convolutional Transformer for Medical Image Segmentation
abstract
We propose a novel transformer, capable of segmenting medical images of varying modalities. Challenges posed by the fine-grained nature of medical image analysis mean that the adaptation of the transformer for their analysis is still at nascent stages. The overwhelming success of the UNet lay in its ability to appreciate the fine-grained nature of the segmentation task, an ability which existing transformer based models do not currently posses. To address this shortcoming, we propose The Fully Convolutional Transformer (FCT), which builds on the proven ability of Convolutional Neural Networks to learn effective image representations, and combines them with the ability of Transformers to effectively capture long-term dependencies in its inputs. The FCT is the first fully convolutional Transformer model in medical imaging literature. It processes its input in two stages, where first, it learns to extract long range semantic dependencies from the input image, and then learns to capture hierarchical global attributes from the features. FCT is compact, accurate and robust. Our results show that it outperforms all existing transformer architectures by large margins across multiple medical image segmentation datasets of varying data modalities without the need for any pre-training. FCT outperforms its immediate competitor on the ACDC dataset by 1.3%, on the Synapse dataset by 4.4%, on the Spleen dataset by 1.2% and on ISIC 2017 dataset by 1.1% on the dice metric, with up to five times fewer parameters. On the ACDC Post-2017-MICCAI-Challenge online test set, our model sets a new state-of-the-art on unseen MRI test cases out-performing large ensemble models as well as nnUNet with considerably fewer parameters. Our code, environments and models will be available via GitHub†.
Athanasios Tragakis, Chaitanya Kaul, Roderick Murray-Smith, Dirk Husmeier
WACV2
2023 Survey: Leakage and Privacy at Inference Time
abstract
Leakage of data from publicly available Machine Learning (ML) models is an area of growing significance since commercial and government applications of ML can draw on multiple sources of data, potentially including users' and clients' sensitive data. We provide a comprehensive survey of contemporary advances on several fronts, covering involuntary data leakage which is natural to ML models, potential malicious leakage which is caused by privacy attacks, and currently available defence mechanisms. We focus on inference-time leakage, as the most likely scenario for publicly available models. We first discuss what leakage is in the context of different data, tasks, and model architectures. We then propose a taxonomy across involuntary and malicious leakage, followed by description of currently available defences, assessment metrics, and applications. We conclude with outstanding challenges and open questions, outlining some promising directions for future research.
Marija Jegorova, Chaitanya Kaul, Charlie Mayor, Alison O'Neil, Alexander J. Weir, Roderick Murray-Smith, Sotirios A. Tsaftaris
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Prediction of weaning from mechanical ventilation using Convolutional Neural Networks
Yan Jia 0008, Chaitanya Kaul, Tom Lawton, Roderick Murray-Smith, Ibrahim Habli
Artif. Intell. Medicine2
2020 FatNet: A Feature-attentive Network for 3D Point Cloud Processing
abstract
The application of deep learning to 3D point clouds is challenging due to its lack of order. Inspired by the point embeddings of PointNet and the edge embeddings of DGCNNs, we propose three improvements to the task of point cloud analysis. First, we introduce a novel feature-attentive neural network layer, a FAT layer, that combines both global point-based features and local edge-based features in order to generate better embeddings. Second, we find that applying the same attention mechanism across two different forms of feature map aggregation, max pooling and average pooling, gives better performance than either alone. Third, we observe that residual feature reuse in this setting propagates information more effectively between the layers, and makes the network easier to train. Our architecture achieves state-of-the-art results on the task of point cloud classification, as demonstrated on the ModelNet40 dataset, and an extremely competitive performance on the ShapeNet part segmentation challenge.
Chaitanya Kaul, Nick E. Pears, Suresh Manandhar
ICPR1