VLDB 2026 Research / reviewers in the wild / expert
Lala Shakti Swarup Ray
dblp:339/6474
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7133-0205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPECTRA: An Efficient Spectral-Informed Neural Network for Sensor-Based Activity RecognitionabstractReal-time sensor-based applications in pervasive computing require edge-deployable models to ensure low latency, privacy, and efficient interaction. A prime example is sensor-based human activity recognition (HAR), where models must balance accuracy with stringent resource constraints. Yet many deep learning approaches treat temporal sensor signals as black-box sequences, overlooking spectral–temporal structure while demanding excessive computation. We present SPECTRA, a deployment-first, co-designed spectral–temporal architecture that integrates short-time Fourier transform (STFT) feature extraction, depthwise separable convolutions, and channel-wise self-attention to capture spectral–temporal dependencies under real edge runtime and memory constraints. A compact bidirectional GRU with attention pooling summarizes within-window dynamics at low cost, reducing downstream model burden while preserving accuracy. Across five public HAR datasets, SPECTRA matches or approaches larger CNN/LSTM/Transformer baselines while substantially reducing parameters, latency, and energy. Deployments on a Google Pixel 9 smartphone and an STM32L4 microcontroller further demonstrate end-to-end deployable real-time, private, and efficient HAR. Deepika Gurung, Lala Shakti Swarup Ray, Mengxi Liu 0004, Bo Zhou 0005, Paul Lukowicz |
PerCom | 2 |
| 2025 | OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language ModelsabstractUnderstanding human-to-human interactions, especially in contexts like public security surveillance, is critical for monitoring and maintaining safety. Traditional activity recognition systems are limited by fixed vocabularies, predefined labels, and rigid interaction categories that often rely on choreographed videos and overlook concurrent interactive groups. These limitations make such systems less adaptable to real-world scenarios, where interactions are diverse and unpredictable. In this paper, we propose an open vocabulary human-to-human interaction recognition (OV-HHIR) framework that leverages large language models to generate open-ended textual descriptions of both seen and unseen human interactions in open-world settings without being confined to a fixed vocabulary. Additionally, we create a comprehensive, large-scale human-to-human interaction dataset by standardizing and combining existing public human interaction datasets into a unified benchmark. Extensive experiments demonstrate that our method outperforms traditional fixed-vocabulary classification systems and existing cross-modal language models for video understanding, setting the stage for more intelligent and adaptable visual understanding systems in surveillance and beyond. Lala Shakti Swarup Ray, Bo Zhou 0005, Sungho Suh, Paul Lukowicz |
ICASSP | 1 |
| 2025 | MuJo: Multimodal Joint Feature Space Learning for Human Activity RecognitionabstractHuman activity recognition (HAR) is a long-standing problem in artificial intelligence with applications in a broad range of areas, including healthcare, sports and fitness, security, and more. The performance of HAR in real-world settings is strongly dependent on the type and quality of the input signal that can be acquired. Given an unobstructed, high-quality camera view of a scene, computer vision systems, in particular in conjunction with foundation models, can today fairly reliably distinguish complex activities. On the other hand, recognition using modalities such as wearable sensors (which are often more broadly available, e.g., in mobile phones and smartwatches) is a more difficult problem, as the signals often contain less information and labeled training data is more difficult to acquire. To alleviate the need for labeled data, we introduce our comprehensive Fitness Multimodal Activity Dataset (FiMAD) in this work, which can be used with the proposed pre-training method MuJo (Multimodal Joint Feature Space Learning) to enhance HAR performance across various modalities. FiMAD was created using YouTube fitness videos and contains parallel video, language, pose, and simulated IMU sensor data. MuJo utilizes this dataset to learn a joint feature space for these modalities. We show that classifiers pre-trained on FiMAD can increase the performance on real HAR datasets such as MM-Fit, MyoGym, MotionSense, and MHEALTH. For instance, on MM-Fit, we achieve a Macro F1-Score of up to 0.855 when fine-tuning on only 2% of the training data and 0.942 when utilizing the complete training set for classification tasks. We compare our approach with other self-supervised ones and show that, unlike them, ours consistently improves compared to the baseline network performance while also providing better data efficiency. Stefan Fritsch, Cennet Oguz, Vitor F. Rey, Lala Shakti Swarup Ray, Maximilian Kiefer-Emmanouilidis, Paul Lukowicz |
PerCom | 4 |
| 2025 | ChairPose: Pressure-based Chair Morphology Grounded Sitting Pose Estimation through Simulation-Assisted Training
Lala Shakti Swarup Ray, Vitor F. Rey, Bo Zhou 0005, Paul Lukowicz, Sungho Suh |
UIST | 1 |
| 2025 | Contrastive-representation IMU-based fitness activity recognition enhanced by bio-impedance sensing
Mengxi Liu 0004, Vitor F. Rey, Lala Shakti Swarup Ray, Bo Zhou 0005, Paul Lukowicz |
Pervasive Mob. Comput. | 3 |
| 2024 | A Novel Local-Global Feature Fusion Framework for Body-Weight Exercise Recognition with Pressure Mapping SensorsabstractWe present a novel local-global feature fusion framework for body-weight exercise recognition with floor-based dynamic pressure maps. One step further from the existing studies using deep neural networks mainly focusing on global feature extraction, the proposed framework aims to combine local and global features using image processing techniques and the YOLO object detection to localize pressure profiles from different body parts and consider physical constraints. The proposed local feature extraction method generates two sets of high-level local features consisting of cropped pressure mapping and numerical features such as angular orientation, location on the mat, and pressure area. In addition, we adopt a knowledge distillation for regularization to preserve the knowledge of the global feature extraction and improve the performance of the exercise recognition. Our experimental results demonstrate a notable 11 percent improvement in F1 score compared to the baseline 3DCNN for exercise recognition while preserving label-specific features. Davinder Pal Singh, Lala Shakti Swarup Ray, Bo Zhou 0005, Sungho Suh, Paul Lukowicz |
ICASSP | 2 |
| 2024 | ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-Based Human Activity Recognition
Lala Shakti Swarup Ray, Daniel Geißler, Mengxi Liu 0004, Bo Zhou 0005, Sungho Suh, Paul Lukowicz |
ICPR (29) | 1 |
| 2024 | A Synthetic Benchmarking Pipeline to Compare Camera Calibration Algorithms
Lala Shakti Swarup Ray, Bo Zhou 0005, Lars Krupp, Sungho Suh, Paul Lukowicz |
ICPR (32) | 1 |
| 2024 | iMove: Exploring Bio-Impedance Sensing for Fitness Activity RecognitionabstractAutomatic and precise fitness activity recognition can be beneficial in aspects from promoting a healthy lifestyle to personalized preventative healthcare. While IMUs are currently the prominent fitness tracking modality, through iMove, we show bio-impedence can help improve IMU-based fitness tracking through sensor fusion and contrastive learning. To evaluate our methods, we conducted an experiment including six upper body fitness activities performed by ten subjects over five days to collect synchronized data from bio-impedance across two wrists and IMU on the left wrist. The contrastive learning framework uses the two modalities to train a better IMU-only classification model, where bio-impedance is only required at the training phase, by which the average Macro F1 score with the input of a single IMU was improved by 3.22 % reaching 84.71 % compared to the 81.49 % of the IMU baseline model. We have also shown how bio-impedance can improve human activity recognition (HAR) directly through sensor fusion, reaching an average Macro F1 score of 89.57 % (two modalities required for both training and inference) even if Bio-impedance alone has an average macro F1 score of 75.36 %, which is outperformed by IMU alone. In addition, similar results were obtained in an extended study on lower body fitness activity classification, demonstrating the generalisability of our approach.Our findings underscore the potential of sensor fusion and contrastive learning as valuable tools for advancing fitness activity recognition, with bio-impedance playing a pivotal role in augmenting the capabilities of IMU-based systems. Mengxi Liu 0004, Vitor F. Rey, Yu Zhang 0171, Lala Shakti Swarup Ray, Bo Zhou 0005, Paul Lukowicz |
PerCom | 4 |
| 2023 | ClothFit: Cloth-Human-Attribute Guided Virtual Try-on Network Using 3D Simulated DatasetabstractOnline clothing shopping has become increasingly popular, but the high rate of returns due to size and fit issues has remained a major challenge. To address this problem, virtual try-on systems have been developed to provide customers with a more realistic and personalized way to try on clothing. In this paper, we propose a novel virtual try-on method called ClothFit, which can predict the draping shape of a garment on a target body based on the actual size of the garment and human attributes. Unlike existing try-on models, ClothFit considers the actual body proportions of the person and available cloth sizes for clothing virtualization, making it more appropriate for current online apparel outlets. The proposed method utilizes a U-Net-based network architecture that incorporates cloth and human attributes to guide the realistic virtual try-on synthesis. Specifically, we extract features from a cloth image using an auto-encoder and combine them with features from the user’s height, weight, and cloth size. The features are concatenated with the features from the U-Net encoder, and the U-Net decoder synthesizes the final virtual try-on image. Our experimental results demonstrate that ClothFit can significantly improve the existing state-of-the-art methods in terms of photo-realistic virtual try-on results. Yunmin Cho, Lala Shakti Swarup Ray, Kundan Sai Prabhu Thota, Sungho Suh, Paul Lukowicz |
ICIP | 2 |