Jayesh Rajkumar Vachhani

dblp:289/1065 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-0267-4474ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ContextGraph: Lifelog Intelligence Framework for Contextual Subgraph Evolution
abstract
Lifelogging involves the continuous and comprehensive recording of a user’s daily activities, behaviors, and interactions, offering valuable insights for personalized healthcare, event retrieval, and lifestyle analysis. However, extracting meaningful patterns from lifelog data requires models to capture deeper temporal contexts beyond simple retrieval. To address this, we introduce ContextGraph, a lifelog intelligence framework that models lifelogs as a Temporal Knowledge Graph (TKG) to reason about the user’s evolving life patterns over time. ContextGraph computes Day Context Embeddings (DCE) to encode the temporal spread and social scene context of user's daily behavior. Then a novel Lens module extracts semantically meaningful subgraph snapshots around an anchor node in the TKG, representing specific personal contexts in the user’s life. The Lens module also computes an evolution signature for each subgraph, indicating whether it is growing, decaying, or remaining static. By analyzing these evolution signatures, ContextGraph provides actionable insights into the user’s lifelogs such as stable routines, behavioral drifts, or lifestyle changes. Our experiments showcase DCE's versatility, outperforming baselines in graph/node classification and reasoning on the Enzyme and DBLP datasets.
Anil Sharma, Gunturi Venkata Sai Phani Kiran, Jayesh Rajkumar Vachhani, Sourabh Vasant Gothe, Ayon Chattopadhyay, Yashwant Saini, Parameswaranath Vadackupurath Mani, Barath Raj Kandur Raja
AAAI3
2025 PhysID: Physics-based Interactive Dynamics from a Single-view Image
abstract
Transforming static images into interactive experiences remains a challenging task in computer vision. Tackling this challenge holds the potential to elevate mobile user experiences, notably through interactive and AR/VR applications. Current approaches aim to achieve this either using pre-recorded video responses or requiring multi-view images as input. In this paper, we present PhysID, that streamlines the creation of physics-based interactive dynamics from a single-view image by leveraging large generative models for 3D mesh generation and physical property prediction. This significantly reduces the expertise required for engineering-intensive tasks like 3D modeling and intrinsic property calibration, enabling the process to be scaled with minimal manual intervention. We integrate an on-device physics-based engine for physically plausible real-time rendering with user interactions. PhysID represents a leap forward in mobile-based interactive dynamics, offering real-time, non-deterministic interactions and user-personalization with efficient on-device memory consumption. Experiments evaluate the zero-shot capabilities of various Multimodal Large Language Models (MLLMs) on diverse tasks and the performance of 3D reconstruction models. These results demonstrate the cohesive functioning of all modules within the end-to-end framework, contributing to its effectiveness.
Sourabh Vasant Gothe, Ayon Chattopadhyay, Gunturi Venkata Sai Phani Kiran, Pratik, Vibhav Agarwal, Jayesh Rajkumar Vachhani, Sourav Ghosh 0001, Parameswaranath VM, Barath Raj Kandur Raja
ICASSP6
2024 SAM-GEBD: Zero-Cost Approach for Generic Event Boundary Detection
abstract
Generic Event Boundary Detection (GEBD) [1] is a crucial task in video analysis, aiming to identify class-agnostic event boundaries. Traditional supervised or unsupervised methods for GEBD rely on expensive data annotation and time-consuming training, often leading to limited generalization across diverse data distributions. In this paper, we introduce SAM-GEBD, a novel, zero-cost approach for GEBD in videos by leveraging the Segment Anything Model (SAM). While SAM has shown its impressive zero-shot capabilities across many domains and tasks, we repurposed it to address the challenge of GEBD. The proposed method involves two stages, a zero-cost method for computing temporal residual Self Similarity Matrix (SSM), and an algorithm for identifying event boundaries by decoding SSM. Our method exhibits superior performance, achieving an [email protected] score of 0.724 on the Kinetics-GEBD and 0.38 on TAPOS, surpassing the current state-of-the-art unsupervised techniques [2], [1]. Additionally, we assess SAM-GEBD’s individual components by integrating them with neural methods to demonstrate their versatility.
Pranay Kashyap, Sourabh Vasant Gothe, Vibhav Agarwal, Jayesh Rajkumar Vachhani
ICASSP4
2024 What's in the Flow? Exploiting Temporal Motion Cues for Unsupervised Generic Event Boundary Detection
abstract
Generic Event Boundary Detection (GEBD) task aims to recognize generic, taxonomy-free boundaries that segment a video into meaningful events. Current methods typically involve a neural model trained on a large volume of data, demanding substantial computational power and storage space. We explore two pivotal questions pertaining to GEBD: Can non-parametric algorithms outperform unsupervised neural methods? Does motion information alone suffice for high performance? This inquiry drives us to algorithmically harness motion cues for identifying generic event boundaries in videos. In this work, we propose FlowGEBD, a non-parametric, unsupervised technique for GEBD. Our approach entails two algorithms utilizing optical flow: (i) Pixel Tracking and (ii) Flow Normalization. By conducting thorough experimentation on the challenging Kinetics-GEBD and TAPOS datasets, our results establish FlowGEBD as the new state-of-the-art (SOTA) among unsupervised methods. FlowGEBD exceeds the neural models on the Kinetics-GEBD dataset by obtaining an [email protected] score of 0.713 with an absolute gain of 31.7% compared to the unsupervised baseline and achieves an average F1 score of 0.623 on the TAPOS validation dataset.
Sourabh Vasant Gothe, Vibhav Agarwal, Sourav Ghosh 0001, Jayesh Rajkumar Vachhani, Pranay Kashyap, Barath Raj Kandur Raja
WACV4
2023 Self-Similarity is all You Need for Fast and Light-Weight Generic Event Boundary Detection
abstract
The self-similarity matrix (SSM) is becoming more prevalent in temporal representation understanding; it has been utilized for various video understanding tasks, such as classifying human actions, counting repetitions, and identifying generic event boundaries. Recently proposed methods for Generic Event Boundary Detection (GEBD) [1] based on SSM have obtained impressive results on the Kinetics-GEBD dataset. However, they demand a large model size and an immense number of computations to achieve good performance, making them challenging to realize on edge devices. We introduce a projected SSM with cosine distance that produces an efficient representation of SSM that can be interpreted using lighter transformer decoders. This paper presents a lightweight novel architecture with just 3M trainable parameters that utilize projected SSM to solve GEBD. In addition to the low computation regime of the model, the experiments demonstrate that the architecture is invariant to the feature extractor model while inferencing. The proposed method achieves a boost of 13.92% on the Kinetics-GEBD validation dataset with 3.5X fewer model parameters and 19.5X fewer multi-add operations compared to the baseline [1]. We report competitive F1 results at unprecedented efficiency with 22X fewer model parameters than state-of-the-art methods and achieve the lowest inference time on GPU and mobile device.
Sourabh Vasant Gothe, Jayesh Rajkumar Vachhani, Rishabh Khurana, Pranay Kashyap
ICASSP2
2023 Repetition Counting from Compressed Videos Using Sparse Residual Similarity
abstract
It is common for modern video codecs to reach triple digit compression ratios, which clearly shows the information redundancy and low information density of the ubiquitous RGB frame video representation. We propose an approach that directly utilizes the components of a compressed video for predicting the count of a repeating action occurring in the video. This complete bypassing of the video decoding step offers significant computational benefits. Furthermore, by leveraging intelligent single I-frame encodings and the sparse nature of accumulated residual vectors, we are able to efficiently capture the frame features even with lightweight feature extraction backbones. On the Countix dataset, our method achieves a considerable 91.5% reduction in model size and 91% reduction in FLOPS, with competitive results compared to the state-of-the-art.
Rishabh Khurana, Jayesh Rajkumar Vachhani, Sourabh Vasant Gothe, Pranay Kashyap
ICASSP2
2021 Fontnet: On-Device Font Understanding and Prediction Pipeline
abstract
Fonts are one of the most basic and core design concepts. Numerous use cases can benefit from an in depth understanding of Fonts such as Text Customization which can change text in an image while maintaining the Font attributes like style, color, size. Currently, Text recognition solutions can group recognized text based on line breaks or paragraph breaks, if the Font attributes are known multiple text blocks can be combined based on context in a meaningful manner. In this paper, we propose two engines: Font Detection Engine, which identifies the font style, color and size attributes of text in an image and a Font Prediction Engine, which predicts similar fonts for a query font. Major contributions of this paper are three-fold: First, we developed a novel CNN architecture for identifying font style of text in images. Second, we designed a novel algorithm for predicting similar fonts for a given query font. Third, we have optimized and deployed the entire engine On-Device which ensures privacy and improves latency in real time applications such as instant messaging. We achieve a worst case On-Device inference time of 30ms and a model size of 4.5MB for both the engines.
S. Rakshith, Rishabh Khurana, Vibhav Agarwal, Jayesh Rajkumar Vachhani, Bhanodai Guggilla
ICASSP4