VLDB 2026 Research / reviewers in the wild / expert
P. K. Srijith
dblp:120/8712 · also Srijith P. K
· DBLP profile ↗
35ranked-venue papers
7as first author
19since 2021 · last 2025
0000-0002-2820-0835ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 14 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pseudo-Inverse Prefix Tuning for Effective Unlearning in LLMsabstractLarge Language Models (LLMs) are widely used in many real-world applications, but their deployment raises concerns about data privacy and compliance with regulations such as the right to be forgotten. To address these challenges, we explore the problem of machine unlearning, selectively removing the influence of specific training data from a model. While many existing approaches require retraining the entire model and access to both forget data and retain data, we propose Pseudo-Inverse Prefix Tuning (PI-Prefix), a parameter-efficient fine-tuning method that enables targeted forgetting with minimal overhead. PI-Prefix learns a small set of prefix parameters on the data to be forgotten and then applies pseudo-inverse transformation to unlearn the forget data while maintaining performance on retain data. Our experiments on two sentiment classification tasks (SST-2 and Yelp) demonstrate that PI-Prefix achieves effective and interpretable forgetting, with forget-set performance approaching random prediction. It preserves a strong generalization on the retain set even without requiring it during unlearning. These results highlight PI-Prefix as a promising direction for scalable and compliant unlearning in data removal contexts. Preethi Gurumurthy, P. K. Srijith |
CIKM | 2 |
| 2025 | Neural Wave Equation for Irregularly Sampled Sequence DataabstractSequence labeling problems arise in several real-world applications such as healthcare and robotics. In many such applications, sequence data are irregularly sampled and are of varying complexities. Recently, efforts have been made to develop neural ODE-based architectures to model the evolution of hidden states continuously in time, to address irregularly sampled sequence data. However, they assume a fixed architectural depth and limit their flexibility to adapt to data sets with varying complexities. We propose the neural wave equation, a novel deep learning method inspired by the wave equation, to address this through continuous modeling of depth. Neural Wave Equation models the evolution of hidden states continuously across time as well as depth by using a non-homogeneous wave equation parameterized by a neural network. Through d'Alembert's analytical solution of the wave equation, we also show that the neural wave equation provides denser connections across the hidden states, allowing for better modeling capability. We conduct experiments on several sequence labeling problems involving irregularly sampled sequence data and demonstrate the superior performance of the proposed neural wave equation model. Arkaprava Majumdar, M. Anand Krishna, P. K. Srijith |
ICLR | 3 |
| 2025 | Linked Adapters: Linking Past and Future to Present for Effective Continual Learning
Dupati Srikar Chandra, P. K. Srijith, Dana Rezazadegan, Chris McCarthy |
PAKDD (1) | 2 |
| 2025 | AdaPrefix++: Integrating Adapters, Prefixes and Hypernetwork for Continual Learning
Sayanta Adhikari, Dupati Srikar Chandra, P. K. Srijith, Pankaj Wasnik, Naoyuki Onoe |
WACV | 3 |
| 2025 | EvoCL: Continual Learning over Evolving DomainsabstractContinual Learning aspires to build models capable of learning new tasks, without forgetting previously learnt tasks. In real-world settings, the distributions underlying the tasks are prone to shift. This necessitates a model capable of observing how the task distributions drift with time and adapt proactively. We present a novel framework of continual learning under evolving domains. Our approach employs a hypernetwork with separate embeddings conditioned on both domain and task to address this problem. The hypernetwork generates customised classifier weights corresponding to any domain-task pair. We employ a separate network that is trained end to end along with the hypernetwork to predict the next domain embedding, which in turn helps to generate classifier parameters corresponding to the next future domain in the evolution. We conduct extensive experiments on various datasets with a wide variety of distribution shifts to demonstrate the efficacy of our model in generalizing to future domains across all the tasks. Vishnuprasadh Kumaravelu, P. K. Srijith, Sunil Gupta 0001 |
WACV | 2 |
| 2024 | TL-CL: Task And Language Incremental Continual LearningabstractThis paper introduces and investigates the problem of Task and Language Incremental Continual Learning (TLCL), wherein a multilingual model is systematically updated to accommodate new tasks in previously learned languages or new languages for established tasks.This significant yet previously unexplored area holds substantial practical relevance as it mirrors the dynamic requirements of real-world applications.We benchmark a representative set of continual learning (CL) algorithms for TLCL.Furthermore, we propose Task and Language-Specific Adapters (TLSA), an adapter-based parameter-efficient fine-tuning strategy.TLSA facilitates cross-lingual and cross-task transfer and outperforms other parameter-efficient finetuning techniques.Crucially, TLSA reduces parameter growth stemming from saving adapters to linear complexity from polynomial complexity as it was with parameter isolation-based adapter tuning.We conducted experiments on several NLP tasks arising across several languages.We observed that TLSA outperforms all other parameter-efficient approaches without requiring access to historical data for replay. Shrey Satapara, P. K. Srijith |
EMNLP | 2 |
| 2024 | Transformer based Multitask Learning for Image Captioning and Object Detection
Debolena Basak, P. K. Srijith, Maunendra Sankar Desarkar |
PAKDD (2) | 2 |
| 2024 | Monte Carlo DropBlock for modeling uncertainty in object detection
Sai Harsha Yelleni, Deepshikha Kumari, P. K. Srijith, C. Krishna Mohan |
Pattern Recognit. | 3 |
| 2023 | Time-to-Event Modeling with Hypernetwork based Hawkes ProcessabstractMany real-world applications are associated with collection of events with timestamps, known as time-to-event data. Earthquake occurrences, social networks, and user activity logs can be represented as a sequence of discrete events observed in continuous time. Temporal point process serves as an essential tool for modeling such time-to-event data in continuous time space. Despite having massive amounts of event sequence data from various domains like social media, healthcare etc., real world application of temporal point process faces two major challenges: 1) it is not generalizable to predict events from unseen event sequences in dynamic environment 2) they are not capable of thriving in continually evolving environment with minimal supervision while retaining previously learnt knowledge. To tackle these issues, we propose HyperHawkes, a hypernetwork based temporal point process framework which is capable of modeling time of event occurrence for unseen sequences and consequently, zero-shot learning for time-to-event modeling. We also develop a hypernetwork based continually learning temporal point process for continuous modeling of time-to-event sequences with minimal forgetting. HyperHawkes augments the temporal point process with zero-shot modeling and continual learning capabilities. We demonstrate the application of the proposed framework through our experiments on real-world datasets. Our results show the efficacy of the proposed approach in terms of predicting future events under zero-shot regime for unseen event sequences. We also show that the proposed model is able to learn the time-to-event sequences continually while retaining information from previous event sequences, mitigating catastrophic forgetting in neural temporal point process. Manisha Dubey, P. K. Srijith, Maunendra Sankar Desarkar |
KDD | 2 |
| 2023 | DialoGen: Generalized Long-Range Context Representation for Dialogue Systems
Maunendra Sankar Desarkar Suvodip Dey, Asif Ekbal, P. K. Srijith |
PACLIC | 3 |
| 2023 | Continuous Depth Recurrent Neural Differential Equations
Srinivas Anumasa, Geetakrishnasai Gunapati, P. K. Srijith |
ECML/PKDD (2) | 3 |
| 2023 | Continual Learning with Dependency Preserving HypernetworksabstractHumans learn continually throughout their lifespan by accumulating diverse knowledge and fine-tuning it for future tasks. When presented with a similar goal, neural networks suffer from catastrophic forgetting if data distributions across sequential tasks are not stationary over the course of learning. An effective approach to address such continual learning (CL) problems is to use hypernetworks which generate task dependent weights for a target network. However, the continual learning performance of existing hypernetwork based approaches are affected by the assumption of independence of the weights across the layers in order to maintain parameter efficiency. To address this limitation, we propose a novel approach that uses a dependency preserving hypernetwork to generate weights for the target network while also maintaining the parameter efficiency. We propose to use recurrent neural network (RNN) based hypernetwork that can generate layer weights efficiently while allowing for dependencies across them. In addition, we propose novel regularisation and network growth techniques for the RNN based hypernetwork to further improve the continual learning performance. To demonstrate the effectiveness of the proposed methods, we conducted experiments on several image classification continual learning tasks and settings. We found that the proposed methods based on the RNN hypernetworks outperformed the baselines in all these CL settings and tasks. Dupati Srikar Chandra, Sakshi Varshney, P. K. Srijith, Sunil Gupta 0001 |
WACV | 3 |
| 2022 | Latent Time Neural Ordinary Differential EquationsabstractNeural ordinary differential equations (NODE) have been proposed as a continuous depth generalization to popular deep learning models such as Residual networks (ResNets). They provide parameter efficiency and automate the model selection process in deep learning models to some extent. However, they lack the much-required uncertainty modelling and robustness capabilities which are crucial for their use in several real-world applications such as autonomous driving and healthcare. We propose a novel and unique approach to model uncertainty in NODE by considering a distribution over the end-time T of the ODE solver. The proposed approach, latent time NODE (LT-NODE), treats T as a latent variable and apply Bayesian learning to obtain a posterior distribution over T from the data. In particular, we use variational inference to learn an approximate posterior and the model parameters. Prediction is done by considering the NODE representations from different samples of the posterior and can be done efficiently using a single forward pass. As T implicitly defines the depth of a NODE, posterior distribution over T would also help in model selection in NODE. We also propose, adaptive latent time NODE (ALT-NODE), which allow each data point to have a distinct posterior distribution over end-times. ALT-NODE uses amortized variational inference to learn an approximate posterior using inference networks. We demonstrate the effectiveness of the proposed approaches in modelling uncertainty and robustness through experiments on synthetic and several real-world image classification data. Srinivas Anumasa, P. K. Srijith |
AAAI | 2 |
| 2022 | Hawkes Process Classification through Discriminative Modeling of TextabstractSocial media such as Twitter has provided a platform for users to gather and share information and stay updated with the news. However, restriction on the length, informal grammar and vocabulary of the posts pose challenges to perform classification from textual content alone. We propose models based on the Hawkes process (HP) which can naturally incorporate additional cues such as the temporal features and past labels of the posts, along with the textual features for improving short text classification. In particular, we propose a discriminative approach to model text in HP, where the text features parameterize the base intensity and the triggering kernel of the intensity function. This allows textual content to determine influence from past posts and consequently determine the intensity function and class label. Another major contribution is to model the kernel as a neural network function of both time and text, permitting more complex influence functions for Hawkes process. This will maintain the interpretability of Hawkes process models along with the improved function learning capability of the neural networks. The proposed HP models can easily consider pretrained word embeddings to represent text for classification. Experiments on the rumour stance classification problems in social media demonstrate the effectiveness of the proposed HP models. Rohan Tondulkar, Manisha Dubey, P. K. Srijith, Michal Lukasik |
IJCNN | 3 |
| 2021 | Multi-view hypergraph convolution network for semantic annotation in LBSNsabstractSemantic characterization of the Point-of-Interest (POI) plays an important role for modeling location-based social networks and various related applications like POI recommendation, link prediction etc. However, semantic categories are not available for many POIs which makes this characterization difficult. Semantic annotation aims to predict such missing categories of POIs. Existing approaches learn a representation of POIs using graph neural networks to predict semantic categories. However, LBSNs involve complex and higher order mobility dynamics. These higher order relations can be captured effectively by employing hypergraphs. Moreover, visits to POIs can be attributed to various reasons like temporal characteristics, spatial context etc. Hence, we propose a Multi-view Hypergraph Convolution Network (Multi-HGCN) where we learn POI representations by considering multiple hypergraphs across multiple views of the data. We build a comprehensive model to learn the POI representation capturing temporal, spatial and trajectory-based patterns among POIs by employing hypergraphs. We use hypergraph convolution to learn better POI representation by using spectral properties of hypergraph. Experiments conducted on three real-world datasets show that the proposed approach outperforms the state-of-the-art approaches. Manisha Dubey, P. K. Srijith, Maunendra Sankar Desarkar |
ASONAM | 2 |
| 2021 | CAM-GAN: Continual Adaptation Modules for Generative Adversarial NetworksabstractWe present a continual learning approach for generative adversarial networks (GANs), by designing and leveraging parameter-efficient feature map transformations. Our approach is based on learning a set of global and task-specific parameters. The global parameters are fixed across tasks whereas the task-specific parameters act as local adapters for each task, and help in efficiently obtaining task-specific feature maps. Moreover, we propose an element-wise addition of residual bias in the transformed feature space, which further helps stabilize GAN training in such settings. Our approach also leverages task similarities based on the Fisher information matrix. Leveraging this knowledge from previous tasks significantly improves the model performance. In addition, the similarity measure also helps reduce the parameter growth in continual adaptation and helps to learn a compact model. In contrast to the recent approaches for continually-learned GANs, the proposed approach provides a memory-efficient way to perform effective continual data generation. Through extensive experiments on challenging and diverse datasets, we show that the feature-map-transformation approach outperforms state-of-the-art methods for continually-learned GANs, with substantially fewer parameters. The proposed method generates high-quality samples that can also improve the generative-replay-based continual learning for discriminative tasks. Sakshi Varshney, Vinay Kumar Verma, P. K. Srijith, Lawrence Carin, Piyush Rai |
NeurIPS | 3 |
| 2021 | Subset-of-data variational inference for deep Gaussian-processes regressionabstractDeep Gaussian Processes (DGPs) are multi-layer, flexible extensions of Gaussian Processes but their training remains challenging. Most existing methods for inference in DGPs use sparse approximation which require optimization over a large number of inducing inputs and their locations across layers. In this paper, we simplify the training by setting the locations to a fixed subset of data and sampling the inducing inputs from a variational distribution. This reduces the trainable parameters and computation cost without any performance degradation, as demonstrated by our empirical results on regression data sets. Our modifications simplify and stabilize DGP training methods while making them amenable to sampling schemes such as leverage score and determinantal point processes. P. K. Srijith, Mohammad Emtiyaz Khan |
UAI | 2 |
| 2021 | Traffic Incident Duration Prediction using BERT Representation of TextabstractOwing to the diverse nature of traffic incidents, accepting and storing relevant data in the form of natural language is more convenient than in constrained value fields. Textual information in such cases can be rich enough for traffic incident analysis and modelling even in the absence of certain fixed set of parameters. However limited studies considered the complexity in processing such information to predict traffic incident duration. In this paper, we propose to represent the textual data from incident reports using BERT word embeddings. These text representations are then inputted into various regressors such as LSTM, XGBoost, RF and SVR to predict traffic incident duration. To demonstrate the significance of this approach, the method is compared with the state-of-the-art approach using LDA representation. Dataset used for the experiment is the Caltrans Performance Measurement System (PeMS). Result analysis indicates that the BERT- LSTM hybrid model is effective in capturing the contextual meaning of textual incident reports to predict the traffic incident duration and outperforms LDA topic modelling with MAE around 11.16 minutes. Prashansa Agrawal, A. Antony Franklin, Digvijay S. Pawar, P. K. Srijith |
VTC Fall | 4 |
| 2021 | Improving Robustness and Uncertainty Modelling in Neural Ordinary Differential EquationsabstractDeep learning models such as Resnets have resulted in state-of-the-art accuracy in many computer vision problems. Neural ordinary differential equations (NODE) provides a continuous depth generalization of Resnets and overcome drawbacks of Resnet such as model selection and parameter complexity. Though NODE is more robust than Resnet, we find that NODE based architectures are still far away from providing robustness and uncertainty handling required for many computer vision problems. We propose novel NODE models which address these drawbacks. In particular, we propose Gaussian processes (GPs) to model the fully connected neural networks in NODE (NODE-GP) to improve robustness and uncertainty handling capabilities of NODE. The proposed model is flexible to accommodate different NODE architectures, and further improves the model selection capabilities in NODEs. We also find that numerical techniques play an important role in modelling NODE robustness, and propose to use different numerical techniques to improve NODE robustness. We demonstrate the superior robustness and uncertainty handling capabilities of proposed models on adversarial attacks and out-of-distribution experiments for the image classification tasks. Srinivas Anumasa, P. K. Srijith |
WACV | 2 |
| 2020 | HAP-SAP: Semantic Annotation in LBSNs using Latent Spatio-Temporal Hawkes ProcessabstractThe prevalence of location-based social networks (LBSNs) has eased the understanding of human mobility patterns. However, categories which act as semantic characterization of the location, might be missing for some check-ins and can adversely affect modelling the mobility dynamics of users. At the same time, mobility patterns provide cues on the missing semantic categories. In this paper, we simultaneously address the problem of semantic annotation of locations and location adoption dynamics of users. We propose our model HAP-SAP, a latent spatio-temporal multivariate Hawkes process, which considers latent semantic category influences, and temporal and spatial mobility patterns of users. The inferred semantic categories can supplement our model on predicting the next check-in events by users. Our experiments on real datasets demonstrate the effectiveness of the proposed model for the semantic annotation and location adoption modelling tasks. Manisha Dubey, P. K. Srijith, Maunendra Sankar Desarkar |
SIGSPATIAL/GIS | 2 |
| 2020 | Improving Adaptive Bayesian Optimization with Spectral Mixture Kernel
Suvodip Dey, Hiransh Gupta, P. K. Srijith |
ICONIP (5) | 4 |
| 2020 | STM-GAN: Sequentially Trained Multiple Generators for Mitigating Mode Collapse
Sakshi Varshney, P. K. Srijith, Vineeth N. Balasubramanian |
ICONIP (5) | 2 |
| 2020 | Evaluation of Deep Gaussian Processes for Text ClassificationabstractWith the tremendous success of deep learning models on computer vision tasks, there are various emerging works on the Natural Language Processing (NLP) task of Text Classification using parametric models. However, it constrains the expressability limit of the function and demands enormous empirical efforts to come up with a robust model architecture. Also, the huge parameters involved in the model causes over-fitting when dealing with small datasets. Deep Gaussian Processes (DGP) offer a Bayesian non-parametric modelling framework with strong function compositionality, and helps in overcoming these limitations. In this paper, we propose DGP models for the task of Text Classification and an empirical comparison of the performance of shallow and Deep Gaussian Process models is made. Extensive experimentation is performed on the benchmark Text Classification datasets such as TREC (Text REtrieval Conference), SST (Stanford Sentiment Treebank), MR (Movie Reviews), R8 (Reuters-8), which demonstrate the effectiveness of DGP models. P. Jayashree 0001, P. K. Srijith |
LREC | 2 |
| 2020 | Modeling Implicit Communities from Geo-Tagged Event Traces Using Spatio-Temporal Point Processes
Ankita Likhyani, P. K. Srijith, Deepak P 0001, Srikanta J. Bedathur |
WISE (1) | 3 |
| 2018 | Classification of Short-Texts Generated During Disasters: A Deep Neural Network Based ApproachabstractMicro-blogging sites provide a wealth of resources during disaster events in the form of short texts. Correct classification of these text data into various actionable classes can be of great help in shaping the means to rescue people in disaster-affected places. The process of classification of these text data poses a challenging problem because the texts are usually short and very noisy and finding good features that can distinguish these texts into different classes is time consuming, tedious and often requires a lot of domain knowledge. We propose a deep learning based model to classify tweets into different actionable classes such as resource need and availability, activities of various NGO etc. Our model requires no domain knowledge and can be used in any disaster scenario with little to no modification. Shamik Kundu, P. K. Srijith, Maunendra Sankar Desarkar |
ASONAM | 2 |
| 2018 | A Bayesian Point Process Model for User Return Time Prediction in Recommendation SystemsabstractIn order to sustain the user-base for a web service, it is important to know the return time of a user to the service. We propose a Bayesian point process, log Gaussian Cox process (LGCP), to model and predict return time of users. It allows encoding the prior domain knowledge and non-parametric estimation of latent intensity functions capturing user behaviour. We capture the similarities among the users in their return time by using a multi-task learning approach. We show the effectiveness of the proposed approaches on predicting the return time of users to last.fm music service. Sherin Thomas, P. K. Srijith, Michal Lukasik |
UMAP | 2 |
| 2017 | Longitudinal Modeling of Social Media with Hawkes Process Based on Users and NetworksabstractOnline social media provide a platform for rapid network propagation of information at an unprecedented scale. In this paper, we study the evolution of information cascades in Twitter using a point process model of user activity. Twitter is rich with heterogenous information on users and network structure. We develop several Hawkes process models considering various properties of Twitter including conversational structure, users' connections and general features of users including the textual information, and show how they are helpful in modeling the social network activity. Evaluation on Twitter data sets shows that incorporating richer properties improves the performance in predicting future activity of users and memes. P. K. Srijith, Michal Lukasik, Kalina Bontcheva, Trevor Cohn |
ASONAM | 1 |
| 2017 | Sub-story detection in Twitter with hierarchical Dirichlet processesabstractSocial media has now become the de facto information source on real world events. The challenge, however, due to the high volume and velocity nature of social media streams, is in how to follow all posts pertaining to a given event over time – a task referred to as story detection. Moreover, there are often several different stories pertaining to a given event, which we refer to as sub-stories and the corresponding task of their automatic detection – as sub-story detection. This paper proposes hierarchical Dirichlet processes (HDP), a probabilistic topic model, as an effective method for automatic sub-story detection. HDP can learn sub-topics associated with sub-stories which enables it to handle subtle variations in sub-stories. It is compared with state-of-the-art story detection approaches based on locality sensitive hashing and spectral clustering. We demonstrate the superior performance of HDP for sub-story detection on real world Twitter data sets using various evaluation measures. The ability of HDP to learn sub-topics helps it to recall the sub-stories with high precision. This has resulted in an improvement of up to 60% in the F-score performance of HDP based sub-story detection approach compared to standard story detection approaches. A similar performance improvement is also seen using an information theoretic evaluation measure proposed for the sub-story detection task. Another contribution of this paper is in demonstrating that considering the conversational structures within the Twitter stream can bring up to 200% improvement in sub-story detection performance. P. K. Srijith, Mark Hepple, Kalina Bontcheva, Daniel Preotiuc-Pietro |
Inf. Process. Manag. | 1 |
| 2016 | Studying the Temporal Dynamics of Word Co-occurrences: An Application to Event Detection
Daniel Preotiuc-Pietro, P. K. Srijith, Mark Hepple, Trevor Cohn |
LREC | 2 |
| 2016 | Gaussian Process Pseudo-Likelihood Models for Sequence Labeling
P. K. Srijith, P. Balamurugan 0001, Shirish K. Shevade |
ECML/PKDD (1) | 1 |
| 2015 | Modeling Tweet Arrival Times using Log-Gaussian Cox ProcessesabstractResearch on modeling time series text corpora has typically focused on predicting what text will come next, but less well studied is predicting when the next text event will occur.In this paper we address the latter case, framed as modeling continuous inter-arrival times under a log-Gaussian Cox process, a form of inhomogeneous Poisson process which captures the varying rate at which the tweets arrive over time.In an application to rumour modeling of tweets surrounding the 2014 Ferguson riots, we show how interarrival times between tweets can be accurately predicted, and that incorporating textual features further improves predictions. Michal Lukasik, P. K. Srijith, Trevor Cohn, Kalina Bontcheva |
EMNLP | 2 |
| 2014 | Gaussian Process Multi-task Learning Using Joint Feature Selection
P. K. Srijith, Shirish K. Shevade |
ECML/PKDD (3) | 1 |
| 2013 | Semi-supervised Gaussian Process Ordinal Regression
P. K. Srijith, Shirish K. Shevade, S. Sundararajan |
ECML/PKDD (3) | 1 |
| 2012 | Multi-Task Learning Using Shared and Task Specific Information
P. K. Srijith, Shirish K. Shevade |
ICONIP (3) | 1 |
| 2012 | Validation Based Sparse Gaussian Processes for Ordinal Regression
P. K. Srijith, Shirish K. Shevade, S. Sundararajan |
ICONIP (2) | 1 |