Dakshak Keerthi Chandra

dblp:230/4705 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Reliable retrieval-augmented feature generation with large language model reasoning
abstract
Abstract Feature generation can significantly enhance learning outcomes, particularly for tasks with limited data. An effective way to improve feature generation is to expand the current feature space using existing features and enriching the informational content. However, generating new, interpretable features usually requires domain-specific knowledge on top of the existing features. In this paper, we introduce a Retrieval-Augmented Feature Generation method, RAFG, to generate useful and explainable features specific to domain classification tasks. To increase the interpretability of the generated features, we conduct knowledge retrieval among the existing features in the domain to identify potential feature associations. These associations are expected to help generate useful features. Moreover, we develop a framework based on large language models (LLMs) for feature generation with reasoning to evaluate their semantic relevance, causal alignment, and expected utility for the downstream task. To mitigate the risk of overconfident or unsupported reasoning, we further introduce a counterfactual validation mechanism that compares reasoning-based predictions with observed performance changes. Experiments across several datasets in medical, economic, and geographic domains show that our RAFG method can produce high-quality, meaningful features and significantly improve classification performance compared with baseline methods.
Jinghan Zhang 0002, Fengran Mo, Dakshak Keerthi Chandra, Yu-Zhong Chen, Kunpeng Liu 0001
Knowl. Inf. Syst.4
2025 Retrieval-Augmented Feature Generation for Domain-Specific Classification
abstract
Feature generation can significantly enhance learning outcomes, particularly for tasks with limited data. An effective way to improve feature generation is to expand the current feature space using existing features and enriching the informational content. However, generating new, interpretable features usually requires domain-specific knowledge on top of the existing features. In this paper, we introduce a Retrieval-Augmented Feature Generation method, RAFG, to generate useful and explainable features specific to domain classification tasks. To increase the interpretability of the generated features, we conduct knowledge retrieval among the existing features in the domain to identify potential feature associations. These associations are expected to help generate useful features. Moreover, we develop a framework based on large language models (LLMs) for feature generation with reasoning to verify the quality of the features during their generation process. Experiments across several datasets in medical, economic, and geographic domains show that our RAFG method can produce high-quality, meaningful features and significantly improve classification performance compared with baseline methods.
Jinghan Zhang 0002, Fengran Mo, Dakshak Keerthi Chandra, Yu-Zhong Chen
ICDM4
2021 NodeSense2Vec: Spatiotemporal Context-Aware Network Embedding for Heterogeneous Urban Mobility Data
abstract
The problem of learning latent representations of heterogeneous networks with spatial and temporal attributes has been gaining traction in recent years, given its myriad of real-world applications. Most systems with applications in the field of transportation, urban economics, medical information, online e-commerce, etc., handle big data that can be structured into Spatiotemporal Heterogeneous Networks (SHNs), thereby making efficient analysis of these networks extremely vital.In this paper, we propose a spatiotemporal context-aware network embedding framework that jointly captures the spatial regularities between objects and the sequential transition patterns of human mobility. First, we model the heterogeneous urban mobility data collected from multiple sources as an SHN using a probabilistic weighted degree centrality measure. To learn the sequential transition patterns of human mobility in urban regions, we perform meta-path constrained random walks (MPCRWs) on the constructed SHN, which captures the proximities between multi-typed objects via their rich spatiotemporal links. By treating the generated meta-path instances as sentences, we capture multiple contrastive context senses associated with nodes in an SHN produced due to multiplex of spatial and temporal dependencies between objects in urban mobility data by performing spectral graph clustering. We then map the learned contrastive contextual node senses with respective meta-path instances. Finally, we learn latent embeddings of the mapped meta-path instances by using the word2vec model Skip-gram. We evaluate the performance of our proposed model on real-world application problems. Experimental results demonstrate the effectiveness of our model over state-of-the-art alternatives.
Dakshak Keerthi Chandra, Jennifer L. Leopold, Yanjie Fu
IEEE BigData1
2020 Collective Embedding with Feature Importance: A Unified Approach for Spatiotemporal Network Embedding
abstract
In the last decade, there has been great progress in the field of machine learning and deep learning. These models have been instrumental in addressing a great number of problems. However, they have struggled when it comes to dealing with high dimensional data. In recent years, representation learning models have proven to be quite efficient in addressing this problem as they are capable of capturing effective lower-dimensional representations of the data. However, most of the existing models are quite ineffective when it comes to dealing with high dimensional spatiotemporal data as they encapsulate complex spatial and temporal relationships that exist among real-world objects. High-dimensional spatiotemporal data of cities represent urban communities. By learning their social structure we can better quantitatively depict them and understand factors influencing rapid growth, expansion, and changes.
Dakshak Keerthi Chandra, Pengyang Wang, Jennifer L. Leopold, Yanjie Fu
CIKM1
2019 Collective Representation Learning on Spatiotemporal Heterogeneous Information Networks
abstract
Representation learning is a technique that is used to capture the underlying latent features of complex data. Representation learning on networks has been widely implemented for learning network structure and embedding it in a low dimensional vector space. In recent years, network embedding using representation learning has attracted increasing attention, and many deep architectures have been widely proposed. However, existing network embedding techniques ignore the multi-class spatial and temporal relationships that crucially reflect the complex nature among vertices and links in spatiotemporal heterogeneous information networks(SHINs).
Dakshak Keerthi Chandra, Pengyang Wang, Jennifer L. Leopold, Yanjie Fu
SIGSPATIAL/GIS1