VLDB 2026 Research / reviewers in the wild / expert
Rishab Sharma
dblp:234/7726
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning Low-Rank Latent Spaces with Simple Deterministic Autoencoder: Theoretical and Empirical InsightsabstractThe autoencoder is an unsupervised learning paradigm that aims to create a compact latent representation of data by minimizing the reconstruction loss. However, it tends to overlook the fact that most data (images) are embedded in a lower-dimensional latent space, which is crucial for effective data representation. To address this limitation, we propose a novel approach called Low-Rank Autoencoder (LoRAE). In LoRAE, we incorporated a low-rank regularizer to adaptively learn a low-dimensional latent space while preserving the basic objective of an autoencoder. This helps embed the data in a lower-dimensional latent space while preserving important information. It is a simple autoencoder extension that learns low-rank latent space. Theoretically, we establish a tighter error bound for our model. Empirically, our model’s superiority shines through various tasks such as image generation and downstream classification. Both theoretical and practical outcomes highlight the importance of acquiring low-dimensional embeddings. Alokendu Mazumder, Tirthajit Baruah, Bhartendu Kumar, Rishab Sharma, Vishwajeet Pattanaik, Punit Rathore |
WACV | 4 |
| 2022 | An exploratory study on code attention in BERTabstractMany recent models in software engineering introduced deep neural models based on the Transformer architecture or use transformer-based Pre-trained Language Models (PLM) trained on code. Although these models achieve the state of the arts results in many downstream tasks such as code summarization and bug detection, they are based on Transformer and PLM, which are mainly studied in the Natural Language Processing (NLP) field. The current studies rely on the reasoning and practices from NLP for these models in code, despite the differences between natural languages and programming languages. There is also limited literature on explaining how code is modeled. Rishab Sharma, Fuxiang Chen, Fatemeh Hendijani Fard, David Lo 0001 |
ICPC | 1 |
| 2022 | LAMNER: code comment generation using character language model and named entity recognitionabstractCode comment generation is the task of generating a high-level natural language description for a given code method/function. Although researchers have been studying multiple ways to generate code comments automatically, previous work mainly considers representing a code token in its entirety semantics form only (e.g., a language model is used to learn the semantics of a code token), and additional code properties such as the tree structure of a code are included as an auxiliary input to the model. There are two limitations: 1) Learning the code token in its entirety form may not be able to capture information succinctly in source code, and 2) The code token does not contain additional syntactic information, inherently important in programming languages. Rishab Sharma, Fuxiang Chen, Fatemeh Hendijani Fard |
ICPC | 1 |
| 2022 | Self-admitted technical debt in R: detection and causesabstractAbstract Self-Admitted Technical Debt (SATD) is primarily studied in Object-Oriented (OO) languages and traditionally commercial software. However, scientific software coded in dynamically-typed languages such as R differs in paradigm, and the source code comments’ semantics are different (i.e., more aligned with algorithms and statistics when compared to traditional software). Additionally, many Software Engineering topics are understudied in scientific software development, with SATD detection remaining a challenge for this domain. This gap adds complexity since prior works determined SATD in scientific software does not adjust to many of the keywords identified for OO SATD, possibly hindering its automated detection. Therefore, we investigated how classification models (traditional machine learning, deep neural networks, and deep neural Pre-Trained Language Models (PTMs)) automatically detect SATD in R packages. This study aims to study the capabilities of these models to classify different TD types in this domain and manually analyze the causes of each in a representative sample. Our results show that PTMs (i.e., RoBERTa) outperform other models and work well when the number of comments labelled as a particular SATD type has low occurrences. We also found that some SATD types are more challenging to detect. We manually identified sixteen causes, including eight new causes detected by our study. The most common cause was failure to remember, in agreement with previous studies. These findings will help the R package authors automatically identify SATD in their source code and improve their code quality. In the future, checklists for R developers can also be developed by scientific communities such as rOpenSci to guarantee a higher quality of packages before submission. Rishab Sharma, Ramin Shahbazi, Fatemeh Hendijani Fard, Zadia Codabux, Melina C. Vidoni |
Autom. Softw. Eng. | 1 |
| 2021 | Retrieval Enhanced Ensemble Model Framework For Rumor Detection On Micro-blogging PlatformsabstractAutomatic rumor detection is the task of finding rumors on social networks. Previous techniques leveraged the propagation structure of tweets to detect the rumors, which makes the propagation of tweets necessary to detect rumors. However, current text-based works provide sub-optimal results as compared to propagation-based techniques. This work presents a retrieval-based framework that leverages the similar tweets from the given train set and chooses the best model from an ensemble of models to predict the test tweet label. Our proposed framework is based on transformers-based pre-trained models (PTM's). Experiments on two public data sets used in previous works, show that our framework can detect the tweets with equivalent accuracy as propagation-based techniques. The primary advantage of this work is in early rumor detection. The proposed framework can detect rumors in few minutes compared to propagation-based works, which requires a significant amount of propagation of tweets that can take hours before they can be detected. Rishab Sharma, Fatemeh Hendijani Fard, Apurva Narayan |
ICMLA | 1 |
| 2021 | API2Com: On the Improvement of Automatically Generated Code Comments Using API DocumentationsabstractCode comments can help in program comprehension and are considered as important artifacts to help developers in software maintenance. However, the comments are mostly missing or are outdated, specially in complex software projects. As a result, several automatic comment generation models are developed as a solution. The recent models explore the integration of external knowledge resources such as Unified Modeling Language class diagrams to improve the generated comments. In this paper, we propose API2Com, a model that leverages the Application Programming Interface Documentations (API Docs) as a knowledge resource for comment generation. The API Docs include the description of the methods in more details and therefore, can provide better context in the generated comments. The API Docs are used along with the code snippets and Abstract Syntax Trees in our model.We apply the model on a large Java dataset of over 130,000 methods and evaluate it using both Transformer and RNN- base architectures. Interestingly, when API Docs are used, the performance increase is negligible. We therefore run different experiments to reason about the results. For methods that only contain one API, adding API Docs improves the results by 4% BLEU score on average (BLEU score is an automatic evaluation metric used in machine translation). However, as the number of APIs that are used in a method increases, the performance of the model in generating comments decreases due to long documentations used in the input. Our results confirm that the API Docs can be useful in generating better comments, but, new techniques are required to identify the most informative ones in a method rather than using all documentations simultaneously. Ramin Shahbazi, Rishab Sharma, Fatemeh Hendijani Fard |
ICPC | 2 |
| 2020 | A Semantic-Based Framework for Analyzing App Users' FeedbackabstractThe competitive market of mobile apps requires app developers to consider the users' feedback frequently. This feedback, when comes from different resources, e.g. App Stores and Twitter, will provide a broader picture of the state of the app, as the users discuss different topics on each platform. Automated tools are developed to filter the informative comments for app developers. However, to integrate the feedbacks from different platforms, one should evaluate the similarities and/or differences of the text from each one. Different meaning of the words in various context, makes this evaluation a challenging task for automated processes. For example, Request night theme and Add dark mode are two comments that are requesting the same feature. This similarity cannot be identified automatically if the semantics of the words are not embedded in the analysis. In this paper, we propose a new framework to analyze the users' feedback by embedding their semantics. As a case study, we investigate whether our approach can identify the similar/different comments from Google Play Store and Twitter, in the two well studied classes of bug reports and feature requests from literature. The initial results, validated by expert evaluation and statistical analysis, shows that this framework can automatically measure the semantic differences among users' comments in both groups. The framework can be used to build intelligent tools to integrate the users' feedback from other platforms, as well as providing ways to analyze the reviews in more detail automatically. Aman Yadav, Rishab Sharma, Fatemeh Hendijani Fard |
SANER | 2 |