Shradha Maharjan

dblp:423/8551 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 SYNC: SYnergistic aNnotation Collaboration between Humans and LLMs for Enhanced Model Training
abstract
Large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing tasks, highlighting their potential as effective data annotators. While LLM-generated annotations tend to be costeffective, they are often error-prone and may inadvertently introduce bias. It is advantageous to harness the strengths of both LLMs and humans to ensure higher accuracy and reliability in the annotation process. In this paper, we present a multi-step, human-LLM collaborative approach to optimize data annotation for Stack Overflow datasets. We begin by applying TF-IDF to rank and prioritize relevant elements. Next, we utilize NLP Transformer and UniXcoder to leverage their deep contextual understanding for handling the code-related queries and discussions typical of Stack Overflow, resulting in more consistent automated labels. Finally, human annotators re-annotate to correct potential errors and mitigate bias introduced during earlier stages. To support human-LLM collaboration, we developed SYNC a research prototype that implements SYnergistic aNnotation Collaboration through an intuitive graphical user interface, enabling real-time interaction between human annotators and the LLM, allowing for iterative refinements. Overall, our approach is designed to integrate automated efficiency with human oversight to improve annotation outcomes, particularly for complex or domain-specific tasks such as those found in Stack Overflow datasets.
Tammy Le, Will Taylor, Shradha Maharjan, Myoungkyu Song
SERA3
2025 Intelligent Code Completion by a Unified Multi-task Learning with a Large Language Model
abstract
Code completion has become an essential tool in modern software development. It helps developers by predicting the next token (e.g., an API function call) based on the current coding context. Its widespread use highlights the need for efficient and context-aware solutions that streamline the development process. Despite ongoing efforts to enhance code completion performance, many existing approaches remain limited, offering ranked lists primarily based on alphabetical order or usage frequency from partially typed code fragments. While these studies have seen incremental improvements, the level of meaningful assistance provided to developers has not advanced in parallel. To address this limitation, we propose CODECOM, a deep learning-based code completion technique that leverages a large language model (LLM) with multi-task learning. Our approach processes sequences of source code tokens, their corresponding abstract syntax trees (ASTs), and program dependencies, enabling more context-aware and accurate code predictions. In our case study, CODECOM demonstrates state-of-the-art performance in the code completion downstream task. The evaluation results highlight significant improvements over the baseline, achieving 29.49% in BLEU, 71.16% in Acc@1, and 67.19% in Acc@5. These advancements are further validated by the Wilcoxon signed-rank test, which confirms strong statistical significance across all metrics. These findings indicate that Code Com can accelerate software development and assist developers in reducing potential errors effectively. Index Terms-software maintenance, automated code completion, deep learning, large language model.
Shradha Maharjan, Tae-Hyuk Ahn, Myoungkyu Song
SERA1
2025 Automated Code Summarization by Training Large Language Models with Crowdsourced Knowledge
abstract
In modern software development, efficient program comprehension is essential for maintaining and evolving software systems. Developers often dedicate over $50 \%$ of their time to understanding code due to the complexity and time demands involved. Code summarization-generating concise natural language descriptions of source code-has emerged as a potential solution. However, existing automated summarization techniques frequently produce summaries that are incomplete or lack accuracy. Moreover, as software evolves, documentation often becomes outdated, leading to discrepancies between code and comments. To address these challenges, we present DeepKnowCode, an automated approach that utilizes a DEEP learning technique based on a large language model and crowdsourced KNOWledge for CODE summarization. This approach aids developers by generating summaries that elucidate (1) the internal behavior of the code, (2) the rationale behind its implementation, and (3) practical guidelines for its use. We implemented a research prototype to assess real-world applicability and rigorously evaluated DeepKnowCode against state-of-the-art approaches. In our evaluation, DeepKnowCode demonstrates performance improvements in BLEU scores of 38.8% and 39.2% over two baselines, with statistical validation underscoring its effectiveness. By incorporating crowdsourced knowledge, DeepKnowCode captures essential elements of code semantics, context, and patterns, enhancing its ability to produce accurate and contextually relevant code summaries that facilitate program comprehension.
Shradha Maharjan, Myoungkyu Song
SERA2