Suranjan Chakraborty

dblp:13/6135 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0002-2357-6099ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Data-Constrained File Fragment Classification Across Heterogeneous File Types using Large Language Models
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC4
2026 On-demand generation of high-quality software engineering datasets using large language models and ontologies
abstract
Recent advances in generative artificial intelligence (AI) and machine learning (ML) have renewed interest in realizing the long-standing goal of computer-aided software engineering by improving software quality and productivity. Although these techniques have been applied across many software engineering (SE) tasks, their effectiveness depends heavily on access to large, high-quality, labeled, domain-specific datasets, which remain limited, particularly in requirements engineering (RE) where research often relies on natural language artifacts. Existing, public datasets are typically small, contain labeling ambiguities, and show substantial class imbalance, which restricts the development, evaluation, and reproducibility of AI-driven SE approaches. To address these challenges, this paper presents the O3DG approach, a repeatable method for generating on-demand, high-quality, ontology-aligned datasets using large language models (LLMs). O3DG integrates prompt engineering strategies, domain-specific seed examples, and ML-based validation to synthesize diverse and cohesive datasets suitable for SE research. The approach is demonstrated through two representative RE case studies involving the classification of non-functional requirements and the detection of ambiguity in software requirements. For each case, the paper details the O3DG pipeline, ontology mappings, and validation steps that ensure dataset reliability and practical utility. Results show that O3DG produces datasets with strong category cohesion, improved balance across classes, and effective support for ML training. More broadly, the study illustrates how LLM-assisted dataset synthesis can help overcome persistent data limitations and provides a transferable process for producing high-quality datasets across additional SE domains.
George Bishop, Suranjan Chakraborty, Honghe Zhou, Josh Dehlinger, Lin Deng 0001, Jonah Lin, Benjamin Kist
Autom. Softw. Eng.2
2025 Forensic Intelligence Graphs: An LLM Approach to Digital Evidence Extraction and Relationship Analysis
abstract
Digital forensics often requires deriving meaningful and investigative intelligence from vast amounts of evidence scattered across various artifacts. In this study, we propose an automated approach to gain insights about criminal incidents using digital evidence networks constructed with the aid of Large Language Models (LLMs). Our method utilizes LLMs to extract evidence entities from mobile devices and infers relationships among them. Using this information, the model enables the generation of Forensic Intelligence Graphs (FIGs). These graphs visually represent evidence entities and their interrelations, providing an intelligence-driven approach to forensic data analysis. Using evidence extracted from Android mobile devices, an empirical evaluation demonstrates that the LLM-aided FIG achieves 93.33% coverage of evidence entities and 86.96% coverage of evidence relationships, effectively uncovering all relevant suspect scenarios. Moreover, our approach uncovered 27 additional evidence entities and 83 relationships beyond those recorded in the official documentation, highlighting its ability to reveal previously overlooked forensic artifacts.
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC4
2023 Reconstructing Android User Behavior through Timestamped State Models
Honghe Zhou, Phuong Dinh Nguyen, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
COMPSAC6
2023 Experimental Evaluation of Adversarial Attacks Against Natural Language Machine Learning Models
abstract
Machine learning models are being increasingly relied on for many natural language processing tasks. However, these models are vulnerable to adversarial attacks, i.e., inputs designed to target models into making a wrong prediction. Among different methods of attacking a model, it is important to understand what attacks are effective, so that we can design countermeasures to protect the models. In this paper, we design and implement six adversarial attacks against natural language machine learning models. Then, we evaluate the effectiveness of these attacks using a fine-tuned distilled BERT model and 5,000 sample sentences from the SST-2 dataset. Our results indicate that the Word-replace attack affected the model the most, which reduces the F1-score of the model by 34%. The Word-delete attack is the least effective, but still reduces the model’s accuracy by 17%. Based on the experimental results, we discuss our insights and provide our recommendations for building robust natural language machine learning models.
Jonathan Li 0007, Steven Pugh, Honghe Zhou, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
SERA6
2022 Towards Internet of Things (IoT) Forensics Analysis on Intelligent Robot Vacuum Systems
abstract
With the rapid advancement of information tech-nology, the Internet of Things (IoT) has significantly impacted people's daily life. IoT devices not only bring comfort and convenience to every aspect of the world, but also appear to be a new target of cybercrimes. Thus, IoT forensics becomes a critical step in forensics investigation. Intelligent robot vacuums are one of the most popular IoT devices. As robot vacuums can connect to the Internet and be operated through mobile apps, a large amount of data may be stored and transmitted among the vacuums, mobile apps, and the network. The data may include the history of the robot's operation, network and user credentials, and layouts of the floor plan of a house. From the perspective of digital forensics, these data can be critical while collecting necessary evidence, investigating suspects and victims, and reconstructing crime scenes. To this end, this paper makes an initial attempt to conduct a digital forensic analysis on intelligent robot vacuum systems. Specifically, this paper retrieves and analyzes a robot vacuum's operation log, the installation details of the robot vacuum's control system, and the usage record of the application from the memory of a smartphone.
Honghe Zhou, Lin Deng 0001, Wei Yu 0002, Josh Dehlinger, Suranjan Chakraborty
SERA6
2021 Automatic Identification of Vulnerable Code: Investigations with an AST-Based Neural Network
abstract
The increasing complexity of software applications and the necessity for minimizing software vulnerabilities has given rise to the use of machine learning techniques that can identify software vulnerabilities in source code. However, many of these techniques lack the accuracy needed for industrial practice. The contribution of this work is the novel use of an Abstract Syntax Tree Neural Network (ASTNN) to identify and classify software vulnerabilities in the Common Weakness Enumeration (CWE) types. We make two fundamental claims in this work. First, the use of an ASTNN performs better than prior machine learning neural network architectures. Second, the benchmark data set commonly used for machine learning vulnerability classification is flawed for this use. To illustrate these claims, we describe our ASTNN architecture and evaluate it with more than 44,000 test cases across 29 CWEs in the NIST Juliet Test Suite data set. Results show a minimum of 88% accuracy across all CWEs.
Garrett Partenza, Trevor Amburgey, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
COMPSAC5
2020 Information Technology and organizational innovation: Harmonious information technology affordance and courage-based actualization
Sutirtha Chatterjee, Greg D. Moody, Paul Benjamin Lowry, Suranjan Chakraborty, Andrew M. Hardin
J. Strateg. Inf. Syst.4
2019 Automatic Multi-class Non-Functional Software Requirements Classification Using Neural Networks
abstract
Advances in machine learning (ML) algorithms, graphics processing units, and readily available ML libraries have enabled the application of ML to open software engineering challenges. Yet, the use of ML to enable decision-making during the software engineering lifecycle is not well understood as there are various ML models requiring parameter tuning. In this paper, we leverage ML techniques to develop an effective approach to classify software requirements. Specifically, we investigate the design and application of two types of neural network models, an artificial neural network (ANN) and a convolutional neural network (CNN), to classify non-functional requirements (NFRs) into the following five categories: maintainability, operability, performance, security and usability. We illustrate and experimentally evaluate this work through two widely used datasets consisting of nearly 1,000 NFRs. Our results indicate that our CNN model can effectively classify NFRs by achieving precision ranging between 82% and 94%, recall ranging between 76% and 97% with an F-score ranging between 82% and 92%.
Cody Baker, Lin Deng 0001, Suranjan Chakraborty, Josh Dehlinger
COMPSAC (2)3
2011 Offshore Vendors' Software Development Team Configurations: An Exploratory Study
abstract
This research uses configuration theory and data collected from a major IT vendor organization to examine primary configurations of distributed teams in a global off-shoring context. The study indicates that off-shoring vendor organizations typically deploy three different types of configurations, which the authors term as thin-at-client, thick-at-client, and hybrid. These configurations differ in terms of the size of the sub-teams in the different distributed locations and the nature of the ISD-related tasks performed by the distributed team members. In addition, the different configurations were compared on their inherent process-related and resource-related flexibilities. The thick-at-client configuration emerged as the one that offers superior flexibility (in all dimensions).However, additional analysis also revealed contingencies apart from flexibility that may influence the appropriateness of the distributed ISD team configuration, including the volatility of the client organization’s environment and the extent to which the ISD tasks can be effortlessly moved to the vendor’s home location.
Suranjan Chakraborty, Saonee Sarker, Sudhanshu Rai, Suprateek Sarker, Ranganadhan Nadadhur
J. Glob. Inf. Manag.1
2009 Examining the success factors for mobile work in healthcare: A deductive study
Sutirtha Chatterjee, Suranjan Chakraborty, Saonee Sarker, Suprateek Sarker, Francis Y. Lau
Decis. Support Syst.2
2009 Assessing the relative contribution of the facets of agility to distributed systems development success: an Analytic Hierarchy Process approach
abstract
Recent studies have sought to identify different types/facets of agility that can potentially contribute to distributed Information Systems Development (ISD) project success. However, prior research has not attempted to assess the relative importance of the various types of agility with respect to different ISD success measures. We believe that such an assessment is critical, since this information can enable organizations to direct scarce organizational resources to the types of agility that are most relevant. To this end, we use the Analytic Hierarchy Process to unearth, from the perspectives of two stakeholder groups of distributed software development projects, managers, and technical staff members, as to which agility facets facilitate (and to what degree) on-time completion of projects and effective collaboration in distributed ISD teams. Furthermore, noting that there is a need for an overall set of prioritized agility facets (by integrating managerial and technically oriented perspectives), we present three ways to aggregate the preferences of the two groups.
Saonee Sarker, Charles L. Munson, Suprateek Sarker, Suranjan Chakraborty
Eur. J. Inf. Syst.4