VLDB 2026 Research / reviewers in the wild / expert
Suprakash Datta
dblp:93/945
· DBLP profile ↗
21ranked-venue papers
4as first author
3since 2021 · last 2026
0009-0006-4698-9709ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorTheory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do multimodal LLMs understand programming screenshots? Inferring questions and extracting relevant content
Faiz Ahmed, Xuchen Tan, Folajinmi Adewole, Suprakash Datta, Maleknaz Nayebi |
Empir. Softw. Eng. | 4 |
| 2025 | Inferring Questions from Programming ScreenshotsabstractThe integration of generative AI into developer forums like Stack Overflow presents an opportunity to enhance problem-solving by allowing users to post screenshots of code or Integrated Development Environments (IDEs) instead of traditional text-based queries. This study evaluates the effectiveness of various large language models (LLMs)—specifically LLAMA, GEMINI, and GPT-4o in interpreting such visual inputs. We employ prompt engineering techniques, including in-context learning, chain-of-thought prompting, and few-shot learning, to assess each model’s responsiveness and accuracy. Our findings show that while GPT-4o shows promising capabilities, achieving over $60 \%$ similarity to baseline questions for $51.75 \%$ of the tested images, challenges remain in obtaining consistent and accurate interpretations for more complex images. This research advances our understanding of the feasibility of using generative AI for image-centric problem-solving in developer communities, highlighting both the potential benefits and current limitations of this approach while envisioning a future where visual-based debugging copilot tools become a reality. Faiz Ahmed, Xuchen Tan, Folajinmi Adewole, Suprakash Datta, Maleknaz Nayebi |
MSR | 4 |
| 2024 | Negative Results of Image Processing for Identifying Duplicate Questions on Stack OverflowabstractIn the rapidly evolving landscape of developer communities, Q&A platforms serve as crucial resources for crowdsourcing developers’ knowledge. A notable trend is the increasing use of images to convey complex queries more effectively. However, the current state-of-the-art method of duplicate question detection has not kept pace with this shift, which predominantly concentrates on text-based analysis. Inspired by advancements in image processing and numerous studies in software engineering illustrating the promising future of image-based communication on social coding platforms, we delved into image-based techniques for identifying duplicate questions on Stack Overflow. When focusing solely on text analysis of Stack Overflow questions and omitting the use of images, our automated models overlook a significant aspect of the question. Previous research has demonstrated the complementary nature of images to text. To address this, we implemented two methods of image analysis: first, integrating the text from images into the question text, and second, evaluating the images based on their visual content using image captions. After a rigorous evaluation of our model, it became evident that the efficiency improvements achieved were relatively modest, approximately an average of 1%. This marginal enhancement falls short of what could be deemed a substantial impact. As an encouraging aspect, our work lays the foundation for easy replication and hypothesis validation, allowing future research to build upon our approach and explore novel solutions for more effective image-driven duplicate question detection. Faiz Ahmed, Suprakash Datta, Maleknaz Nayebi |
ESEM | 2 |
| 2014 | Efficient and accurate sensor network localization
Tareq Adnan, Suprakash Datta, Stuart MacLean |
Pers. Ubiquitous Comput. | 2 |
| 2014 | Reducing the Positional Error of Connectivity-Based Positioning Algorithms Through Cooperation Between NeighborsabstractThe information available to connectivity-based positioning algorithms is the radio range of sensor devices and the position estimates of neighbors and neighbors of neighbors. This information creates special graph theoretic structures which impose new constraints on the positions of sensor devices. The new constraints sometimes lead to a feasible set of positions with disconnected regions. These properties can be used to reduce the set of feasible positions for a node. In this paper, a new fully distributed positioning algorithm, called Orbit, which exploits these properties is presented for mobile sensor networks. The algorithm uses additional constraints and trims disconnected regions. These new constraints are generated through cooperation between neighbors. The performance of Orbit is examined for many communication and mobility models, including a probabilistic communication model generated from radio experiments. Computer simulation experiments demonstrate that Orbit outperforms a recently proposed positioning algorithm in terms of positional accuracy under different models with a wide range of parameter values. Orbit is implemented on resource limited sensor devices. This implementation demonstrates the feasibility of the algorithm for sensor devices. The algorithm is tested on deployments of the sensor devices in a field and the results are comparable to those from the simulation experiments. Stuart MacLean, Suprakash Datta |
IEEE Trans. Mob. Comput. | 2 |
| 2013 | Evolved Features for DNA Sequence Classification and Their Fitness LandscapesabstractA key problem in genomics is the classification and annotation of sequences in a genome. A major challenge is identifying good sequence features. Evolutionary algorithms have the potential to search a large space of features and automatically generate useful ones. This paper proposes a two-stage method that generates features using multiple replicates of a genetic algorithm operating on an augmented finite state machine, called a side effect machine (SEM), and then selects a small diverse feature set using several methods, including a novel method called dissimilarity clustering. We apply our method to three problems related to transposable elements and compare the results to those usingk-mer features. We are able to produce a small set of interesting and comprehensible features that create random forest classifiers more accurate and less prone to overfitting than those created usingk-mer features. We analyze the SEM fitness landscapes and discuss the use of different fitness functions. Wendy Ashlock, Suprakash Datta |
IEEE Trans. Evol. Comput. | 2 |
| 2012 | Distinguishing Endogenous Retroviral LTRs from SINE Elements Using Features Extracted from Evolved Side Effect MachinesabstractSide effect machines produce features for classifiers that distinguish different types of DNA sequences. They have the, as yet unexploited, potential to give insight into biological features of the sequences. We introduce several innovations to the production and use of side effect machine sequence features. We compare the results of using consensus sequences and genomic sequences for training classifiers and find that more accurate results can be obtained using genomic sequences. Surprisingly, we were even able to build a classifier that distinguished consensus sequences from genomic sequences with high accuracy, suggesting that consensus sequences are not always representative of their genomic counterparts. We apply our techniques to the problem of distinguishing two types of transposable elements, solo LTRs and SINEs. Identifying these sequences is important because they affect gene expression,genome structure, and genetic diversity, and they serve as genetic markers. They are of similar length, neither codes for protein, and both have many nearly identical copies throughout the genome. Being able to efficiently and automatically distinguish them will aid efforts to improve annotations of genomes. Our approach reveals structural characteristics of the sequences of potential interest to biologists. Wendy Ashlock, Suprakash Datta |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Improving the accuracy of connectivity-based positioning for mobile sensor networksabstractIn this paper we present a new positioning algorithm, called Orbit, for mobile sensor networks. The algorithm runs in a fully distributed manner at all nodes. The information available to the algorithm is the radio range of the sensors and, the latest location estimates of neighbours and neighbours of neighbours. The algorithm has several novel aspects. First, we use geometric structures and results from unit disk graphs to impose tighter constraints on sensor positions. Second, we separate the set of possible positions of a node into connected subsets and possibly eliminate some subsets. Using simulations and data from radio experiments, we demonstrate that Orbit outperforms a recently proposed positioning algorithm WMCL-B in terms of positional accuracy for several communication and mobility models with a wide range of parameter values. Stuart MacLean, Suprakash Datta |
PIMRC | 2 |
| 2011 | The minimum positional error incurred by any connectivity-based positioning algorithm for mobile wireless systems
Stuart MacLean, Suprakash Datta |
Theor. Comput. Sci. | 2 |
| 2010 | Detecting retroviruses using reading frame information and side effect machinesabstractThis paper addresses the problem of distinguishing retroviruses from non-coding DNA sequences. Retroviruses have a distinctive reading frame structure that includes multiple reading frames that often overlap. This paper uses reading frame information generated from Fourier spectral analysis as input for Side Effect Machines (SEMs) that are evolved to create clusterings which separate the two types of sequences. The output from these SEMs is then used to train Support Vector Machines (SVMs) to perform the classification. The best classifier out of 100 replicates achieves 100% accuracy using complete retroviral genomes and the average classifier achieves 85% accuracy. Using endogenous retroviral data that includes many mutations, the best classifier achieves 86% accuracy; the average achieves an accuracy of 71%. The method also was able to distinguish lentiviruses from other types of retroviruses with a best accuracy of 100% (average 93%). In order to better understand the evolved SEMs, classifiers trained on SEMs evolved using endogenous retroviral data were used to classify the complete unmutated retroviral genomes and vice versa. It was found that, regardless of which type of data was used to create the classifiers, their performance on the test data sets was similar. This suggests that SEMs are able to extract the distinctive retroviral reading frame structure from the Fourier spectra, but that in some of the endogenous retroviruses in our data set there were too many mutations for this structure to be discernable from the data using this method. Wendy Ashlock, Suprakash Datta |
CIBCB | 2 |
| 2010 | Swift: Scalable weighted iterative sampling for flow cytometry clusteringabstractFlow cytometry (FC) is a powerful technology for rapid multivariate analysis and functional discrimination of cells. Current FC platforms generate large, high-dimensional datasets which pose a significant challenge for traditional manual bivariate analysis. Automated multivariate clustering, though highly desirable, is also stymied by the critical requirement of identifying rare populations that form rather small clusters, in addition to the computational challenges posed by the large size and dimensionality of the datasets. In this paper, we address these twin challenges by developing a two-stage scalable multivariate parametric clustering algorithm. In the first stage, we model the data as a mixture of Gaussians and use an iterative weighted sampling technique to estimate the mixture components successively in order of decreasing size. In the second stage, we apply a graph-based hierarchical merging technique to combine Gaussian components with significant overlaps into the final number of desired clusters. The resulting algorithm offers a reduction in complexity over conventional mixture modeling while simultaneously allowing for better detection of small populations. We demonstrate the effectiveness of our method both on simulated data and actual flow cytometry datasets. Iftekhar Naim, Suprakash Datta, Gaurav Sharma 0001, James S. Cavenaugh, Tim R. Mosmann |
ICASSP | 2 |
| 2010 | TCP is Competitive with Resource Augmentation
Jeff Edmonds, Suprakash Datta, Patrick W. Dymond |
Theory Comput. Syst. | 2 |
| 2009 | Device temporal forensics: An information theoretic approachabstractBy formulating the problem of ordering the outputs observed from a device over time, we pose a new problem in forensics and propose a framework for addressing this problem of device temporal forensics. Our proposed framework is based on a two-stage approach wherein time-dependent device parameters are first estimated from observed outputs and the resulting estimates are then temporally ordered by employing a Markov model for the temporal evolution of device parameters and exploiting the data processing inequality in information theory. We demonstrate and evaluate a simple realization of the framework for digital camera forensics based on photo-response non-uniformity. Results obtained over a database of online images indicate that the method provides accurate temporal ordering. Junwen Mao, Orhan Bulan, Gaurav Sharma 0001, Suprakash Datta |
ICIP | 4 |
| 2007 | Localization in wireless sensor networksabstractA fundamental problem in wireless sensor networks is localization -- the determination of the geographical locations of sensors. Most existing localization algorithms were designed to work well either in networks of static sensors or networks in which all sensors are mobile. In this paper, we propose two localization algorithms, MSL and MSL*, that work well when any number of sensors are static or mobile. MSL and MSL* are range-free algorithms -- they do not require that sensors are equipped with hardware to measure signal strengths, angles of arrival of signals or distances to other sensors. We present simulation results to demonstrate that MSL and MSL* outperform existing algorithms in terms of localization error in very different mobility conditions. MSL* outperforms MSL in most scenarios, but incurs a higher communication cost. MSL outperforms MSL* when there is significant irregularity in the radio range. We also point out some problems with a well known lower bound for the error in any range-free localization algorithm in static sensor networks. Masoomeh Rudafshani, Suprakash Datta |
IPSN | 2 |
| 2006 | A Fast Algorithm for Detecting Frame Shifts in DNA sequencesabstractSequencing technologies used to generate long strands of DNA are susceptible to laboratory errors that may result in several DNA nucleotides being deleted from the genome. Detecting such deletions in the protein coding regions is of utmost importance. Missing even a single nucleotide may lead to frame shifts with all the following codons (and consequently the encoded amino acids) being identified incorrectly. In addition to the deletion of nucleotides during sequencing, frame shifts can occur because of a variety of other reasons including mutations. In this paper, we present a fast computational technique to identify frame shifts in protein coding regions in DNA sequences. Our technique is based on Fourier spectral characteristics of coding regions in DNA sequences. We provide two applications of our technique - detecting deletions in DNA sequences in coding regions and also detecting frame shifts in viral DNA Hassan Masoom, Suprakash Datta, Amir Asif, Lesley Cunningham, Gillian Wu |
CIBCB | 2 |
| 2006 | Distributed localization in static and mobile sensor networksabstractSensor networks are expected to revolutionize information gathering, processing and dissemination in many diverse environments. In this paper, we address a fundamental problem in designing sensor networks: localization, or determining the locations of nodes. We assume that a small fraction of the sensor nodes (called seeds) know their locations. We propose an algorithm that enables other nodes to estimate their locations by exchanging information between nodes and seeds. Unlike most existing work, in our algorithm, a node uses the location information of all its neighbors, not just the seed nodes. Unlike most existing algorithms our algorithm works for both static and mobile sensor networks. Using simulation experiments, we demonstrate that our algorithm significantly outperforms comparable existing algorithms like DV-hop [1] and MCL [2] Suprakash Datta, Chris Klinowski, Masoomeh Rudafshani, Shaker Khaleque |
WiMob | 1 |
| 2006 | A Low-Maintenance Energy-Aware Clustering Algorithm for Wireless Ad-hoc NetworksabstractClustering has often been used to impose structure in wireless ad hoc networks. In this work, we propose a modified lowest-ID clustering algorithm that tries to increase the stability of the created clusters. A stability factor is associated with nodes to improve the stability of clusters produced. The stability parameter is a measure of the time that a cluster head starts its leadership role. In our algorithm, nodes use periodic beacons as the only means of communications with its neighbors. The stability parameter is defined in one of the fields of the beacons. Nodes contend to become cluster head; the node with a lower ID and larger stability factor wins the contention. Since cluster heads have extra functionality and therefore consume more energy compared to the other nodes in the network, we propose an energy efficient load balancing mechanism on the created clusters based on their energy levels. To balance the energy consumption among the nodes, a cluster head retires after some time and hands over its role to another neighbor cluster head with higher energy levels. This is useful for prolonging the network lifetime. We demonstrate using simulations that our algorithm improves the average residual energy of the network as well as the stability of the clusters produced Foroohar Foroozan, Suprakash Datta |
WiMob | 2 |
| 2005 | A fast DFT based gene prediction algorithm for identification of protein coding regionsabstractThe paper provides theoretical justification for the "3-periodicity property" observed in protein coding regions within genomic DNA sequences. We propose a new classification criteria improving upon traditional frequency based approaches for identification of coding regions. Experimental studies indicate superior performance compared with other algorithms that use the 3-periodicity property. Suprakash Datta, Amir Asif |
ICASSP (5) | 1 |
| 2003 | TCP is competitive against a limited adversaryabstractThe well-known Transport Control Protocol (TCP) is a crucial component of the TCP/IP architecture on which the Internet is built, and is a de facto standard for reliable communication on the Internet. At the heart of the TCP protocol is its congestion control algorithm. While most practitioners believe that TCP congestion control algorithm performs very well, a complete analysis of the congestion control algorithm is yet to be done. A lot of effort has, therefore, gone into the evaluation of different performance metrics like throughput and average latency under TCP. In this paper, we approach the problem from a different perspective and use the the competitive analysis framework to provide some answers to the question “how good is the TCP/IP congestion control algorithm? ” First, we prove that for networks with a single bottleneck (or point of congestion), TCP is competitive to the optimal centralized (global) algorithm in minimizing the user-perceived latency or flow time of the sessions, provided we limit the adversary by giving it strictly less resources than TCP. Specifically, we show that with O(1) times as much bandwidth and O(1) extra time per job, TCP is O(1)-competitive against an optimal global algorithm. We motivate the need for allowing TCP to have extra resources by observing that existing lower bounds for nonclairvoyant scheduling algorithms imply that no online, distributed, non-clairvoyant algorithm can be competitive with an optimal offline algorithm if both algorithms were given the same resources. Second, we show that TCP is fair by proving that it converges quickly to allocations where every session gets its fair share of network bandwidth. 1 Jeff Edmonds, Suprakash Datta, Patrick W. Dymond |
SPAA | 2 |
| 1999 | Convergence and Concentration Results for packets Routing Networks
Suprakash Datta, Ramesh K. Sitaraman |
SIROCCO | 1 |
| 1997 | The Performance of Simple Routing Algorithms That Drop PacketsabstractSeveral modern high-speed networks implement routing algorithms that resolve contention for resources such as buffer space by dropping (i.e., deleting) packets.In this paper, we analyze the performance of such routing algorithms for the commonly-used butterfly network.We assume that each switch of the butterfly has a buffer that can hold a bounded number of packets, and any packet attempting to enter a switch with a full buffer is simply dropped from the network.We study three significant metrics that characterize routing performance:expected throughput of the network, packet loss rate, and expected delay of a packet.Our main results are analytic expressions for these three performance metrics in terms of the network-size, size of the buffer at each switch, and the packet arrival rate.Our analyses for the throughput and packet loss rate hold for any non-predictive queuing protocol, including simple, often-implemented protocols such as i%st-in fist-out (FIFO) and fixed-priority scheduling.Our delay expressions hold for the FIFO protocol.Several facts of interest to a network designer fall out of our analysis.Further, our results provide quantitative insights into how the three performance metrics tradeoff against each other.Also, we present simulation results to bolster the results of our analysis.Finally, we outline preliminary results for routing on other networks such aa the crossbar."The authors are supported in part by NSF Grant CCR-94-1OO77.Permission 10 nmkc digilillhrd copIcs otall or pml Ol-lhIS m;IINI:Il Ior personal or clmsroom mse is gmntc[i u'ithoul lte pmwdtd 1]1o1 (he copies are NOImade or distrihu{ed for protit o!-comnmrci:ll :Idlmltagc.the mp,vright iwlice.(he litle o!'dw puh(ictlllon Jnd its dale JPPMI'.and notice IS given 11111 copyright is by pwmwslon ol'lhr .+4Chi.inc."1'(copyo!hmwsc.10republish, !0 posl on scnvrs or (0 redislrlhu[c 10 1]s[s,rcqultm spwllic pmnissm atd/or lee STA4 97 NwpoII, Rhode lslw)d 1ISA Copyright 1997 ACh4 0-89791-8°0-8/97/06 .$3.5(1 Suprakash Datta, Ramesh K. Sitaraman |
SPAA | 1 |