VLDB 2026 Research / reviewers in the wild / expert
Sudipta Acharya
dblp:117/5592
· DBLP profile ↗
24ranked-venue papers
12as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intent2QoS: Language Model-Driven Automation of Traffic Shaping ConfigurationsabstractTraffic shaping and Quality of Service (QoS) enforcement are critical for managing bandwidth, latency, and fairness in networks. These tasks often rely on low-level traffic control settings, which require manual setup and technical expertise. This paper presents an automated framework that converts high-level traffic shaping intents in natural or declarative language into valid and correct traffic control rules. To the best of our knowledge, we present the first end-to-end pipeline that ties intent translation in a queuing-theoretic semantic model and, with a rule-based critic, yields deployable Linux traffic control configuration sets. The framework has three steps: (1) a queuing simulation with priority scheduling and Active Queue Management (AQM) builds a semantic model; (2) a language model, using this semantic model and a traffic profile, generates sub-intents and configuration rules; and (3) a rule-based critic checks and adjusts the rules for correctness and policy compliance. We evaluate multiple language models by generating traffic control commands from business intents that comply with relevant standards for traffic control protocols. Experimental results on 100 intents show significant gains, with LLaMA3 reaching 0.88 semantic similarity and 0.87 semantic coverage, outperforming other models by over 30\. A thorough sensitivity study demonstrates that AQM-guided prompting reduces variability threefold compared to zero-shot baselines. Sudipta Acharya, Burak Kantarci |
ICC | 1 |
| 2025 | PulseAnomaly: Unsupervised Anomaly Detection on Avionic Platforms With Seasonality and Trend Modeling in Transformer NetworksabstractFor communication within military avionic platforms (e.g., F-15 and F-35), the US Department of Defense established MIL-STD-1553 military standard. It has been released for more than 50 years and is still used in platforms other than military avionics. It was originally produced to be used with military avionics, but in the following decades, it was adopted into all branches of the armed forces, as well as spacecraft and commercial avionics. However, potential attacks against the MIL-STD-1553 may exist due to the demand for internet communication between planes and the lack of security. The current study presentsPulseAnomaly, a novel unsupervised anomaly detection model for the MIL-STD-1553 bus that utilizes time-feature and message sequences. Our model demonstrates better performance compared to baseline models in the test, achieving a higher F1-score and showing excellent AUROC compared to existing methods. Additionally, we have used data from a recently developed open-source MIL-STD-1553 real-time bus simulator, which features a more diverse range of attacks and data points that more closely resemble real-world scenarios. Evaluation results show that our model outperforms existing unsupervised solutions. Hanbo Yu, Sudipta Acharya, Steven H. H. Ding, Mohammad Zulkernine |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Toward a Robust Detection of PowerShell Malware against Code Mixing and Obfuscation by Using Sentence Transformer and Similarity LearningabstractEmbedded PowerShell commands or scripts are among the most popular malware payloads. For malware that prioritizes stealthiness, such as fileless malware, PowerShell’s access to Windows API functions without additional libraries makes it useful for evading detection. Detecting malicious PowerShell scripts and commands is an open challenge for proactive endpoint protection due to three major issues: (1) The malicious commands are usually hidden in a long script beyond the processing limit of typical machine learning models. (2) They are usually mixed with bulky benign scripts. (3) Script obfuscation can easily conceal their potential matching signatures. In this article, we introduce a novel model addressing these challenges. It incorporates similarity learning, sentence transformer, sliding window method, and stochastic gradient descent (SGD) classifier. Our key insight is that malicious PowerShell code, particularly when obfuscated, exhibits semantic and statistical deviations from benign administrative usage, and these deviations can be captured by contrastive sentence embeddings without the need for de-obfuscation or handcrafted features. We operate this insight through a Siamese similarity learning framework that improves robustness against Out-of-Vocabulary tokens due to unseen code obfuscation methods. The sliding window method enables the model to handle long scripts, and the SGD classifier evaluates segment-level maliciousness. Our model achieves accuracies of 99.01%, 97.59%, 98.70%, and 99.73% across multiple obfuscated and mixed script benchmarks, outperforming existing baselines by over 30% in all cases. This work demonstrates a scalable and effective strategy for robust PowerShell malware detection in real-world scenarios. Zhiwei Fu, Leo Song, Steven H. H. Ding, Furkan Alaca, Sudipta Acharya |
ACM Trans. Priv. Secur. | 5 |
| 2023 | Variational Autoencoder with Temporal Condition for Effective Shape-based Calcium Imaging Neuron RegistrationabstractThanks to recent advances in optical imaging techniques, calcium imaging can now record the activities of thousands of neurons simultaneously, through several sessions and over long periods. Neuron registration assumes a vital role in this process, enabling the monitoring of neurons across multiple movies and extended timeframes by aligning their spatial patterns. Previous approaches often relied on clusters or probabilistic models based on simple distance metrics like overlapping pixels or center-to-center distances, neglecting crucial neuron shape and spatial relationship details. In this paper, we introduce a novel technique for cell registration. Our investigation revealed that a neuron’s shape is influenced by its temporal behaviors, leading us to suggest a temporal-conditional variational autoencoder (tcVAE) for precise shape modeling. A comprehensive evaluation demonstrates that incorporating shape-related details can significantly enhance the quality of neuron registration. Cyrus Y. H. Fung, Sudipta Acharya, Tak Pan Wong, Steven H. H. Ding |
BIBM | 2 |
| 2023 | MMCo-Clus - An Evolutionary Co-clustering Algorithm for Gene Selection (Extended abstract)abstractDimensionality reduction through feature selection becomes inevitable to overcome the problem of the Curse of dimensionality. In this article, we propose a feature (gene) selection method for high dimensional gene expression (GE) data through a Multi-objective optimization-based Multi-view Co-Clustering algorithm (named MMCo-Clus). A thorough comparative analysis with existing feature selection algorithms using external/internal evaluation metrics supports our proposed method’s potency. Laizhong Cui, Sudipta Acharya, Sumit Mishra, Yi Pan 0001, Joshua Zhexue Huang |
ICDE | 2 |
| 2023 | TDRLM: Stylometric learning for authorship verification by Topic-Debiasing
Weihan Ou, Sudipta Acharya, Steven H. H. Ding, Ryan D'Gama, Hanbo Yu |
Expert Syst. Appl. | 3 |
| 2022 | A Refined 3-in-1 Fused Protein Similarity Measure: Application in Threshold-Free Hub DetectionabstractAn exhaustive literature survey shows that finding protein/gene similarity is an important step towards solving widespread bioinformatics problems, such as predicting protein-protein interactions, analyzing Protein-Protein Interaction Networks (PPINs), gene prioritization, and disease gene/protein detection. In this article, we have proposed an improved 3-in-1 fused protein similarity measure called FuSim-II. It is built upon combining the weighted average of biological knowledge extracted from three potential genomic/ proteomic resources such as Gene Ontology (GO), PPIN, and protein sequence. Furthermore, we have shown the application of the proposed measure in detecting potential hub-proteins from a given PPIN. Aiming that, we have proposed a multi-objective clustering-based protein hub detection framework with FuSim-II working as the underlying proximity measure. The PPINs of H. Sapiens and M. Musculus organisms are chosen for experimental purposes. Unlike most of the existing hub-detection methods, the proposed technique does not require to follow any protein degree cut-off or threshold to define hubs. A thorough assessment of efficiency between proposed and existing eight protein similarity measures along with eight single/multi-objective clustering methods has been carried out. Internal cluster validity indices like Silhouette and Davies Bouldin (DB) are deployed to accomplish analytical study. Also, a comparative performance analysis between proposed and five existing hub-proteins detection algorithms is conducted through the enrichment of essentiality study. The reported results show the improved performance of FuSim-II over existing protein similarity measures in terms of identifying functionally related proteins as well as relevant hub-proteins. Supplementary material is available at http://csse.szu.edu.cn/staff/cuilz/eng/index.html. Sudipta Acharya, Laizhong Cui, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | MMCo-Clus - An Evolutionary Co-clustering Algorithm for Gene SelectionabstractIn the era of Big Data, cluster analysis of high-dimensional data sets often suffers from theCurse of dimensionality. To overcome this problem, the dimensionality reduction throughfeature selectionbecomes inevitable. Co-clustering or two-way clustering is considered to be a more sophisticated tool than conventional one-way clustering. Moreover, the advent of multi-view learning shows that the subjects of a data set can be interpreted in many ways. Interestingly, a minimal number of existing feature selection algorithms take advantage of the co-clustering method and are designed to consider multi-view data. Motivated by this, in the current article, we propose a feature (gene) selection method for high dimensional gene expression (GE) data through amulti-objective optimization basedmulti-viewCo-Clustering algorithm (namedMMCo-Clus). A popular evolutionary technique – Non-dominated Sorting Genetic Algorithm-II (NSGA-II) has been utilized as the proposed method's underlying optimization strategy. First, we construct two views of a chosen data set, utilizing knowledge from two different biological data sources. Next, we develop the MMCo-Clusalgorithm considering the constructed views to identify a set of “good” co-clustering solutions. Finally, based on a concept ofconsensus operationon the co-clustering outcome, a small number of most relevant and non-redundant features are extracted from the original feature-space. The reduced dimension formed by new feature-space causes to decrease the computational burden and noise level of original data. For experimental analysis, we have chosen three benchmark GE data sets. Our feature selection method's effectiveness is evaluated through sample-classification accuracy, accompanied by the cluster profile plot/Eisen plot/t-SNE plot, and biological/statistical significance test. A thorough comparative analysis with existing feature selection algorithms using external and internal evaluation metrics supports our proposed method's potency. Laizhong Cui, Sudipta Acharya, Sumit Mishra, Yi Pan 0001, Joshua Zhexue Huang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | A consensus multi-view multi-objective gene selection approach for improved sample classificationabstractBACKGROUND: In the field of computational biology, analyzing complex data helps to extract relevant biological information. Sample classification of gene expression data is one such popular bio-data analysis technique. However, the presence of a large number of irrelevant/redundant genes in expression data makes a sample classification algorithm working inefficiently. Feature selection is one such high-dimensionality reduction technique that helps to maximize the effectiveness of any sample classification algorithm. Recent advances in biotechnology have improved the biological data to include multi-modal or multiple views. Different 'omics' resources capture various equally important biological properties of entities. However, most of the existing feature selection methodologies are biased towards considering only one out of multiple biological resources. Consequently, some crucial aspects of available biological knowledge may get ignored, which could further improve feature selection efficiency. RESULTS: In this present work, we have proposed a Consensus Multi-View Multi-objective Clustering-based feature selection algorithm called CMVMC. Three controlled genomic and proteomic resources like gene expression, Gene Ontology (GO), and protein-protein interaction network (PPIN) are utilized to build two independent views. The concept of multi-objective consensus clustering has been applied within our proposed gene selection method to satisfy both incorporated views. Gene expression data sets of Multiple tissues and Yeast from two different organisms (Homo Sapiens and Saccharomyces cerevisiae, respectively) are chosen for experimental purposes. As the end-product of CMVMC, a reduced set of relevant and non-redundant genes are found for each chosen data set. These genes finally participate in an effective sample classification. CONCLUSIONS: The experimental study on chosen data sets shows that our proposed feature-selection method improves the sample classification accuracy and reduces the gene-space up to a significant level. In the case of Multiple Tissues data set, CMVMC reduces the number of genes (features) from 5565 to 41, with 92.73% of sample classification accuracy. For Yeast data set, the number of genes got reduced to 10 from 2884, with 95.84% sample classification accuracy. Two internal cluster validity indices - Silhouette and Davies-Bouldin (DB) and one external validity index Classification Accuracy (CA) are chosen for comparative study. Reported results are further validated through well-known biological significance test and visualization tool. Sudipta Acharya, Laizhong Cui, Yi Pan 0001 |
BMC Bioinform. | 1 |
| 2020 | Multi-view feature selection for identifying gene markers: a diversified biological data driven approachabstractBACKGROUND: In recent years, to investigate challenging bioinformatics problems, the utilization of multiple genomic and proteomic sources has become immensely popular among researchers. One such issue is feature or gene selection and identifying relevant and non-redundant marker genes from high dimensional gene expression data sets. In that context, designing an efficient feature selection algorithm exploiting knowledge from multiple potential biological resources may be an effective way to understand the spectrum of cancer or other diseases with applications in specific epidemiology for a particular population. RESULTS: In the current article, we design the feature selection and marker gene detection as a multi-view multi-objective clustering problem. Regarding that, we propose an Unsupervised Multi-View Multi-Objective clustering-based gene selection approach called UMVMO-select. Three important resources of biological data (gene ontology, protein interaction data, protein sequence) along with gene expression values are collectively utilized to design two different views. UMVMO-select aims to reduce gene space without/minimally compromising the sample classification efficiency and determines relevant and non-redundant gene markers from three cancer gene expression benchmark data sets. CONCLUSION: A thorough comparative analysis has been performed with five clustering and nine existing feature selection methods with respect to several internal and external validity metrics. Obtained results reveal the supremacy of the proposed method. Reported results are also validated through a proper biological significance test and heatmap plotting. Sudipta Acharya, Laizhong Cui, Yi Pan 0001 |
BMC Bioinform. | 1 |
| 2020 | Multi-Factored Gene-Gene Proximity Measures Exploiting Biological Knowledge Extracted from Gene Ontology: Application in Gene ClusteringabstractTo describe the cellular functions of proteins and genes, a potential dynamic vocabulary is Gene Ontology (GO), which comprises of three sub-ontologies namely, Biological-process, Cellular-component, and Molecular-function. It has several applications in the field of bioinformatics like annotating/measuring gene-gene or protein-protein semantic similarity, identifying genes/proteins by their GO annotations for disease gene and target discovery, etc. To determine semantic similarity between genes, several semantic measures have been proposed in literature, which involve information content of GO-terms, GO tree structure, or the combination of both. But, most of the existing semantic similarity measures do not consider different topological and information theoretic aspects of GO-terms collectively. Inspired by this fact, in this article, we have first proposed three novel semantic similarity/distance measures for genes covering different aspects of GO-tree. These are further implanted in the frameworks of well-known multi-objective and single-objective based clustering algorithms to determine functionally similar genes. For comparative analysis, 10 popular existing GO based semantic similarity/distance measures and tools are also considered. Experimental results on Mouse genome, Yeast, and Human genome datasets evidently demonstrate the supremacy of multi-objective clustering algorithms in association with proposed multi-factored similarity/distance measures. Clustering outcomes are further validated by conducting some biological/statistical significance tests. Supplementary information is available at https://www.iitp.ac.in/sriparna/journals.html. Sudipta Acharya, Sriparna Saha 0001, Prasanna Pradhan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Automated Hub-Protein Detection via a New Fused Similarity Measure-Based Multi-objective Clustering Framework
Sudipta Acharya, Laizhong Cui, Yi Pan 0001 |
ISBRA | 1 |
| 2019 | Bi-clustering of microarray data using a symmetry-based multi-objective optimization framework
Sudipta Acharya, Sriparna Saha 0001, Pracheta Sahoo |
Soft Comput. | 1 |
| 2018 | Fusion of stability and multi-objective optimization for solving cancer tissue classification problem
Sayantan Mitra, Sriparna Saha 0001, Sudipta Acharya |
Expert Syst. Appl. | 3 |
| 2018 | Simultaneous Clustering and Feature Weighting Using Multiobjective Optimization for Identifying Functionally Similar miRNAsabstractMicroRNAs (miRNAs) are a type of RNAs, which are responsible for monitoring the gene expression values. Recent research asserts that miRNAs form some clustering on chromosomes. The miRNAs belonging to a particular cluster are highly similar in terms of their activity and they are termed as "coregulated" miRNAs. The current paper presents an approach that simultaneously performs two tasks: i) clustering of miRNAs into different categories based on some similarity measures ii) identification of proper weight values for different time points with respect to which expression values are available. In general, a large number of expression values are available for a given miRNA data set. All these values may not be suitable to be used equally to measure the similarity between two miRNAs. In the current study, the problem of proper selection of weight values for different time points and then determining the proper partitioning from the given miRNA data set utilizing the similarity computed using the new set of weight values is formulated as an optimization problem where several cluster validity indices are optimized as the goodness measures. To that end, a multiobjective differential evolution based optimization technique is utilized. The supremacy of the proposed technique is tested on three miRNA data sets in comparison to some recent approaches in terms of some popular performance measures like Silhouette index and DB-index. The observations are further supported by statistical and biological significance tests. Supplementary information is available at https://www.iitp.ac.in/~sriparna/journals.html. Sriparna Saha 0001, Sudipta Acharya, Kavya K, Saisree Miriyala |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Unsupervised gene selection using biological knowledge : application in sample clusteringabstractBACKGROUND: Classification of biological samples of gene expression data is a basic building block in solving several problems in the field of bioinformatics like cancer and other disease diagnosis and making a proper treatment plan. One big challenge in sample classification is handling large dimensional and redundant gene expression data. To reduce the complexity of handling this high dimensional data, gene/feature selection plays a major role. RESULTS: The current paper explores the use of biological knowledge acquired from Gene Ontology database in selecting the proper subset of genes which can further participate in clustering of samples. The proposed feature selection technique is unsupervised in nature as it does not utilize any class label information in the process of gene selection. At the end, a multi-objective clustering approach is deployed to cluster the available set of samples in the reduced gene space. CONCLUSIONS: Reported results show that consideration of biological knowledge in gene selection technique not only reduces the feature space dimensionality in great extent but also improves the accuracy of sample classification. The obtained reduced gene space is validated using strong biological significance tests. In order to prove the supremacy of our proposed gene selection based sample clustering technique, a thorough comparative analysis has also been performed with state-of-the-art techniques. Sudipta Acharya, Sriparna Saha 0001, N. Nikhil |
BMC Bioinform. | 1 |
| 2016 | Automatic generation of biclusters from gene expression data using multi-objective simulated annealing approachabstractThe invention of microarray technology aids in the successful monitoring of the gene expression patterns. Biclustering is a method in which a number of co-regulated genes are identified over subset of conditions. Our aim is to detect all the non trivial biclusters having low mean squared residue(MSR) and high row variance. In this paper, we have proposed a multi-objective simulated annealing based solution framework to solve the biclustering problem from gene expression data sets. Two objective functions MSR and row-variance capturing two important properties of biclusters are optimized in parallel using the search capability of multi-objective simulated annealing based optimization technique, AMOSA. A new encoding strategy and several different search operators are defined for fast convergence of the algorithm. We have done experiment on two real-life data sets and obtained results are quantified by using several cluster validity indices. We have compared our obtained results with some state-of-the-art biclustering techniques. Pracheta Sahoo, Sudipta Acharya, Sriparna Saha 0001 |
ICPR | 2 |
| 2016 | Multi-objective Word Sense Induction Using Content and Interlink Connections
Sudipta Acharya, Asif Ekbal, Sriparna Saha 0001, Prabhakaran Santhanam, José G. Moreno 0001, Gaël Dias |
NLDB | 1 |
| 2016 | Multi-objective semi-supervised clustering of tissue samples for cancer diagnosis
Sriparna Saha 0001, Kuldeep Kaushik, Abhay Kumar Alok, Sudipta Acharya |
Soft Comput. | 4 |
| 2016 | Use of line based symmetry for developing cluster validity indices
Sudipta Acharya, Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Soft Comput. | 1 |
| 2016 | Multiobjective Simulated Annealing-Based Clustering of Tissue Samples for Cancer DiagnosisabstractIn the field of pattern recognition, the study of the gene expression profiles of different tissue samples over different experimental conditions has become feasible with the arrival of microarray-based technology. In cancer research, classification of tissue samples is necessary for cancer diagnosis, which can be done with the help of microarray technology. In this paper, we have presented a multiobjective optimization (MOO)-based clustering technique utilizing archived multiobjective simulated annealing(AMOSA) as the underlying optimization strategy for classification of tissue samples from cancer datasets. The presented clustering technique is evaluated for three open source benchmark cancer datasets [Brain tumor dataset, Adult Malignancy, and Small Round Blood Cell Tumors (SRBCT)]. In order to evaluate the quality or goodness of produced clusters, two cluster quality measures viz, adjusted rand index and classification accuracy ( % CoA) are calculated. Comparative results of the presented clustering algorithm with ten state-of-the-art existing clustering techniques are shown for three benchmark datasets. Also, we have conducted a statistical significance test called t-test to prove the superiority of our presented MOO-based clustering technique over other clustering techniques. Moreover, significant gene markers have been identified and demonstrated visually from the clustering solutions obtained. In the field of cancer subtype prediction, this study can have important impact. Sudipta Acharya, Sriparna Saha 0001, Yamini Thadisina |
IEEE J. Biomed. Health Informatics | 1 |
| 2014 | Multi-Objective Search Results Clustering
Sudipta Acharya, Sriparna Saha 0001, José G. Moreno 0001, Gaël Dias |
COLING | 1 |
| 2013 | Virtual Medical Board: A Distributed Bayesian Agent Based Approach (S)
Animesh Dutta, Sudipta Acharya, Aneesh Krishna, Swapan Bhattacharya |
SEKE | 2 |
| 2012 | Requirement Analysis and Automated Verification: A Semantic Approach
Animesh Dutta, Prajna Upadhyay, Sudipta Acharya |
SEKE | 3 |