VLDB 2026 Research / reviewers in the wild / expert
Soumya Sen 0001
dblp:57/3852-1
· DBLP profile ↗
17ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-9178-6410ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 3Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Interpretable Rule-Based Framework with Language Model Support for Placement Prediction and Skill Enhancement
Anjan Dutta 0002, Bhargav Prasad Das, Soumya Sen 0001, Punyasha Chatterjee, Subhankar Dhar |
COMPSAC | 3 |
| 2025 | Lagged Co-movement Prediction of Sectoral Indices in Stock Market using Frequent Itemset MiningabstractStock price prediction has become a critical area of interest for investors and market analysts, though forecasting stock market trends remains a challenging endeavor due to the inherent volatility and unpredictability of the market. The process of stock price prediction typically involves estimating future prices based on historical data, market trends, and various socioeconomic factors. However, factors like market fluctuations, incomplete or erroneous data, and investor behavior add complexity to these predictions. Several methods are employed for stock price forecasting, including fundamental analysis, technical analysis, and machine learning approaches such as Linear Regression, Random Forest, and Long Short-Term Memory (LSTM) networks. This study focuses on using sectoral indices as benchmarking tools to evaluate sector performance. Specifically, it explores the co-movements of thirteen NSE sectoral indices, with one index chosen as the target. The analysis centers on using closing prices to measure sector performance and calculates the correlations between the target index and others. The six most highly correlated indices are identified, and association rule mining is used to uncover the relationships between these indices and the target index. The study aims to: (i) examine the interdependencies between the target sector and other sectors, and (ii) generate predictive rules for a sector’s performance based on the behavior of correlated sectors, providing valuable insights for making informed investment decisions. Anjan Dutta 0002, Giridhar Maji, Partha Ghosh, Punyasha Chatterjee, Takaaki Goto, Soumya Sen 0001 |
SERA | 6 |
| 2024 | A Machine Learning Based Automated Model for Managing Student DropoutabstractAddressing the persistent challenge of student dropout, particularly prevalent in developing countries like India, Bangladesh, etc. are of paramount importance. Factors such as poverty, natural calamities, and early marriages exacerbate this issue. High student dropout rates can negatively impact a country by diminishing its economic productivity, increasing social inequalities, and perpetuating a cycle of poverty. Addressing dropout issues requires comprehensive strategies to ensure a skilled and educated workforce, fostering societal well-being and global competitiveness. This research focuses on analysing comprehensive data on students who have dropped out. Thereafter, a machine learning based methodology is used to discern the underlying causes of student attrition in various schools. Furthermore, it allows for efficient monitoring of the state's educational landscape, with the ability to drill down to granular levels when necessary to identify specific regional challenges. The effectiveness of this approach is validated through the utilization of real-world datasets. Partha Ghosh, Arnab Charit, Hindol Banerjee, Debanwesa Bandhu, Agniv Ghosh, Ankita Pal, Takaaki Goto, Soumya Sen 0001 |
SERA | 8 |
| 2024 | Sectoral recommendation system for medium term investment using technical indicators
Giridhar Maji, Soumya Sen 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Efficient OLAP query processing across cuboids in distributed data warehousing environment
Saikat Raj, Tamal Chakraborty, Anirban Chakrabarty, Agostino Cortesi, Soumya Sen 0001 |
Expert Syst. Appl. | 6 |
| 2023 | Scientific Organization of Blood Donation Camp Through Lexicographic Optimization and Taxicab Path ComputationabstractBlood is the indispensable circulating fluid for sustaining human life. On demand supply of quality blood is a big challenge for every government in all developing countries. Specially, in festive seasons and winter, supplying quality blood on time is a big medical challenge. On the other hand, the consequences of mismanaged blood donation camp may lead to excess supply of human blood units. Also, in some cases, it is being noticed that human blood units are getting corrupted in transit from the blood donation camp to the blood bank. Hence, several units of human blood are getting spoiled over the time due to mismanagement and/or maintenance. In this research, we have applied a lexicographic optimization based model for finding best available blood bank from the point of blood donation camp. Alternative taxicab geometry based paths are used for finding best possible shortest path from the blood donation camp to the blood bank. Partha Ghosh, Takaaki Goto, Leena Jana Ghosh, Soumya Sen 0001 |
SERA | 4 |
| 2023 | Customer Segmentation Using Credit Card Data AnalysisabstractCustomer segmentation is a separation of a market into multiple distinct groups of consumers who share the similar characteristics. Segmentation of market is an effective way to define and meet Customer needs and also to identify the future business plan. Unsupervised machine learning algorithms are suitable to analyze and identify the possible set of customers when the labeled data about the customers are no available. In this research work the spending of different customers who have credit cards are analyzed to segment them into different clusters and also to plan further business improvements based on the different characteristics of these identified clusters. Saikat Raj, Surajit Jana, Soumyadip Roy, Takaaki Goto, Soumya Sen 0001 |
SERA | 6 |
| 2023 | Identification of City Hotspots by Analyzing Telecom Call Detail Records Using Complex Network Modeling
Giridhar Maji, Sharmistha Mandal, Soumya Sen 0001 |
Expert Syst. Appl. | 3 |
| 2021 | Identifying and ranking super spreaders in real world complex networks without influence overlap
Giridhar Maji, Animesh Dutta, Mariana Curado Malta, Soumya Sen 0001 |
Expert Syst. Appl. | 4 |
| 2021 | Cover independent image steganography in spatial domain using higher order pixel bits
Giridhar Maji, Sharmistha Mandal, Soumya Sen 0001 |
Multim. Tools Appl. | 3 |
| 2020 | A systematic survey on influential spreaders identification in complex networks with a focus on K-shell based techniques
Giridhar Maji, Sharmistha Mandal, Soumya Sen 0001 |
Expert Syst. Appl. | 3 |
| 2020 | Dual Image-Based Dictionary Encoded Data Hiding in Spatial DomainabstractIn the modern digital era, the privacy of personal communication is a serious concern to all netizens. A better way to preserve privacy could be to hide the secret message inside some innocent looking digital object such as image, audio, video, etc., which is known as steganography. A new steganographic scheme using a reference image along with the cover image has been proposed in this article. It enhances the robustness and security by increasing the obscurity of the hidden message. It also employs an additional dictionary-based encoding module to increase the hiding capacity as well as security. Experiments show that bit changes in the reference image are very few and undetectable to human perception, it also evades common statistical tests. Evaluation of standard quality parameters such as MSE, PSNR, UIQI, SSIM along with chi-squared statistics based embedding probability testing has been performed. When dictionary-based encoding is applied it further improves the quality parameters. Giridhar Maji, Sharmistha Mandal, Soumya Sen 0001 |
Int. J. Inf. Secur. Priv. | 3 |
| 2019 | Pixel Value Difference Based Image Steganography with One Time Pad EncryptionabstractPixel value differencing (PVD) and Least Significant Bit (LSB) embedding are well known spatial domain steganographic techniques. PVD utilizes the sharp changes of intensities among adjacent pixels where a large number of secret bits could be embedded without any perceptible change. one-time pad (OTP) symmetric cryptography is known for its security. In the proposed scheme one or more LSB bits of the selected pixels are used depending on the pixel intensity difference with neighboring pixels in 2 × 2 image blocks of the cover image. Secret bits are encrypted using OTP with randomly generated pre-shared key. Such encrypted bits are completely random and resemble noise hence make the scheme robust against different statistical attacks. Comparative simulations with some well-known PVD-based techniques show good results in terms of visual imperceptibility and different quality metrics such as MSE, PSNR, SSIM etc. Giridhar Maji, Sharmistha Mandal, Narayan C. Debnath, Soumya Sen 0001 |
INDIN | 4 |
| 2019 | AFARTICA: A Frequent Item-Set Mining Method Using Artificial Cell Division AlgorithmabstractFrequent item-set mining has been exhaustively studied in the last decade. Several successful approaches have been made to identify the maximal frequent item-sets from a set of typical item-sets. The present work has introduced a novel pruning mechanism which has proved itself to be significant time efficient. The novel technique is based on the Artificial Cell Division (ACD) algorithm which has been found to be highly successful in solving tasks that involve a multi-way search of the search space. The necessity conditions of the ACD process have been modified accordingly to tackle the pruning procedure. The proposed algorithm has been compared with the apriori algorithm implemented in WEKA. Accurate experimental evaluation has been conducted and the experimental results have proved the superiority of AFARTICA over apriori algorithm. The results have also indicated that the proposed algorithm can lead to better performance when the support threshold value is more for the same set of item-sets. Saubhik Paladhi, Sankhadeep Chatterjee, Takaaki Goto, Soumya Sen 0001 |
J. Database Manag. | 4 |
| 2018 | Data Warehouse Based Analysis with Integrated Blood Donation Management SystemabstractBlood donation is an important issue throughout the world to save and manage lives. In order to make this process more effective an automated system could be built to monitor and organize blood donation camps. In this paper we developed an integrated framework with all related but isolated web based sub systems of a blood management system. We propose a data warehouse (DW) as an integral part of the integrated framework to store historical blood donation data in a centralized database for analytical processing. Proposed system would enable the authorities to take informed blood donation camping decision based on the analytical reports from the DW for some area for a particular time and citizen demography. Finally, we introduce a new measure of humanity (scoring system for good deeds) of citizens called Philanthropy Score (PS) and Philanthropy League (PL) derived from PS. In our proposed system after blood collection at donation camps, few health related details would be updated in the blood management system database and most importantly PS of the donor would be updated into the national citizens’ database. A well-advertised communication would allow the citizens to know about the PS points they can accrue for every possible good deed. Blood donation is merely a prototype use of the proposed PS. Giridhar Maji, Narayan C. Debnath, Soumya Sen 0001 |
INDIN | 3 |
| 2017 | Water quality prediction: Multi objective genetic algorithm coupled artificial neural network based approachabstractDomestic and industrial pollutions affected the water quality to a greater extent. Polluted water became a major reason behind several community diseases, mainly in undeveloped and developing countries. The public health condition is deteriorating and putting an extra burden of countermeasures to prevent such water borne diseases from spreading. Detecting the drinking water quality can prevent such scenarios prior to the critical stage. Recent research works have achieved reasonable success in predicting the water quality. However, the accuracy levels of already proposed models are to be improved, keeping in mind the sensitivity of the problem domain. In the current work, multi-objective genetic algorithm was employed to train the artificial neural network (NN-MOGA) to improve its performance over its traditional counterparts. The proposed model gradually minimizes two different objective functions; namely the root mean square error (RMSE) and Maximum Error in order to find the optimal weight vector for the artificial neural network (ANN). The proposed model was compared with three other, well established models namely NN-GA (ANN trained with Genetic Algorithm), NN-PSO (ANN trained with Particle Swarm Optimization) and SVM in terms of accuracy, precision, recall, F-Measure, Matthews correlation coefficient (MCC) and Fowlkes-Mallows index (FM index). The simulation results established superior accuracy of NN-MOGA over the other models. Sankhadeep Chatterjee, Sarbartha Sarkar, Nilanjan Dey, Soumya Sen 0001, Takaaki Goto, Narayan C. Debnath |
INDIN | 4 |
| 2013 | Dynamic query path selection from lattice of cuboids using memory hierarchyabstractData warehouse represents multi-dimensional data suitable for analytical processing and logically data are organized in the form of data cube or cuboid. Data warehouse actually represents a business theme which is called fact table. The cuboid that identifies the complete fact table is called base cuboid. The all possible combination of the cuboids that could be generated from base cuboid corresponds to lattice structure. A lattice consists of numbers of cuboids. In real life, all these cuboids may not be important for business analysis. Thus all of them are not always called during business processing. The cuboids that are referred in different applications are fetched from diverse memory hierarchy such as cache memory, primary memory and secondary memory. The different execution speed of the respective memory element is taken into account which forms a memory hierarchy. The focus of this research work is to dynamically identify the most cost effective path within the lattice structure of cuboids to minimize the query access time having the knowledge of existing cuboid location at different memory elements. Soumya Sen 0001, Anirban Sarkar 0002, Nabendu Chaki, Narayan C. Debnath |
ISCC | 2 |