EDBT 2026 Demo / reviewers in the wild / expert
Partha Pakray
dblp:27/7928
· DBLP profile ↗
27ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-3834-5154ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning as Supportive Context for Machine Translation: A Case Study on Hindi to Bengali Language PairabstractWe investigate whether reasoning information can enhance machine translation when incorporated as supportive context during training and inference. Using Hindi-Bengali translation as a case study, we define five reasoning components: Key Terms, Syntactic, Semantic, Pragmatic, and Paraphrase. We conduct a complete ablation across all 31 possible combinations using Gemma-3-1B-Instruct and evaluate on multi-domain benchmark with BLEU, chrF, and TER. Evaluation results show that reasoning effectiveness depends on its type and composition rather than quantity. Combining multiple heterogeneous signals causes objective diffusion, degrading performance. The compact Semantic and Paraphrase combination proves optimal, and providing it during inference yields 23.86 BLEU compared to 22.12 from standard fine-tuning a +1.74 BLEU gain across eight domains. These findings demonstrate that targeted semantic guidance consistently and meaningfully improves the compact translation models. Kshetrimayum Boynao Singh, Saksham Singh, Partha Pakray, Asif Ekbal |
EAMT (2) | 3 |
| 2026 | Quantum recurrent neural network for sequential labeling
Shyambabu Pandey, Partha Pakray, Fabio Massimo Zanzotto |
Knowl. Inf. Syst. | 2 |
| 2026 | Synergizing linguistic features and transformer networks for detecting AI-generated text
Annepaka Yadagiri, Pratik Kumar, Yashraj Poddar, Partha Pakray, Chukhu Chunka |
Knowl. Inf. Syst. | 4 |
| 2025 | A deep dive into automated sexism detection using fine-tuned deep learning and large language models
Advaitha Vetagiri, Partha Pakray, Amitava Das 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Large language models: a survey of their development, capabilities, and applications
Annepaka Yadagiri, Partha Pakray |
Knowl. Inf. Syst. | 2 |
| 2025 | Extractive single document summarization using multi-objective modified cat swarm optimization approach: ESDS-MCSO
Dipanwita Debnath, Ranjita Das, Partha Pakray |
Neural Comput. Appl. | 3 |
| 2024 | A binary grey wolf optimizer to solve the scientific document summarization problem
Ranjita Das, Dipanwita Debnath, Partha Pakray, Naga Chaitanya Kumar |
Multim. Tools Appl. | 3 |
| 2024 | Text summary evaluation based on interpretable semantic textual similarity
Goutam Majumder, Vikrant Rajput, Partha Pakray, Sivaji Bandyopadhyay, Benoît Favre |
Multim. Tools Appl. | 3 |
| 2024 | MizBERT: A Mizo BERT ModelabstractThis research investigates the utilization of pre-trained BERT transformers within the context of the Mizo language. BERT, an abbreviation for Bidirectional Encoder Representations from Transformers, symbolizes Google’s forefront neural network approach to Natural Language Processing (NLP), renowned for its remarkable performance across various NLP tasks. However, its efficacy in handling low-resource languages such as Mizo remains largely unexplored. In this study, we introduce MizBERT , a specialized Mizo language model. Through extensive pre-training on a corpus collected from diverse online platforms, MizBERT has been tailored to accommodate the nuances of the Mizo language. Evaluation of MizBERT’s capabilities is conducted using two primary metrics: masked language modeling and perplexity, yielding scores of 76.12% and 3.2565, respectively. Additionally, its performance in a text classification task is examined. Results indicate that MizBERT outperforms both the Multilingual BERT model and the Support Vector Machine algorithm, achieving an accuracy of 98.92%. This underscores MizBERT’s proficiency in understanding and processing the intricacies inherent in the Mizo language. Robert Lalramhluna, Sandeep Kumar Dash, Partha Pakray |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Single document text summarization addressed with a cat swarm optimization approach
Dipanwita Debnath, Ranjita Das, Partha Pakray |
Appl. Intell. | 3 |
| 2023 | English-Assamese neural machine translation using prior alignment and pre-trained language model
Sahinur Rahman Laskar, Bishwaraj Paul, Pankaj Dadure, Riyanka Manna, Partha Pakray, Sivaji Bandyopadhyay |
Comput. Speech Lang. | 5 |
| 2023 | An extract-then-abstract based method to generate disaster-news headlines using a DNN extractor followed by a transformer abstractor
Sumanta Banerjee, Shyamapada Mukherjee, Sivaji Bandyopadhyay, Partha Pakray |
Inf. Process. Manag. | 4 |
| 2022 | A formula embedding approach for semantic similarity and relatedness between formulasabstractSummary The mathematical formula is one of the most vital components in a scientific document, which can explicitly describe various complex concepts and ideas. In addition to numerical calculations, they are also used to clarify definitions and disambiguate explanations transcribed in natural language. Nevertheless, the formulas have a noteworthy impact in the scientific documents, the existing information retrieval systems have limited access to scientific documents based on formulas‐based queries. To accomplish this, in this research, we have studied and implemented the formula embedding approach, which encodes the formula into the embedded vector. For encoding the formula, we have used pretrained sentence bidirectional encoder representations from transformers model. The proposed embedding model takes the latex formula as input and generates an upshot as a fixed dimensional embedding representation. In addition to this, the Siamese network is used to reform the semantic meaning of the formulas. Furthermore, the embedding of the formulas and the queried formula are compared, and cosine similarity is estimated. The performance of the suggested methodology is verified using a math stack exchange corpus of ARQMath 2020, and obtained results have shown a remarkable contribution in the task of formula retrieval. Pankaj Dadure, Partha Pakray, Sivaji Bandyopadhyay |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Part-of-Speech (POS) Tagging Using Deep Learning-Based Approaches on the Designed Khasi POS CorpusabstractPart-of-speech (POS) tagging is one of the research challenging fields in natural language processing (NLP). It requires good knowledge of a particular language with large amounts of data or corpora for feature engineering, which can lead to achieving a good performance of the tagger. Our main contribution in this research work is the designed Khasi POS corpus. Till date, there has been no form of any kind of Khasi corpus developed or formally developed. In the present designed Khasi POS corpus, each word is tagged manually using the designed tagset. Methods of deep learning have been used to experiment with our designed Khasi POS corpus. The POS tagger based on BiLSTM, combinations of BiLSTM with CRF, and character-based embedding with BiLSTM are presented. The main challenges of understanding and handling Natural Language toward Computational linguistics to encounter are anticipated. In the presently designed corpus, we have tried to solve the problems of ambiguities of words concerning their context usage, and also the orthography problems that arise in the designed POS corpus. The designed Khasi corpus size is around 96,100 tokens and consists of 6,616 distinct words. Initially, while running the first few sets of data of around 41,000 tokens in our experiment the taggers are found to yield considerably accurate results. When the Khasi corpus size has been increased to 96,100 tokens, we see an increase in accuracy rate and the analyses are more pertinent. As results, accuracy of 96.81% is achieved for the BiLSTM method, 96.98% for BiLSTM with CRF technique, and 95.86% for character-based with LSTM. Concerning substantial research from the NLP perspectives for Khasi, we also present some of the recently existing POS taggers and other NLP works on the Khasi language for comparative purposes. Sunita Warjri, Partha Pakray, Saralin Lyngdoh, Arnab Kumar Maji |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2021 | Interpretable semantic textual similarity of sentences using alignment of chunks with classification and regression
Goutam Majumder, Partha Pakray, Ranjita Das, David Pinto 0001 |
Appl. Intell. | 2 |
| 2021 | An Improved English-to-Mizo Neural Machine TranslationabstractMachine Translation is an effort to bridge language barriers and misinterpretations, making communication more convenient through the automatic translation of languages. The quality of translations produced by corpus-based approaches predominantly depends on the availability of a large parallel corpus. Although machine translation of many Indian languages has progressively gained attention, there is very limited research on machine translation and the challenges of using various machine translation techniques for a low-resource language such as Mizo. In this article, we have implemented and compared statistical-based approaches with modern neural-based approaches for the English–Mizo language pair. We have experimented with different tokenization methods, architectures, and configurations. The performance of translations predicted by the trained models has been evaluated using automatic and human evaluation measures. Furthermore, we have analyzed the prediction errors of the models and the quality of predictions based on variations in sentence length and compared the model performance with the existing baselines. Candy Lalrempuii, Badal Soni, Partha Pakray |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Classification of Microarray Gene Expression Data using Weighted Grey Wolf Optimizer based Fuzzy ClusteringabstractWith the emergence of DNA microarray technology, scientists are continuously studying the expression levels of a large number of genes over different instances of time points. Analyzing the DNA microarray data confirm that the expression levels of two different genes vary simultaneously with the effect of external stimuli exhibiting different patterns. Therefore soft fuzzy clustering plays an important role in detecting patterns belonging to multiple clusters at the same time. Therefore this article proposes a novel fuzzy clustering technique utilizing weighted distance measure instead of the Euclidean distance using Grey Wolf Optimizer (GWO) as the global optimization techniques. Here clustering of microarray data is posed as a single objective optimization problem where the objective is to minimize the variability within the cluster and simultaneously maximizes the variability between the cluster. The newly proposed fuzzy-based weighted GWO clustering technique (Fuzzy-WDGWO) is then compared with some of the existing clustering techniques. Four different artificial datasets and three different real-life gene expression datasets have been considered to verify the efficiency of the proposed Fuzzy-WDGWO clustering technique both numerically as well as pictorially. Experimental analysis and performance evaluation of the proposed Fuzzy-WDGWO clustering technique show the superiority over the other existing clustering techniques such as PSO-FCM, FA-FCM, GWO-FCM, GWO-FCM, DE-FCM. Amika Achom, Ranjita Das, Partha Pakray, Sriparna Saha 0001 |
TENCON | 3 |
| 2019 | English-Mizo Machine Translation using neural and statistical approaches
Amarnath Pathak, Partha Pakray, Jereemi Bentham |
Neural Comput. Appl. | 2 |
| 2018 | Addressing the Issue of Unavailability of Parallel Corpus Incorporating Monolingual Corpus on PBSMT System for English-Manipuri Translation
Amika Achom, Partha Pakray, Alexander F. Gelbukh |
CICLing (1) | 2 |
| 2018 | An Abstractive Text Summarization Using Recurrent Neural Network
Dipanwita Debnath, Partha Pakray, Ranjita Das, Alexander F. Gelbukh |
CICLing (2) | 2 |
| 2018 | Aggregation of multi-objective fuzzy symmetry-based clustering techniques for improving gene and cancer classification
Sriparna Saha 0001, Ranjita Das, Partha Pakray |
Soft Comput. | 3 |
| 2017 | Designing an Ontology for Physical Exercise Actions
Sandeep Kumar Dash, Partha Pakray, Robert Porzel, Jan D. Smeddinck, Rainer Malaka, Alexander F. Gelbukh |
CICLing (1) | 2 |
| 2016 | Multiword Expressions (MWE) for Mizo Language: Literature Survey
Goutam Majumder, Partha Pakray, Zoramdinthara Khiangte, Alexander F. Gelbukh |
CICLing (1) | 2 |
| 2015 | Mining Parallel Resources for Machine Translation from Comparable Corpora
Santanu Pal, Partha Pakray, Alexander F. Gelbukh, Josef van Genabith |
CICLing (1) | 2 |
| 2011 | Answer Validation Using Textual Entailment
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay |
CICLing (2) | 1 |
| 2011 | Answer Validation through Textual Entailment
Partha Pakray |
NLDB | 1 |
| 2010 | A Syntactic Textual Entailment System Based on Dependency Parser
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay |
CICLing | 1 |