Sudeshna Sarkar

dblp:61/3197 · DBLP profile ↗
← Back
54ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-3439-4282ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 7 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Time series prediction of multi-spectral images using self-supervised learning and its applications in cloud removal and land use analysis
Shankho Subhra Pal, Jayanta Mukhopadhyay, Sudeshna Sarkar
Eng. Appl. Artif. Intell.3
2025 Subclass-Aware Inclusive Classifier via Repulsive Hidden Strata
abstract
Classification models in machine learning are typically trained using coarse-grained class labels. Although these models often achieve strong overall accuracy, their performance is asymmetric across the subclasses that arise out of a common phenomenon called hidden stratification. Generally, the latent subclasses within each class differ substantially in distribution and characteristics, resulting in poor generalization for underrepresented groups. Moreover, imbalanced subclass distributions lead to majority subclasses dominating training, resulting in biased and less reliable models, especially for safety-critical applications (such as medical). To address these challenges, we propose a novel framework that attempts to uncover hidden subclasses via a repulsive point process. Our approach then leverages these fine-grained labels to make the classifier more inclusive across the subclasses. Our approach identifies subclasses without requiring additional supervision, thereby promoting diversity and reducing sensitivity to subclass imbalance. Extensive experiments on four benchmark datasets demonstrate consistent and significant improvements over state-of-the-art baselines across both balanced and imbalanced subclass distributions, underscoring the effectiveness and generalizability of our approach.
Namita Bajpai, Jiaul H. Paik, Sudeshna Sarkar
CIKM3
2025 Dysarthric speech recognition: an investigation on using depthwise separable convolutions and residual connections
Seyed Reza Shahamiri, Krishnendu Mandal, Sudeshna Sarkar
Neural Comput. Appl.3
2025 Balanced seed selection for K-means clustering with determinantal point process
Namita Bajpai, Jiaul H. Paik, Sudeshna Sarkar
Pattern Recognit.3
2024 A Novel Multi-Stage Prompting Approach for Language Agnostic MCQ Generation Using GPT
Subhankar Maity, Aniket Deroy, Sudeshna Sarkar
ECIR (3)3
2024 How Ready Are Generative Pre-trained Large Language Models for Explaining Bengali Grammatical Errors?
Subhankar Maity, Aniket Deroy, Sudeshna Sarkar
EDM3
2024 Exploring the Capabilities of Prompted Large Language Models in Educational and Assessment Applications
Subhankar Maity, Aniket Deroy, Sudeshna Sarkar
EDM3
2024 From Text to Context: An Entailment Approach for News Stakeholder Classification
abstract
Navigating the complex landscape of news articles involves understanding the various actors or entities involved, referred to as news stakeholders. These stakeholders, ranging from policymakers to opposition figures, citizens, and more, play pivotal roles in shaping news narratives. Recognizing their stakeholder types, reflecting their roles, political alignments, social standing, and more, is paramount for a nuanced comprehension of news content. Despite existing works focusing on salient entity extraction, coverage variations, and political affiliations through social media data, the automated detection of stakeholder roles within news content remains an underexplored domain. In this paper, we bridge this gap by introducing an effective approach to classify stakeholder types in news articles. Our method involves transforming the stakeholder classification problem into a natural language inference task, utilizing contextual information from news articles and external knowledge to enhance the accuracy of stakeholder type detection. Moreover, our proposed model showcases efficacy in zero-shot settings, further extending its applicability to diverse news contexts.
Alapan Kuila, Sudeshna Sarkar
SIGIR2
2024 Finding hierarchy of clusters
Shankho Subhra Pal, Jayanta Mukhopadhyay, Sudeshna Sarkar
Pattern Recognit. Lett.3
2023 An Approach to Battery Pack Balancing Control Optimizing the Usable Capacity of the Battery Pack
abstract
Lithium-ion batteries are widely used in electric vehicles and energy storage systems because of their high energy density, high power density and long service life. However, the degradation of available capacity caused by the consistency difference of batteries has always been a key technical problem limiting the long-term stable operation of battery packs. In this paper, a balancing control strategy considering the maximum available capacity of the battery pack is proposed. The balancing operation is conducted in the process of charging and discharging respectively, thus the available capacity of the battery pack can be optimized. Firstly, the influence of Coulomb efficiency on the imbalance of battery quantity is analyzed theoretically. Then, the realization way of maximizing the available capacity of battery pack is sought by deducing the formula of available capacity of battery pack. Finally, artificial fish swarm algorithm was used to obtain the balanced electric quantity during charging and discharging respectively, and experimental verification was carried out. The results show that the charged amount of electric quantity for battery pack increases by 6.1%, and that for discharging increases by 7.9%, as compared to the capacity-based equalization method.
Sudeshna Sarkar
IECON2
2023 Semi-Supervised Semantic Segmentation of Hyper-Spectral Images
abstract
Hyper-spectral images have hundreds of bands, which provide rich information compared to popular RGB images. This information can be leveraged to accurately segment images into multiple fine-grained classes. However, processing high-dimensional data poses challenges. In this work, we propose two methods for land cover classification using hyper-spectral images, aiming to achieve accurate classification into multiple fine-grained classes. One of the main challenges in land cover classification is the availability of labeled data. To address this, we employ semi-supervised methods that require less labeled data and effectively utilize the vast amount of un-labeled data. We evaluate our methods on the Indian Pine dataset, while varying the amount of labeled pixels used for training.
Shankho Subhra Pal, Jayanta Mukhopadhyay, Sudeshna Sarkar
IGARSS3
2023 Meta-ED: Cross-lingual Event Detection Using Meta-learning for Indian Languages
abstract
Lack of annotated data is a major concern in Event Detection (ED) tasks for low-resource languages. Cross-lingual ED seeks to address this issue by transferring information across various languages to improve overall performance. In this article, we propose a method for cross-lingual ED with a few training instances. We present a model agnostic meta-learning approach for few-shot cross-lingual ED that is able to find good parameter initialization and enables fast adaptation to new low-resource languages. We evaluate our model on four Indian languages. The results show that our approach significantly outperforms the base model.
Aniruddha Roy, Isha Sharma, Sudeshna Sarkar, Pawan Goyal 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 Does Meta-learning Help mBERT for Few-shot Question Generation in a Cross-lingual Transfer Setting for Indic Languages?
abstract
Few-shot Question Generation (QG) is an important and challenging problem in the Natural Language Generation (NLG) domain. Multilingual BERT (mBERT) has been successfully used in various Natural Language Understanding (NLU) applications. However, the question of how to utilize mBERT for few-shot QG, possibly with cross-lingual transfer, remains. In this paper, we try to explore how mBERT performs in few-shot QG (cross-lingual transfer) and also whether applying meta-learning on mBERT further improves the results. In our setting, we consider mBERT as the base model and fine-tune it using a seq-to-seq language modeling framework in a cross-lingual setting. Further, we apply the model agnostic meta-learning approach to our base model. We evaluate our model for two low-resource Indian languages, Bengali and Telugu, using the TyDi QA dataset. The proposed approach consistently improves the performance of the base model in few-shot settings and even works better than some heavily parameterized models. Human evaluation also confirms the effectiveness of our approach.
Aniruddha Roy, Rupak Kumar Thakur, Isha Sharma, Ashim Gupta, Amrith Krishna, Sudeshna Sarkar, Pawan Goyal 0002
COLING6
2021 Towards Reducing the Pendency of Cases at Court: Automated Case Analysis of Supreme Court Judgments in India
abstract
The Indian court system generates huge amounts of data relating to administration, pleadings, litigant behaviour, and court decisions on a regular basis. But the existing Judiciary is incapable of managing these vast troves of data efficiently that causes delays and pendency of a large volume of cases in the courts. Some of these time-consuming tasks involve case briefing, examining the legal issues, facts, legal principles, observations, and other significant aspects submitted by the contending parties in the court. In other words, computational methods to understand the underlying structure of a case document will directly aid the lawyers to perform these tasks efficiently and improve the overall efficiency of the Justice delivery system. Application of Computational techniques (such as Natural Language Processing) can help to gather and sift through these vast troves of information, identify patterns, extract the document structure, draft documents and make the information available online. Traditionally lawyers are trained to examine cases using the Case Law Analysis approach for case briefing. In this article, the authors aim to establish the importance and relevance of the automated case analysis problem in the legal domain. They introduce a novel case analysis structure for the supreme court judgment documents and define twelve different case law labels that are used by legal professionals to identify the structure. Finally the authors propose a method for automated case analysis, which will directly aid the lawyers to prepare speedy and efficient case briefs and drastically reduce the time taken by them in litigation.
Shubham Pandey, Ayan Chandra, Sudeshna Sarkar, Uday Shankar
JURIX3
2020 Transform, Combine, and Transfer: Delexicalized Transfer Parser for Low-resource Languages
abstract
Transfer parsing has been used for developing dependency parsers for languages with no treebank by using transfer from treebanks of other languages (source languages). In delexicalized transfer, parsed words are replaced by their part-of-speech tags. Transfer parsing may not work well if a language does not follow uniform syntactic structure with respect to its different constituent patterns. Earlier work has used information derived from linguistic databases to transform a source language treebank to reduce the syntactic differences between the source and the target languages. We propose a transformation method where a source language pattern is transformed stochastically to one of the multiple possible patterns followed in the target language. The transformed source language treebank can be used to train a delexicalized parser in the target language. We show that this method significantly improves the average performance of single-source delexicalized transfer parsers. We also show that, in the multi-source settings, parsers trained using a concatenation of transformed source language treebanks work better when a subset of the source language treebanks is used rather than concatenating all of them or only one. However, the problem of selecting the subset of treebanks whose combination gives the best-performing parser from the set of all the available treebanks is hard. We propose a greedy selection heuristic based on the labelled attachment scores of the corresponding single-source parsers trained using the treebanks after transformation.
Ayan Das 0002, Sudeshna Sarkar
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 A Survey of the Model Transfer Approaches to Cross-Lingual Dependency Parsing
abstract
Cross-lingual dependency parsing approaches have been employed to develop dependency parsers for the languages for which little or no treebanks are available using the treebanks of other languages. A language for which the cross-lingual parser is developed is usually referred to as the target language and the language whose treebank is used to train the cross-lingual parser model is referred to as the source language. The cross-lingual parsing approaches for dependency parsing may be broadly classified into three categories: model transfer, annotation projection, and treebank translation. This survey provides an overview of the various aspects of the model transfer approach of cross-lingual dependency parsing. In this survey, we present a classification of the model transfer approaches based on the different aspects of the method. We discuss some of the challenges associated with cross-lingual parsing and the techniques used to address these challenges. In order to address the difference in vocabulary between two languages, some approaches use only non-lexical features of the words to train the models while others use shared representations of the words. Some approaches address the morphological differences by chunk-level transfer rather than word-level transfer. The syntactic differences between the source and target languages are sometimes addressed by transforming the source language treebanks or by combining the resources of multiple source languages. Besides cross-lingual transfer parser models may be developed for a specific target language or it may be trained to parse sentences of multiple languages. With respect to the above-mentioned aspects, we look at the different ways in which the methods can be classified. We further classify and discuss the different approaches from the perspective of the corresponding aspects. We also demonstrate the performance of the transferred models under different settings corresponding to the classification aspects on a common dataset.
Ayan Das 0002, Sudeshna Sarkar
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Drug-Drug Interactions Prediction Based on Drug Embedding and Graph Auto-Encoder
abstract
Identification of potential Drug-Drug Interactions (DDI) for newly developed drugs is essential in public healthcare. Computational methods of DDI prediction rely on known interactions to learn possible interaction between drug pairs whose interactions are unknown. Past work has used various similarity measures of drugs to predict DDIs. In this paper, we propose an effective approach to DDI Prediction using rich drug representations utilizing multiple knowledge sources. We have used the Drug-Target Interaction (DTI) Network to learn an embedding of drugs by using the metapath2vec algorithm. We have also used drug representation gained from the rich chemical structure representation of drugs using Variational Auto-Encoder. The DDI prediction problem is modeled as a link prediction problem in the DDI network containing known interactions. We represent the nodes in the DDI network as their embeddings. We apply a link prediction algorithm based on Graph Auto-Encoders to predict additional edges in this network, which are potential interactions. We have evaluated our approach on three benchmark DDI datasets, namely DrugBank, SemMedDB, and BioSNAP. Experimental results demonstrate that the proposed method outperforms the prior methods in terms of several performance metrics (AUC, AUPR, and F1-score) on all the datasets. Furthermore, we have also evaluated the role of the individual type of drug representation embeddings in boosting up the performance of DDI Prediction.
Sukannya Purkayastha, Ishani Mondal, Sudeshna Sarkar, Pawan Goyal 0002, Jitesh K. Pillai
BIBE3
2019 MorphBen: A Neural Morphological Analyzer for Bengali Language
Ayan Das 0002, Sudeshna Sarkar
CICLing (1)2
2019 Recommendation for Multi-stakeholders and through Neural Review Mining
abstract
Recommender systems are able to produce a list of recommended items tailored to user preferences, while the end user is the only stakeholder in these traditional system. However, there could be multiple stakeholders in several applications domains (e.g., e-commerce, movies, music). Recommendations are necessary to be produced by balancing the needs of different stakeholders. First session of this tutorial introduces multi-stakeholder recommender systems (MSRS) with several case studies, and discusses the corresponding methods and challenges in MSRS. Reviews in an e-commerce platform may be mined to address cold-start problem and to generate explanations. Our earlier tutorial covered aspect-based sentiment analysis of products and topic models/distributed representations that bridge vocabulary gap between user reviews and product descriptions. Focus in the second session of this tutorial instead is on recent neural methods for review text mining - covering hands-on code for its use to enhance product recommendation. Each section will introduce topics from various mechanism (e.g., attention) and task (e.g., review ranking) perspectives, present cutting-edge research and a walk-through of programs executed on Jupyter notebook using real-world data sets.
Muthusamy Chelliah, Yong Zheng 0001, Sudeshna Sarkar
CIKM3
2019 Fully Contextualized Biomedical NER
Ashim Gupta, Pawan Goyal 0002, Sudeshna Sarkar, Mahanandeeshwar Gattu
ECIR (2)3
2019 Using Communities of Words Derived from Multilingual Word Vectors for Cross-Language Information Retrieval in Indian Languages
abstract
We investigate the use of word embeddings for query translation to improve precision in cross-language information retrieval (CLIR). Word vectors represent words in a distributional space such that syntactically or semantically similar words are close to each other in this space. Multilingual word embeddings are constructed in such a way that similar words across languages have similar vector representations. We explore the effective use of bilingual and multilingual word embeddings learned from comparable corpora of Indic languages to the task of CLIR. We propose a clustering method based on the multilingual word vectors to group similar words across languages. For this we construct a graph with words from multiple languages as nodes and with edges connecting words with similar vectors. We use the Louvain method for community detection to find communities in this graph. We show that choosing target language words as query translations from the clusters or communities containing the query terms helps in improving CLIR. We also find that better-quality query translations are obtained when words from more languages are used to do the clustering even when the additional languages are neither the source nor the target languages. This is probably because having more similar words across multiple languages helps define well-defined dense subclusters that help us obtain precise query translations. In this article, we demonstrate the use of multilingual word embeddings and word clusters for CLIR involving Indic languages. We also make available a tool for obtaining related words and the visualizations of the multilingual word vectors for English, Hindi, Bengali, Marathi, Gujarati, and Tamil.
Paheli Bhattacharya, Pawan Goyal 0002, Sudeshna Sarkar
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Concept to code: learning distributed representation of heterogeneous sources for recommendation
abstract
Recommender Systems fuel e-commerce. Deep Learning techniques have started to make an impact in building recommenders. Many techniques have been proposed recently to create low dimensional embeddings of heterogeneous sources including users, items, text and images that can capture the semantic relationships between them. Such combined embeddings play a very important role in the effectiveness of a Recommender System.
Omprakash Sonie, Sudeshna Sarkar
RecSys2
2017 A deep learning based approach with adversarial regularization for Doppler weather radar ECHO prediction
abstract
Precipitation nowcasting is an important component for accurate weather modeling and Doppler radar data acts as an important input for nowcasting models. In this work, we propose a deep learning based approach for radar echo states prediction. Our approach uses a hybrid structure of convolutions within Long Short Term Memory recurrent network structure and a discriminator network is added in the loss objective to refine the predictions acting as a regularizer. This models the spatio-temporal nature of the problem explicitly in the neural network. We show that this model can be applied for fine grained short term precipitation prediction with improvement in evaluation metrics as compared to strong baselines. The proposed model improves recall score by 11% compared to without adversarial regularization. Results are presented using usual train, test strategy for the task of echo state prediction and derived precipitation based skill scores on the data from Seattle, WA, USA.
Sonam Singh, Sudeshna Sarkar, Pabitra Mitra
IGARSS2
2017 Leveraging Convolutions in Recurrent Neural Networks for Doppler Weather Radar Echo Prediction
Sonam Singh, Sudeshna Sarkar, Pabitra Mitra
ISNN (2)2
2017 Product Recommendations Enhanced with Reviews
abstract
User-written product reviews contain rich information about user preferences for product features and provide helpful explanations that are often used by shoppers to make their purchase decisions. E-commerce recommender systems can benefit enormously by also exploiting experiences of multiple customers captured in product reviews. In this tutorial, we present a range of techniques that allow recommender systems in e-commerce websites to take full advantage of reviews. This includes text mining methods for feature-specific sentiment analysis of products, topic models and distributed representations that bridge the vocabulary gap between user reviews and product descriptions. We present recommender algorithms that use review information to address the cold-start problem and generate recommendations with explanations. We discuss examples and experiences from an online marketplace (i.e., Flipkart).
Muthusamy Chelliah, Sudeshna Sarkar
RecSys2
2017 Recommendation of High Quality Representative Reviews in e-commerce
abstract
Many users of e-commerce portals commonly use customer reviews for making purchase decisions. But a product may have tens or hundreds of diverse reviews leading to information overload on the customer. The main objective of our work is to develop a recommendation system to recommend a subset of reviews that have high content score and good coverage over different aspects of the product along with their associated sentiments. We address the challenge which arises due to the fact that similar aspects are mentioned in different reviews using different natural language expressions. We use vector representations to identify mentions of similar aspects and map them with aspects mentioned in product features specifications. Review helpfulness score may act as a proxy for the quality of reviews, but new reviews do not have any helpfulness score. We address the cold start problem by using a dynamic convolutional neural network to estimate the quality score from review content. The system is evaluated on datasets from Amazon and Flipkart and is found to be more effective than the competing methods.
Debanjan Paul, Sudeshna Sarkar, Muthusamy Chelliah, Chetan Kalyan, Prajit Prashant Sinai Nadkarni
RecSys2
2016 Enhancing Neural Network Based Dependency Parsing Using Morphological Information for Hindi
Agnivo Saha, Sudeshna Sarkar
CICLing (1)2
2016 Preference relations based unsupervised rank aggregation for metasearch
Maunendra Sankar Desarkar, Sudeshna Sarkar, Pabitra Mitra
Expert Syst. Appl.2
2015 Improving Cross Language Information Retrieval Using Corpus Based Query Suggestion Approach
Rajendra Prasath, Sudeshna Sarkar, Philip O'Reilly
CICLing (2)2
2015 Assisting web document retrieval with topic identification in tourism domain
abstract
In this work, we present a domain specific Information Retrieval (IR) system that identifies query and document topics and use them for better documents retrieval. We focus on retrieving documents having the specific types of information as that of the user query related to the tourism domain. Base d on our past experience in handling tourism specific information, we observed that the query intent in the tourism domain largely span over a few major types. Based on this observation, we present an approach for document retrieval based on query and documents type identification. To do this, we have identified the major types (topics) in the tourism domain and built an ontology of the tourism domain. We developed a document classifier to identify the topic of web documents, and a query classifier to identify the topic of the user query, both pertaining to the tourism domain. The proposed IR system performs document retrieval by matching the type of user query with the matching type of documents. The experimental results show that the tourism specific topic identification of queries and documents improves the retrieval of documents having more specific information to satisfy user queries in the tourism domain.
Rajendra Prasath, Vijai Kumar, Sudeshna Sarkar
Web Intell.3
2013 A Mobility Simulation Framework Of Humans With Group Behavior Modeling
abstract
We present a mobility simulation framework that simulates the movement behaviors of people to generate spatiotemporal movement data. There is a growing interest in applications that make use of patterns mined from spatio-temporal data. However, since the availability of actual spatio-temporal movement data in the public domain is limited, it is useful to have simulation frameworks that generate data close to the real life behavior of people, so that data mining techniques can be tested. We argue that modeling group behavior effectively is a key element of any real-life simulation framework, because there are many applications that require the knowledge of groups and events. In this work, we propose generic models to represent individual and group movement behaviors. We present an algorithm that takes various behaviors created using the proposed models, and generates spatio-temporal movement data for as many individuals as needed. Experimental analysis shows the efficacy of the proposed framework handling a broad spectrum of behaviors with high scalability.
Aurosish Mishra, Satya Gautam Vadlamudi, P. P. Chakrabarti 0001, Sudeshna Sarkar, Tridib Mukherjee, Nathan Gnanasambandam
ICDM5
2012 Preference Relation Based Matrix Factorization for Recommender Systems
Maunendra Sankar Desarkar, Roopam Saxena, Sudeshna Sarkar
UMAP3
2012 A comparative study on feature reduction approaches in Hindi and Bengali named entity recognition
Sujan Kumar Saha, Pabitra Mitra, Sudeshna Sarkar
Knowl. Based Syst.3
2011 Prenominal Modifier Ordering in Bengali Text Generation
Sumit Das 0001, Anupam Basu, Sudeshna Sarkar
CICLing (1)3
2010 Unsupervised Feature Generation using Knowledge Repositories for Effective Text Categorization
abstract
We propose an unsupervised feature generation algorithm using the repositories of human knowledge for effective text categorization. Conventional bag of words (BOW) depends on the presence / absence of keywords to classify the documents. To understand the actual context behind these keywords, we use knowledge concepts / hyperlinks from external knowledge sources through content and structure mining on Wikipedia. Then, the features of knowledge concepts are clustered to generate knowledge cluster vectors with which the input text documents are mapped into a high dimensional feature space and the classification is performed. The simulation results show that the proposed approach identifies associated features in the text collection and yields an improved classification accuracy.
Rajendra Prasath, Sudeshna Sarkar
ECAI2
2010 Aggregating preference graphs for collaborative rating prediction
abstract
Collaborative filtering is a widely used technique for rating prediction in recommender systems. Memory based collaborative filtering algorithms assign weights to the users to capture similarities between them. The weighted average of similar users' ratings for the test item is output as prediction. We propose a memory based algorithm that is markedly different from the existing approaches. We use preference relations instead of absolute ratings for similarity calculations, as preference relations between items are generally more consistent than ratings across like-minded users. Each user's ratings are viewed as a preference graph. Similarity weights are learned using an iterative method motivated by online learning. These weights are used to create an aggregate preference graph. Ratings are inferred to maximally agree with this aggregate graph. The use of preference relations allows the rating of an item to be influenced by other items, which is not the case in the weighted-average approaches of the existing techniques. This is very effective when the data is sparse, specially for the items rated by few users. Our experiments show that the our method outperforms other methods in the sparse regions. However, for dense regions, sometimes our results are comparable to the competing approaches, and sometimes worse.
Maunendra Sankar Desarkar, Sudeshna Sarkar, Pabitra Mitra
RecSys2
2010 A composite kernel for named entity recognition
Sujan Kumar Saha, Shashi Narayan, Sudeshna Sarkar, Pabitra Mitra
Pattern Recognit. Lett.3
2009 Stylometric Analysis of Bloggers' Age and Gender
Sumit Goswami, Sudeshna Sarkar, Mayur Rustagi
ICWSM2
2009 Feature selection techniques for maximum entropy based biomedical named entity recognition
Sujan Kumar Saha, Sudeshna Sarkar, Pabitra Mitra
J. Biomed. Informatics2
2008 Word Clustering and Word Selection Based Feature Reduction for MaxEnt Based Hindi NER
Sujan Kumar Saha, Pabitra Mitra, Sudeshna Sarkar
ACL3
2008 Bengali and Hindi to English CLIR Evaluation
Debasis Mandal, Sandipan Dandapat, Mayank Gupta 0001, Pratyush Banerjee, Sudeshna Sarkar
IJCNLP5
2008 A Hybrid Named Entity Recognition System for South and South East Asian Languages
Sujan Kumar Saha, Sanjay Chatterji, Sandipan Dandapat, Sudeshna Sarkar, Pabitra Mitra
IJCNLP4
2008 A Hybrid Feature Set based Maximum Entropy Hindi Named Entity Recognition
Sujan Kumar Saha, Sudeshna Sarkar, Pabitra Mitra
IJCNLP2
2007 Automatic Part-of-Speech Tagging for Bengali: An Approach for Morphologically Rich Languages in a Poor Resource Scenario
Sandipan Dandapat, Sudeshna Sarkar, Anupam Basu
ACL2
2007 Sahayika: A framework for participatory authoring of knowledge structures for education domain
abstract
In countries like India, a great deal of diversity exists in language, culture and socio-economic conditions. In order to deliver computer aided education, efforts have to be put for proper management of concerned domain and learning materials keeping in mind this diversity. Participatory authoring of domain knowledge structure is of immense importance in this regard. In this paper, we describe a knowledge structure which is effective in capturing the structure of the learning materials of school level subjects. Ontologies have gained importance in representing the knowledge of the domain in a formal and machine understandable form in areas like intelligent information processing. Thus it can provide the platform for effective extraction of information and many other applications. We describe different aspects of manual ontology engineering in developing application specific domain ontology. The domain of our interest is the education domain where we are interested in retrieving relevant Web documents for the curriculum related requirements of school students. We identify an effective way of structuring the knowledge about these domains, which allows us to clearly demarcate the roles of topics, concepts, and actual words. We also describe applications in the area of information retrieval and in indexing document repositories in connection with e-learning where our ontology plays an important role. In this paper, we have provided a framework, Sahayika, for building knowledge structures in education domain.
Plaban Kumar Bhowmick, Sudipta Bhowmick, Devshri Roy, Sudeshna Sarkar, Anupam Basu
ICTD4
2007 Samvidha: A ICT system for personalized offline Internet Access for rural schools
abstract
Internet is a huge repository of quality learning materials and continues to grow in a faster rate. The school students may be benefited immensely as these learning materials may well supplement their curricular requirements. But Access to the Internet is costly, because it is very expensive to maintain a persistent Internet connection. For some schools in the developing countries like India, this cost may not be affordable specifically in rural schools. This makes way to a digital divide between the rural and urban schools which is unwanted. For these rural schools, limiting the amount of bandwidth consumed is of paramount importance. It is necessary that the schools be connected to the Internet for the least time, in order to minimize the access cost. In this paper, we present a system Samvidha that allows the rural school students to access the Internet contents in an offline fashion.
Plaban Kumar Bhowmick, Sudeshna Sarkar, Sunandan Chakraborty, Anupam Basu
ICTD2
2007 Shikshak: An Intelligent Tutoring System Authoring tool for rural education
abstract
Low literacy scenario in India and other developing nation demands an alternative learning environment to deal with the problem. Lack of trained teachers, high dropout rates are some of the major problems that need to be addressed. Intelligent Tutoring System (ITS) or ITS Authoring tools (ITSAT) can be thought of as a possible solution to these problems. In this paper we present Shikshak, an ITSAT developed by us and discuss its deployment in the district of Paschim Medinipur, West Bengal along with its sample effect on primary education.
Sunandan Chakraborty, Tamali Bhattacharya, Plaban Kumar Bhowmick, Anupam Basu, Sudeshna Sarkar
ICTD5
2007 Investigation and modeling of the structure of texting language
Monojit Choudhury, Rahul Saraf, Vijit Jain, Animesh Mukherjee 0001, Sudeshna Sarkar, Anupam Basu
Int. J. Document Anal. Recognit.5
2003 Computation of intraprocedural dynamic program slices
Ganga Bishnu Mund, Rajib Mall, Sudeshna Sarkar
Inf. Softw. Technol.3
2002 An efficient dynamic program slicing technique
Ganga Bishnu Mund, Rajib Mall, Sudeshna Sarkar
Inf. Softw. Technol.3
2001 CROWSE: A System for Organizing Repositories and Web Search Results
abstract
No abstract available.
Kinshuman, Sudeshna Sarkar
SIGIR2
2000 Study of methods for model reduction in transition systems
abstract
We consider the reinforcement learning problem where the agent interacts with the environment by taking actions. The agent receives some reward from the environment and changes state as a result. The problem is to find a good policy in large transition systems such that the agent's value function is maximized. If the transition system is Markovian, there exist several algorithms for finding the optimum policy in such systems. However, finding the optimum policy may take considerable time for large systems, and the learning algorithm will not converge if we allow a smaller number of iterations. Our objective is to investigate methods of reducing the state space of large transition systems, and to evaluate the effect of model reduction on getting a good policy in reasonable time. We conjecture that when we have limited time, it may be wise to reduce the transition system and then apply our learning algorithm to the smaller system.
Sudeshna Sarkar, A. V. Subramaniam, Rajdeep Neogi
SMC1
1998 A Framework for Learning in Search-Based Systems
abstract
We provide an overall framework for learning in search based systems that are used to find optimum solutions to problems. This framework assumes that prior knowledge is available in the form of one or more heuristic functions (or features) of the problem domain. An appropriate clustering strategy is used to partition the state space into a number of classes based on the available features. The number of classes formed will depend on the resource constraints of the system. In the training phase, example problems are run using a standard admissible search algorithm. In this phase, heuristic information corresponding to each class is learned. This new information can be used in the problem solving phase by appropriate search algorithms so that subsequent problem instances can be solved more efficiently. In this framework, we also show that heuristic information of forms other than the conventional single valued underestimate value can be used, since we maintain the heuristic of each class explicitly. We show some novel search algorithms that can work with some such forms. Experimental results have been provided for some domains.
Sudeshna Sarkar, P. P. Chakrabarti 0001, Sujoy Ghose
IEEE Trans. Knowl. Data Eng.1
1998 Learning while solving problems in best first search
abstract
We investigate the role of learning in search-based systems for solving optimization problems. We use a learning model, where the values of a set of features can be used to induce a clustering of the problem state space. The feasible set of h* values corresponding to each cluster is called h*set. If we relax the optimality guarantee, and tolerate a risk factor, the distribution of h*set can be used to expedite search and produce results within a given risk of suboptimality. The off-line learning method consists of solving a batch of problems by using A* to learn the distribution of the h*set in the learning phase. This distribution can be used to solve the rest of the problems effectively. We show how the knowledge acquisition phase can be integrated with the problem solving phase. We present a continuous online learning scheme that uses an "anytime" algorithm to learn continuously while solving problems.
Sudeshna Sarkar, P. P. Chakrabarti 0001, Sujoy Ghose
IEEE Trans. Syst. Man Cybern. Part A1