VLDB 2026 Research / reviewers in the wild / expert
Lipika Dey
dblp:09/3737
· DBLP profile ↗
43ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0003-3831-5545ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 22 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MMFood'25: 1st International Workshop on Multi-modal Food ComputingabstractMMFood'25, the 1st International Workshop on Multi-modal Food Computing, brings together researchers and practitioners at the intersection of artificial intelligence, computer vision, natural language processing, and sensory modeling to advance the study of food. The workshop highlights how multimodal methods can be applied to food recognition, recommendation, analysis, and monitoring to address pressing challenges in health, nutrition, sustainability, and food culture. It features a rich program including a keynote, an invited talk, paper presentations, a poster session, and a panel discussion on the role of multimodal AI in preserving cultural heritage, fostering sustainable food futures, and enabling personal well-being. By convening experts from academia, industry, and healthcare, MMFood'25 provides a unique platform for fostering interdisciplinary collaboration and for shaping the emerging field of multimodal food computing. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746264. Lipika Dey, Marianna Obrist, Stavroula G. Mougiakakou |
ACM Multimedia | 1 |
| 2023 | Generating insights about financial asks from Reddit posts and user interactionsabstractAs an increasingly large number of people turn to platforms like Reddit, YouTube, Twitter, Instagram, etc. for financial advice, generating insights about the content generated and interactions taking place within these platforms have become a key research question. This study proposes content and interaction analysis techniques for a large repository created from social media content, where people's interactions are centered around financial information exchange. We propose methods for content analysis that can generate human-interpretable insights using topic-centered clustering and multi-document abstractive summarization. We share details of insights generated from our experiments with a large repository of data gathered from subreddit for personal finance. We have also explored the use of ChatGPT and Vicuna for generating responses to queries and compared them with human responses. The methods proposed in this work are generic and applicable to all large social media platforms. Sachin Thukral, Suyash Sangwan, Vipul Chauhan, Lipika Dey |
ASONAM | 5 |
| 2023 | Deciphering Clinical Narratives - Augmented Intelligence for Decision Making in Healthcare SectorabstractClinical notes that describe details about diseases, symptoms, treatments, and observed reactions of patients to them, are valuable resources to generate insights about the effectiveness of treatments.Their role in designing better clinical decision making systems is being increasingly acknowledged.However, the availability of clinical notes is still an issue due to privacy violation concerns.Hence most of the work done are on small datasets and neither the power of machine learning is fully utilized, nor is it possible to validate the models properly.With the availability of the Medical Information Mart for Intensive Care (MIMIC-III v1.4) dataset for researchers though, the problem has been somewhat eased.In this paper we have presented an overview of our earlier work on designing deep neural models for prediction of outcomes and hospital stay for patients using MIMIC data.We have also presented new work on patient stratification and explanation generation for patient cohorts.This is early work targeted towards studying trajectories for treatment for different cohorts of patients, which can ultimately lead to discovery of low-risk models for individual patients to ensure better outcomes. Lipika Dey, Sudeshna Jana, Tirthankar Dasgupta, Tanay Gupta |
FedCSIS | 1 |
| 2023 | Style Augmented Transformer Architecture for Automatic Essay AssessmentabstractIn this paper, we present a grammar and style aware transformer-based neural network for computing the quality of a text in an automatic essay-scoring task. The proposed model takes into consideration different grammatical error categories and discourse writing styles like, concreteness, uncertainty, conviction and commitment in text along with the pre-trained language models of a text document. We have evaluated the proposed model with the automated student assessment dataset. Our preliminary investigation shows that incorporating such stylistic vectors and grammatical error categories with the BERT based language model can give us a better understanding of improving the overall evaluation of the input essays. Tirthankar Dasgupta, Gaurav K. Singh, Lipika Dey |
ICALT | 3 |
| 2023 | Factors affecting user experience of contact tracing app during COVID-19: an aspect-based sentiment analysis of user-generated reviewabstractThis study aims to identify the critical factors influencing the user experience of contact tracing apps and the sentiments around them. For this purpose, we used Google play reviews of Aarogya Setu, a contact tracing app developed in India. First, we establish the relationship between review sentiment and review rating using regression between sentiment polarity and review rating. Then, we used a hybrid aspect-based sentiment analysis approach that uses unsupervised linguistic techniques to determine statistically significant concepts present in the review texts and cluster them into representative aspects that were then tagged under human supervision. Finally, supervised deep learning methods were applied for exhaustive extraction of the aspects and associated sentiments from the reviews. The final exercise of determining the key influencing factors was done by grouping these aspects under factors identified by marketing experts. A total of nine factors were identified, with the usefulness of the app being the most important factor. The findings of this study are essential for the development team and government to improve the application and increase adoption. Satyabhusan Dash, Avinash Jain, Lipika Dey, Tirthankar Dasgupta, Abir Naskar |
Behav. Inf. Technol. | 3 |
| 2021 | Query Specific Focused Summarization of Biomedical Journal ArticlesabstractDuring COVID-19, a large repository of relevant literature, termed as "CORD-19", was released by Allen Institute of AI.The repository being very large, and growing exponentially, concerned users are struggling to retrieve only required information from the documents.In this paper, we present a framework for generating focused summaries of journal articles.The summary is generated using a novel optimization mechanism to ensure that it definitely contains all essential scientific content.The parameters for summarization are drawn from the variables that are used for reporting scientific studies.We have evaluated our results on the CORD-19 dataset.The approach however is generic. Akshara Rai, Suyash Sangwan, Tushar Goel, Ishan Verma, Lipika Dey |
FedCSIS | 5 |
| 2018 | Analyzing Behavioral Trends in Community Driven Discussion Platforms Like RedditabstractThe aim of this paper is to present methods to systematically analyze individual and group behavioral patterns observed in community driven discussion platforms like Reddit where users exchange information and views on various topics of current interest. We conduct this study by analyzing the statistical behavior of posts and modeling user interactions around them. We have chosen Reddit as an example, since it has grown exponentially from a small community to one of the biggest social network platforms in the recent times. Due to its large user base and popularity, a variety of behavior is present among users in terms of their activity. Our study provides interesting insights about a large number of inactive posts which fail to gather attention despite their authors exhibiting Cyborg-like behavior to draw attention. We also present interesting insights about short-lived but extremely active posts emulating a phenomenon like Mayfly Buzz. Further, we present methods to find the nature of activity around highly active posts to determine the presence of Limelight hogging activity, if any. We analyzed over 2 million posts and more than 7 million user responses to them during entire 2008 and over 63 million posts and over 608 million user responses to them from August 2014 to July 2015 amounting to two one-year periods, in order to understand how social media space has evolved over the years. Sachin Thukral, Hardik Meisheri, Tushar Kataria, Aman Agarwal, Ishan Verma, Lipika Dey |
ASONAM | 7 |
| 2018 | Automatic Extraction of Causal Relations from Text using Linguistically Informed Deep Neural NetworksabstractIn this paper we have proposed a linguistically informed recursive neural network architecture for automatic extraction of cause-effect relations from text.These relations can be expressed in arbitrarily complex ways.The architecture uses word level embeddings and other linguistic features to detect causal events and their effects mentioned within a sentence.The extracted events and their relations are used to build a causal-graph after clustering and appropriate generalization, which is then used for predictive purposes.We have evaluated the performance of the proposed extraction model with respect to two baseline systems,one a rule-based classifier, and the other a conditional random field (CRF) based supervised model.We have also compared our results with related work reported in the past by other authors on SEMEVAL data set, and found that the proposed bidirectional LSTM model enhanced with an additional linguistic layer performs better.We have also worked extensively on creating new annotated datasets from publicly available data, which we are willing to share with the community. Tirthankar Dasgupta, Rupsa Saha, Lipika Dey, Abir Naskar |
SIGDIAL Conference | 3 |
| 2018 | Extraction and Visualization of Occupational Health and Safety Related Information from Open WebabstractIn this paper, we have proposed natural language processing and deep learning based techniques for the automatic extraction and curation of occupational health and safety related information from safety-related articles. Such articles typically contain details of the organizations that have been cited for violating the health and safety regulations, safety-related issues and incidents, the location of the incident, and finally details of the penalties incurred. We have done experiments with a collection of 5400 related articles. The end-product of our work is an occupational risk-register that contains details of safety incidents across geographies and time. This register can be further utilized for analytical and reporting purposes. Such information is extremely valuable to industries which see a high occurrence of occupational injuries. Tirthankar Dasgupta, Abir Naskar, Rupsa Saha, Lipika Dey |
WI | 4 |
| 2018 | Multi-Document Summarization Using Distributed Bag-of-Words ModelabstractAs the number of documents on the web is growing exponentially, multi-document summarization is becoming more and more important since it can provide the main ideas in a document set in short time. In this paper, we present an unsupervised centroid-based document-level reconstruction framework using distributed bag of words model. Specifically, our approach selects summary sentences in order to minimize the reconstruction error between the summary and the documents. We apply sentence selection and beam search, to further improve the performance of our model. Experimental results on two different datasets show significant performance gains compared with the state-of-the-art baselines. Kaustubh Mani, Ishan Verma, Hardik Meisheri, Lipika Dey |
WI | 4 |
| 2018 | Air Pollutant Severity Prediction Using Bi-Directional LSTM NetworkabstractAir pollution has emerged as a universal concern across the globe affecting human health. This increasing danger motivates the study of systems for predicting air pollutant severities ahead of time. In this paper, we have proposed the use of a bi-directional LSTM model to predict air pollutant severity levels ahead of time. We have shown that the predictions can be significantly improved using an ensemble of three Bi-Directional LSTMs (BiLSTM) that model the long-term, short-term and immediate effects of PM2.5 (the key air pollutant) severity levels. Further, weather information data has been taken into account while modelling, since they are found to boost prediction accuracies. Experimental results for multiple locations in New Delhi, India are presented to demonstrate model superiority over earlier techniques. Ishan Verma, Rahul Ahuja, Hardik Meisheri, Lipika Dey |
WI | 4 |
| 2017 | Exploring Linguistic and Graph Based Features for the Automatic Classification and Extraction of Adverse Drug Effects
Tirthankar Dasgupta, Abir Naskar, Lipika Dey |
CICLing (1) | 3 |
| 2017 | CrimeProfiler: crime information extraction and visualization from news mediaabstractNews articles from different sources regularly report crime incidents that contain details of crime, information about accused entities, details of the investigation process and finally details of judgement. In this paper, we have proposed natural language processing techniques for extraction and curation of crime-related information from digitally published News articles. We have leveraged computational linguistics based methods to analyse crime related News documents to extract different crime related entities and events. This includes name of the criminal, name of the victim, nature of crime, geographic location, date and time, and action taken against the criminal. We have also proposed a semi-supervised learning technique to learn different categories of crime events from the News documents. This helps in continuous evolution of the crime dictionaries. Thus the proposed methods are not restricted to detecting known crimes only but contribute actively towards maintaining an updated crime dictionary. We have done experiments with a collection of 3000 crime-reporting News articles. The end-product of our experiments is a crime-register that contains details of crime committed across geographies and time. This register can be further utilized for analytical and reporting purposes. Tirthankar Dasgupta, Abir Naskar, Rupsa Saha, Lipika Dey |
WI | 4 |
| 2017 | Detecting, quantifying and accessing impact of news events on Indian stock indicesabstractThe impact of different types of events reported in News articles on stock market is a widely accepted phenomenon. Market analysts rely heavily on technology to combine data from different sources and generate appropriate insights for predicting stock movements. With plethora of sources reporting news on plentitude of events happening across the world, a combination of text mining techniques and predictive technologies can play a significant role in this arena. In this paper we have presented methodologies to identify and quantify the presence of different types of information that can affect the market from a multitude of web sources, and finally use the information for predicting stock movement direction. We propose the use of PESTEL factors to categorize market-impacting information. We have analyzed large volumes of past available data using Granger causality to understand how these categories impact the market. We propose a paragraph-vector based information classification mechanism. We also present Long-Short term memory Network (LSTM) based prediction model to investigate the prediction capabilities of the information components. The proposed system outperforms state of the art linear SVM on data from different stock indices. Ishan Verma, Lipika Dey, Hardik Meisheri |
WI | 2 |
| 2016 | Enterprise risk analytics: Automatic analysis of risk factors from textual feedbacksabstractThere has been a growing need to automatically identify, extract and analyze risk related statements from textual data. In this paper, we have exploited natural language processing research to develop a risk analytics framework that processes human-reported risk statements to analyzes the enterprise risk description texts to classify them into valid and invalid risk categories, and perform analytics to extract information from the text pertaining to the different categories of risks and their possible cause and impacts. A manual annotation study from management experts using risk descriptions collected for a specific organization was conducted to evaluate the framework. The evaluation showed promising results for automated risk analysis and identification. Tirthankar Dasgupta, Lipika Dey |
SMC | 2 |
| 2015 | Interactive Visual Analysis of Temporal Text DataabstractThis paper presents a novel interactive visualization technique that helps in gathering insights from large volumes of text generated through dyadic communications. The emphasis is specifically on showing content evolution and modification with passage of time. The challenge lies in presenting not only the content as a stand-alone but also understand how the present is related to the past. For example analyzing large volumes of emails can show how communication among a set of people have progressed or evolved over time, may be along with the roles of the communicators. It can also show how the content has changed or evolved. In order to depict the changes, the email repositories are first clustered using a novel algorithm. The clusters are further time-stamped and correlated. User-insights are provided through visualization of these clusters. Results of implementation over two different datasets are presented. Aditeya Pandey, Kunal Ranjan, Geetika Sharma, Lipika Dey |
VINCI | 4 |
| 2013 | A framework to integrate unstructured and structured data for enterprise analytics
Lipika Dey, Ishan Verma, Arpit Khurdiya, Sameera Bharadwaja H. |
FUSION | 1 |
| 2013 | Email Analytics for Activity Management and Insight DiscoveryabstractEmails constitute the bulk of all official communications in any organization. Email repositories are tacit store-houses of knowledge about people, projects and processes. Mining one's own email repository can also provide interesting and valuable insights about his or her engagements and contacts along different dimensions. In this paper, we propose an email analytics framework that combines text-mining, network analysis and data analytics principles to mine email repositories for useful insights. While individuals are more attuned to looking at emails as individual items along with a history that is embedded in the trail, mining the whole collection can also lead to knowledge-discovery about similarities and dissimilarities of different engagements. This in turn can lead to valuable information like comparative status reports on various projects or deeper insights about why certain projects succeed while others don't. Given the volumes, diversity and noisy nature of e-mails, it becomes impossible for human beings to comprehend the impact of all of it unless the task is automated and approached in a structured fashion. We show that combination of text and network analytics along with temporal reasoning can provide valuable insights about task-states, actionable items, recommendations and forecasts. These insights can be exploited very effectively for project-management tasks like automated identification of bottlenecks or their causes, elimination of inefficiencies, early-warnings and suggestions about proactive measures to avoid problems. It is possible to extend the framework quite easily to analyze multiple email repositories of different users, though this work does not address the privacy or security concerns that might exist. Lipika Dey, Sameera Bharadwaja H., G. Meera, Gautam Shroff |
Web Intelligence | 1 |
| 2012 | Discovering regular and consistent behavioral patterns in topical tweeting
Lipika Dey, Bhakti Gaonkar |
ICPR | 1 |
| 2012 | TWIPIX: a web magazine curated from social mediaabstractThis paper describes a method to identify events being vigorously discussed on social media and to present them in the form of a web-based daily magazine. Tweet texts and hyperlinked information sources, such as images and news articles, are analyzed to discover the events. The events are selected for presentation using an "interestingness factor", which combines several facets of the discussions surrounding the events. The events are correlated based on their content similarity. Romil Bansal, Radhika Kumaran, Diwakar Mahajan, Arpit Khurdiya, Lipika Dey, Hiranmay Ghosh |
ACM Multimedia | 5 |
| 2012 | An Ontology-Based Mining of Consumer Feedbacks Using Fuzzy ReasoningabstractText analytics on consumer-generated content has gained significant momentum over last few years. A wide-range of text mining techniques has been proposed which can provide interesting insights about the text content. But, the challenge still exists in consuming the extracted information in form of actionable intelligence. Identifying actionable intelligence is difficult due to differences in consumer and business languages. Since feedbacks rarely talks of a single problem, determining the problems is also challenging. We propose a framework to address some of these challenges. Organizational websites or standard domain-ontologies are rich repositories of domain knowledge. The proposed method utilizes this knowledge to learn a discriminative classifier model for a domain using Fisher's discriminant metric. The consumer feedbacks are classified to different business categories using the learnt model. The output is further fed into a fuzzy reasoning unit where every feedback is assigned confidence values for each category. Initial experiments show that the proposed framework is capable of handling text feedbacks containing customer complaints in various domains. Lipika Dey, Sameera Bharadwaja H., Shefali Bhat |
Web Intelligence | 1 |
| 2012 | Extraction and Compilation of Events and Sub-events from TwitterabstractTwitter has emerged as a great source to provide insights about upcoming planned and unplanned events of social, economic and political relevance. Big events are publicized and known in advance, but smaller, unplanned sub-events around them are not always advertised. These unplanned events may have a large localized impact. If known in advance, knowledge about events like threats, protests, demonstrations etc. or even about large flash mobs can be utilized by planners and event managers. Given the large volumes of tweets floating around at any given time, identifying relevant sub-events is a non-trivial task. In this paper, we explore machine learning techniques to identify, extract and build a map of small sub-events around a big, popular event. We use CRFs to extract event components from tweets. Events are resolved for uniqueness and compiled into a complete calendar. The model is evaluated on tweets around Olympic Games. The framework is generic enough to be adapted to other domains. Arpit Khurdiya, Lipika Dey, Diwakar Mahajan, Ishan Verma |
Web Intelligence | 2 |
| 2012 | A rule-based method for identifying the factor structure in customer satisfaction
Amir Ahmad, Lipika Dey, Sami M. Halawani |
Inf. Sci. | 2 |
| 2011 | Enterprise information fusion for real-time business intelligence
Gautam Shroff, Puneet Agarwal, Lipika Dey |
FUSION | 3 |
| 2011 | Multi-Perspective Linking of News Articles within a Repository
Arpit Khurdiya, Lipika Dey, Nidhi Raj, S. K. Mirajul Haque |
IJCAI | 2 |
| 2011 | Expertise Prediction for Social Network Platforms to Encourage Knowledge SharingabstractKnowledge sharing social platforms where users mutually benefit through question-answering are gaining popularity. The success of these platforms on the web has led to their adoption within the firewalls of enterprises also. In this paper we have presented some in-depth study about two such platforms -- one open on the web and one which is within an enterprise to identify the similarities and dissimilarities of user behavior in the two platforms. We have proposed an algorithm to predict experts to improve the effectiveness of such platforms. Nidhi Raj, Lipika Dey, Bhakti Gaonkar |
Web Intelligence | 2 |
| 2011 | A k-means type clustering algorithm for subspace clustering of mixed numeric and categorical datasets
Amir Ahmad, Lipika Dey |
Pattern Recognit. Lett. | 2 |
| 2010 | Detection and Characterization of Anomalous Entities in Social Communication NetworksabstractSocial networks generated from emails or calls provide enormous geospatial and interaction information about subscribers. These have served as important inputs to intelligence analysts. In this paper, we propose an efficient algorithm for anomaly detection from social networks. Anomalous users are detected based on their behavioral dissimilarity from others. A rich feature set is proposed for outlier detection. A method for providing visual explanation for the results is also proposed. Nithi Gupta, Lipika Dey |
ICPR | 2 |
| 2010 | A concept-driven biomedical knowledge extraction and visualization framework for conceptualization of text corpora
Jahiruddin, Muhammad Abulaish, Lipika Dey |
J. Biomed. Informatics | 3 |
| 2009 | Opinion mining from noisy text data
Lipika Dey, S. K. Mirajul Haque |
Int. J. Document Anal. Recognit. | 1 |
| 2008 | Fuzzy ontologies for handling uncertainties and inconsistencies in domain knowledge descriptionabstractOntologies represent a method of formally expressing a shared understanding of information, and have paved the way for sharing concepts across applications in an unambiguous way. However, these ontologies are assumed to be hand-crafted, pre-defined structures with crisp concept descriptions and inter-concept relations. Crisp definitions however are not sufficient for real-world applications like ontology-based information extraction from unstructured text. In this paper we propose an enhancement of the ontology structure to a fuzzy ontology, which provides a mechanism to store imprecise concept definitions. The proposed fuzzy ontology framework can also help in ascertaining similarities and dissimilarities of concept definitions across distributed ontologies representing the same domain. Our design of fuzzy ontology is motivated by fuzzy set theoretic representation and reasoning, and is entirely different from other fuzzy ontology designs, which use only co-occurrence of concepts to determine their closeness. We have cited examples from various domains to show the necessity and capability of the structure in representing real-world knowledge. Lipika Dey, Muhammad Abulaish |
FUZZ-IEEE | 1 |
| 2008 | Mining Financial News for Major Events and Their Impacts on the MarketabstractIn this paper we have proposed a stock market analysis system that analyzes financial news items to identify and characterize major events that impact the market. The events have been identified using latent Dirichlet allocation (LDA) based topic extraction mechanism. These topics have been thereafter analyzed in conjunction with actual market data to understand their impact on the market. A prediction system has been proposed which can predict whether the stock market will fall or rise, based on news items. Anuj Mahajan, Lipika Dey, S. K. Mirajul Haque |
Web Intelligence | 2 |
| 2007 | Biological relation extraction and query answering from MEDLINE abstracts using ontology-based text mining
Muhammad Abulaish, Lipika Dey |
Data Knowl. Eng. | 2 |
| 2007 | A k-mean clustering algorithm for mixed numeric and categorical data
Amir Ahmad, Lipika Dey |
Data Knowl. Eng. | 2 |
| 2007 | A method to compute distance between two categorical values of same attribute in unsupervised learning for categorical data set
Amir Ahmad, Lipika Dey |
Pattern Recognit. Lett. | 2 |
| 2006 | Interoperability among Distributed Overlapping Ontologies - A Fuzzy Ontology FrameworkabstractOntologies are proposed as a means for knowledge sharing among applications but, it is often not possible to converge to a single unambiguous ontology that is acceptable to all knowledge engineers. Different ontologies vary greatly in terms of the level of detail of their representations, as well as the nature of their underlying logical specifications. Interoperability among different ontologies becomes essential to gain from the power of the existing domain ontologies. In this paper we have proposed a fuzzy ontology framework in which a concept descriptor is represented as a fuzzy relation which encodes the degree of a property value using a fuzzy membership function. Other than concept descriptors, the semantic relations in the ontology like IS-A, HAS-PART etc. are also associated a strength of association. The strength of association between two concepts determines the "uniformity" with which these two concepts have been defined identically across different ontologies. The fuzzy ontology framework provides appropriate support for application integration by identifying the most likely location of a particular term in the ontology Muhammad Abulaish, Lipika Dey |
Web Intelligence | 2 |
| 2006 | Generating Concept Ontologies through Text MiningabstractDesigning mechanisms for creating concept ontologies automatically is an important research problem. In this work we have proposed a rough-set based mechanism to generate concept ontologies with concepts mined from documents. When the concept ontology is mined from preclassified documents, the output signifies the core set of domain concepts and their inter-relationships that define the categories, as well as the inter-category relationships. When the ontology is mined from a heterogeneous collection, the documents are first clustered into homogeneous groups and then mined for concepts. Rough set based lower and upper approximations have been used to identify core concepts and associated concepts for a domain or a group. The scheme has been tested over multiple domains. Lipika Dey, Ashish Chandra Rastogi |
Web Intelligence | 1 |
| 2006 | Information extraction and imprecise query answering from web documents
Muhammad Abulaish, Lipika Dey |
Web Intell. Agent Syst. | 2 |
| 2005 | Biological Ontology Enhancement with Fuzzy Relations: A Text-Mining FrameworkabstractDomain ontology can help in information retrieval from documents. But ontology is a pre-defined structure with crisp concept descriptions and inter-concept relations. However, due to the dynamic nature of the document repository, ontology should be upgradeable with information extracted through text mining of documents in the domain. This also necessitates that concepts, their descriptions and inter-concept relations should be associated with a degree of fuzziness that will indicate the support for the extracted knowledge according to the currently available resources. Supports may be revised with more knowledge coming in future. This approach preserves the basic structured knowledge format for storing domain knowledge, but at the same time allows for update of information. In this paper, we have proposed a mechanism which initiates text mining with a set of ontological concepts, and thereafter extracts fuzzy relations through text mining. Membership values of relations are functions of frequency of co-occurrence of concepts and relations. We have worked on the GENIA corpus and shown how fuzzy relations can be further used for guided information extraction from MEDLINE documents. Muhammad Abulaish, Lipika Dey |
Web Intelligence | 2 |
| 2005 | A rough-fuzzy document grading system for customized text information retrieval
Shailendra Singh 0001, Lipika Dey |
Inf. Process. Manag. | 2 |
| 2005 | A feature selection technique for classificatory analysis
Amir Ahmad, Lipika Dey |
Pattern Recognit. Lett. | 2 |
| 2004 | Using Part-of-Speech Patterns and Domain Ontology to Mine Imprecise Concepts from Text Documents
Muhammad Abulaish, Lipika Dey |
iiWAS | 2 |
| 1999 | An MIMD algorithm for constant curvature feature extraction using curvature based data partitioning
Santanu Chaudhury, Anjana Roy, Lipika Dey |
Pattern Recognit. Lett. | 3 |