Min Song 0001

dblp:99/5721-1 · DBLP profile ↗
← Back
62ranked-venue papers
16as first author
10since 2021 · last 2025
0000-0003-3255-1600ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 37 · 12 first-author · 5 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 4 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Four decades of data & knowledge engineering: A bibliometric analysis and topic evolution study (1985-2024)
Tatsawan Timakum, Soobin Lee, Min Song 0001, Il-Yeol Song
Data Knowl. Eng.4
2025 CoTEL-D3X: A chain-of-thought enhanced large language model for drug-drug interaction triplet extraction
Haotian Hu, Alex Jie Yang, Sanhong Deng, Dongbo Wang, Min Song 0001
Expert Syst. Appl.5
2024 DOLAP: A 25 Year Journey Through Research Trends and Performance (Invited Paper)
Tatsawan Timakum, Soobin Lee, Haotian Hu, Il-Yeol Song, Min Song 0001
DOLAP5
2023 A method of constructing a fine-grained sentiment lexicon for the humanities computing of classical chinese poetry
Hao Wang 0194, Min Song 0001, Sanhong Deng
Neural Comput. Appl.3
2023 Heading Towards Sub-Discipline Rankings for Higher Education Institutions
abstract
Although the system for annual rankings of higher education institutions (HEIs) faces considerable criticism, these rankings are here to stay. Having become competent in assigning holistic ranking scores to HEIs, reputed ranking entities have now started focusing on subject-specific and regional rankings. However, in experts’ opinion, the process of assigning rankings should be more consistent, transparent, and representative. This study focuses on enhancing the credibility of the academic ranking process, by performing fine-grained assessment of the academic data pertaining to the computing discipline. The proposed assessment approach explores the data at the sub-discipline level, analyzing several ranking dimensions, including the research productivity, research impact, and research contribution of influential research scholars affiliated with renowned HEIs in the computing discipline. The analysis considers highly curated data published by three well-known international academic ranking entities, namely, Academic Rankings of World Universities (ARWU), Quacquarelli Symonds (QS), and Times Higher Education (THE), in 2018, 2019, and 2020, respectively. Researchers’ profiles are obtained from the Scopus repository, and the DBpedia repository is used to retrieve information about HEIs and their locations. For a stable comparison of the subject-specific academic rankings, the grand average rank measure is employed, whereas for finding the most influential researchers in computing, the ResRank measure is used. The sub-discipline-specific academic rankings provide more detailed insight into the academic rankings, thereby providing more robust decision support. This analysis, which focuses on the computing sub-discipline, is among the first few such efforts.
Muhammad Sajid Qureshi, Ali Daud, Malik Khizar Hayat, Min Song 0001, Yejin Park
IEEE Trans. Comput. Soc. Syst.4
2022 SQ2SV: Sequential Queries to Sequential Videos retrieval
abstract
Current video retrieval models are one-to-one matching models, which limits them from learning from the sequential context of the videos. While most public datasets for this task are text-video pairs that are contextually independent, datasets such as YouCook2, Video Storytelling, and COIN consist of chronological text-video pair segments. This paper introduces a retrieval task Sequential Queries to Sequential Videos retrieval (SQ2SV) that retrieves multiple sets of sequential videos from sequential queries to utilize such contextual interdependence. To the best of our knowledge, this paper is the first a ttempt to introduce multiple sets of sequential videos retrieval. We not only introduce a new task but also build a task-specific model and its evaluation metric. Our model, UniSeq (UniVL-based sequential videos retrieval), is a sequential as well as a cross representation model. Our new metric ‘Video R@k’ evaluates the performance of a retrieval model in a unit of video, not in a unit of the video segment. Our best model outperforms the UniVL baseline in the original R@1 of YouCook2 by 0.40% and Video Storytelling by 1.09%. Furthermore, comparing the Video R@1 score, our model outperforms the baseline by 0.27% for YouCook2 and 0.94% for Video Storytelling.
Injin Paek, Nayoung Choi, Seongjin Ha, Yuheun Kim, Min Song 0001
IEEE Big Data5
2022 MLM-based typographical error correction of unstructured medical texts for named entity recognition
abstract
BACKGROUND: Unstructured text in medical records, such as Electronic Health Records, contain an enormous amount of valuable information for research; however, it is difficult to extract and structure important information because of frequent typographical errors. Therefore, improving the quality of data with errors for text analysis is an essential task. To date, few prior studies have been conducted addressing this. Here, we propose a new methodology for extracting important information from unstructured medical texts by overcoming the typographical problem in surgical pathology records related to lung cancer. METHODS: We propose a typo correction model that considers context, based on the Masked Language Model, to solve the problem of typographical errors in real-world medical data. In addition, a word dictionary was used for the typo correction model based on PubMed abstracts. After refining the data through typo correction, fine tuning was performed on pre-trained BERT model. Next, deep learning-based Named Entity Recognition (NER) was performed. By solving the quality problem of medical data, we sought to improve the accuracy of information extraction in unstructured text data. RESULTS: We compared the performance of the proposed typo correction model based on contextual information with an existing SymSpell model. We confirmed that our proposed model outperformed the existing model in a typographical correction task. The F1-score of the model improved by approximately 5% and 9% when compared with the model without contextual information in the NCBI-disease and surgical pathology record datasets, respectively. In addition, the F1-score of NER after typo correction increased by 2% in the NCBI-disease dataset. There was a significant performance difference of approximately 25% between the before and after typo correction in the Surgical pathology record dataset. This confirmed that typos influenced the information extraction of the unstructured text. CONCLUSION: We verified that typographical errors in unstructured text negatively affect the performance of natural language processing tasks. The proposed method of a typo correction model outperformed the existing SymSpell model. This study shows that the proposed model is robust and can be applied in real-world environments by focusing on the typos that cause difficulties in analyzing unstructured medical text.
Eunbyul Lee, Go Eun Heo, Chang Min Choi, Min Song 0001
BMC Bioinform.4
2022 Pandemics are catalysts of scientific novelty: Evidence from COVID-19
abstract
Abstract Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First‐time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pretrained on 29 million PubMed articles, and first‐time collaboration increased after the outbreak of COVID‐19, and international collaboration witnessed a sudden decrease. During COVID‐19, papers with more first‐time collaboration were found to be more novel and international collaboration did not hamper novelty as it had done in the normal periods. The findings suggest the necessity of reaching out for distant resources and the importance of maintaining a collaborative scientific community beyond nationalism during a pandemic.
Meijun Liu, Yi Bu 0001, Chongyan Chen, Jian Xu 0003, Daifeng Li, Yan Leng, Richard B. Freeman 0002, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee, Jaewoo Kang, Min Song 0001, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.15
2021 Exploring the research landscape of data warehousing and mining based on DaWaK Conference full-text articles
Tatsawan Timakum, Soobin Lee, Min Song 0001
Data Knowl. Eng.3
2021 BioPREP: Deep learning-based predicate classification with SemMedDB
Gibong Hong, Yuheun Kim, YeonJung Choi, Min Song 0001
J. Biomed. Informatics4
2020 Analyzing the Research Landscape of DaWaK Papers from 1999 to 2019
Tatsawan Timakum, Soobin Lee, Il-Yeol Song, Min Song 0001
DaWaK4
2020 Literature based discovery of alternative TCM medicine for adverse reactions to depression drugs
abstract
BACKGROUND: In recent years, Traditional Chinese Medicine (TCM) and alternative medicine have been widely used along with western drugs as a complementary form of treatment. In this study, we first use the scientific literature to identify western drugs with obvious side effects. Then, we find TCM alternatives for these western drugs to ameliorate their side effects. RESULTS: We used depression as a case study. To evaluate our method, we showed the relation between herb-ingredients-target-disease for representative alternative herbs of western drugs. Further, a protein-protein interaction network of western drugs and alternative herbs was produced, and we performed enrichment analysis of the targets of the active ingredients of the herbs and examined the enrichment of Gene Ontology terms for Biological Process, Cellular Component, and Molecular Function and KEGG Pathway levels, to show how these targets affect different levels of gene expression. CONCLUSION: Our proposed method is able to select herbs that are highly relevant to the target indication (depression) and are able to treat the side effects caused by the target drug. The compounds from our selected alternative herbal medicines can therefore be complementary to the western drugs and ameliorate their side effects, which may help in the development of new drugs.
Qing Xie 0009, Kyoung Min Yang, Go Eun Heo, Min Song 0001
BMC Bioinform.4
2020 Characterizing user interest in NoSQL databases of social question and answer data
Minchul Lee, Sieun Jeon, Min Song 0001
J. Supercomput.3
2020 Developing a supervised learning-based social media business sentiment index
Hyeonseo Lee, Nakyeong Lee, Harim Seo, Min Song 0001
J. Supercomput.4
2019 AutoSense Model for Word Sense Induction
abstract
Word sense induction (WSI), or the task of automatically discovering multiple senses or meanings of a word, has three main challenges: domain adaptability, novel sense detection, and sense granularity flexibility. While current latent variable models are known to solve the first two challenges, they are not flexible to different word sense granularities, which differ very much among words, from aardvark with one sense, to play with over 50 senses. Current models either require hyperparameter tuning or nonparametric induction of the number of senses, which we find both to be ineffective. Thus, we aim to eliminate these requirements and solve the sense granularity problem by proposing AutoSense, a latent variable model based on two observations: (1) senses are represented as a distribution over topics, and (2) senses generate pairings between the target word and its neighboring word. These observations alleviate the problem by (a) throwing garbage senses and (b) additionally inducing fine-grained word senses. Results show great improvements over the stateof-the-art models on popular WSI datasets. We also show that AutoSense is able to learn the appropriate sense granularity of a word. Finally, we apply AutoSense to the unsupervised author name disambiguation task where the sense granularity problem is more evident and show that AutoSense is evidently better than competing models. We share our data and code here: https://github.com/rktamplayo/AutoSense.
Reinald Kim Amplayo, Seung-won Hwang, Min Song 0001
AAAI3
2019 An application of convolutional neural networks with salient features for relation classification
abstract
BACKGROUND: Due to the advent of deep learning, the increasing number of studies in the biomedical domain has attracted much interest in feature extraction and classification tasks. In this research, we seek the best combination of feature set and hyperparameter setting of deep learning algorithms for relation classification. To this end, we incorporate an entity and relation extraction tool, PKDE4J to extract biomedical features (i.e., biomedical entities, relations) for the relation classification. We compared the chosen Convolutional Neural Networks (CNN) based classification model with the most widely used learning algorithms. RESULTS: Our CNN based classification model outperforms the most widely used supervised algorithms. We achieved a significant performance on binary classification with a weighted macro-average F1-score: 94.79% using pre-extracted relevant feature combinations. For multi-class classification, the weighted macro-average F1-score is estimated around 86.95%. CONCLUSIONS: Our results suggest that our proposed CNN based model using the not only single feature as the raw text of the sentences of biomedical literature, but also coupling with multiple and highlighted features extracted from the biomedical sentences could improve the classification performance significantly. We offer hyperparameter tuning and optimization approaches for our proposed model to obtain optimal hyperparameters of the models with the best performance.
Zolzaya Dashdorj, Min Song 0001
BMC Bioinform.2
2018 Relation extraction for biological pathway construction using node2vec
abstract
BACKGROUND: Systems biology is an important field for understanding whole biological mechanisms composed of interactions between biological components. One approach for understanding complex and diverse mechanisms is to analyze biological pathways. However, because these pathways consist of important interactions and information on these interactions is disseminated in a large number of biomedical reports, text-mining techniques are essential for extracting these relationships automatically. RESULTS: In this study, we applied node2vec, an algorithmic framework for feature learning in networks, for relationship extraction. To this end, we extracted genes from paper abstracts using pkde4j, a text-mining tool for detecting entities and relationships. Using the extracted genes, a co-occurrence network was constructed and node2vec was used with the network to generate a latent representation. To demonstrate the efficacy of node2vec in extracting relationships between genes, performance was evaluated for gene-gene interactions involved in a type 2 diabetes pathway. Moreover, we compared the results of node2vec to those of baseline methods such as co-occurrence and DeepWalk. CONCLUSIONS: Node2vec outperformed existing methods in detecting relationships in the type 2 diabetes pathway, demonstrating that this method is appropriate for capturing the relatedness between pairs of biological entities involved in biological pathways. The results demonstrated that node2vec is useful for automatic pathway construction.
Munui Kim, Seung Han Baek, Min Song 0001
BMC Bioinform.3
2018 The landscape of smart aging: Topics, applications, and agenda
Il-Yeol Song, Min Song 0001, Tatsawan Timakum, Su-Ryeon Ryu, Hanju Lee
Data Knowl. Eng.2
2018 Network-based approach to detect novelty of scholarly literature
Reinald Kim Amplayo, SuLyn Hong, Min Song 0001
Inf. Sci.3
2018 Incorporating product description to sentiment topic models for improved aspect-based sentiment analysis
Reinald Kim Amplayo, Seanie Lee, Min Song 0001
Inf. Sci.3
2018 Topic diffusion analysis of a weighted citation network in biomedical literature
abstract
In this study, we propose a framework for detecting topic evolutions in weighted citation networks. Citation networks are important in studying knowledge flows; however, citation network analysis has primarily focused on binary networks in which the individual citation influences of each cited paper in a citing paper are considered identical, even though not all cited papers have a significant influence on the cited publication. Accordingly, it is necessary to build and analyze a citation network comprising scholarly publications that notably impact one another, thus identifying topic evolution in a more precise manner. To measure the strength of citation influence and identify paper topics, we employ a citation influence topic model primarily based on topical inheritance between cited and citing papers. Using scholarly publications in the field of the protein p53 as a case study, we build a citation network, filter it using citation influence values, and examine the diffusion of topics not only in the field but also in the subfields of p53.
Munui Kim, Injun Baek, Min Song 0001
J. Assoc. Inf. Sci. Technol.3
2018 Investigating drug-disease interactions in drug-symptom-disease triples via citation relations
abstract
With the growth in biomedical literature, the necessity of extracting useful information from the literature has increased. One approach to extracting biomedical knowledge involves using citation relations to discover entity relations. The assumption is that citation relations between any two articles connect knowledge entities across the articles, enabling the detection of implicit relationships among biomedical entities. The goal of this article is to examine the characteristics of biomedical entities connected via intermediate entities using citation relations aided by text mining. Based on the importance of symptoms as biomedical entities, we created triples connected via citation relations to identify drug–disease pairs with shared symptoms as intermediate entities. Drug–disease interactions built via citation relations were compared with co‐occurrence‐based interactions. Several types of analyses were adopted to examine the properties of the extracted entity pairs by comparing them with drug–disease interaction databases. We attempted to identify the characteristics of drug–disease pairs through citation relations in association with biomedical entities. The results showed that the citation relation‐based approach resulted in diverse types of biomedical entities and preserved topical consistency. In addition, drug–disease pairs identified only via citation relations are interesting for clinical trials when they are examined using BITOLA.
Min Song 0001, Keun Young Kang, Juyoung An
J. Assoc. Inf. Sci. Technol.1
2017 Analyzing the field of bioinformatics with the multi-faceted topic modeling technique
abstract
BACKGROUND: Bioinformatics is an interdisciplinary field at the intersection of molecular biology and computing technology. To characterize the field as convergent domain, researchers have used bibliometrics, augmented with text-mining techniques for content analysis. In previous studies, Latent Dirichlet Allocation (LDA) was the most representative topic modeling technique for identifying topic structure of subject areas. However, as opposed to revealing the topic structure in relation to metadata such as authors, publication date, and journals, LDA only displays the simple topic structure. METHODS: In this paper, we adopt the Tang et al.'s Author-Conference-Topic (ACT) model to study the field of bioinformatics from the perspective of keyphrases, authors, and journals. The ACT model is capable of incorporating the paper, author, and conference into the topic distribution simultaneously. To obtain more meaningful results, we use journals and keyphrases instead of conferences and bag-of-words.. For analysis, we use PubMed to collected forty-six bioinformatics journals from the MEDLINE database. We conducted time series topic analysis over four periods from 1996 to 2015 to further examine the interdisciplinary nature of bioinformatics. RESULTS: We analyze the ACT Model results in each period. Additionally, for further integrated analysis, we conduct a time series analysis among the top-ranked keyphrases, journals, and authors according to their frequency. We also examine the patterns in the top journals by simultaneously identifying the topical probability in each period, as well as the top authors and keyphrases. The results indicate that in recent years diversified topics have become more prevalent and convergent topics have become more clearly represented. CONCLUSION: The results of our analysis implies that overtime the field of bioinformatics becomes more interdisciplinary where there is a steady increase in peripheral fields such as conceptual, mathematical, and system biology. These results are confirmed by integrated analysis of topic distribution as well as top ranked keyphrases, authors, and journals.
Go Eun Heo, Keun Young Kang, Min Song 0001
BMC Bioinform.3
2017 An adaptable fine-grained sentiment analysis for summarization of multiple short online reviews
Reinald Kim Amplayo, Min Song 0001
Data Knowl. Eng.2
2017 Exploring characteristics of highly cited authors according to citation location and content
abstract
Big Science and cross‐disciplinary collaborations have reshaped the intellectual structure of research areas. A number of works have tried to uncover this hidden intellectual structure by analyzing citation contexts. However, none of them analyzed by document logical structures such as sections. The two major goals of this study are to find characteristics of authors who are highly cited section‐wise and to identify the differences in section‐wise author networks. This study uses 29,158 of research articles culled from the ACL Anthology, which hosts articles on computational linguistics and natural language processing. We find that the distribution of citations across sections is skewed and that a different set of highly cited authors share distinct academic characteristics, according to their citation locations. Furthermore, the author networks based on citation context similarity reveal that the intellectual structure of a domain differs across different sections.
Juyoung An, Namhee Kim, Min-Yen Kan, Muthu Kumar Chandrasekaran, Min Song 0001
J. Assoc. Inf. Sci. Technol.5
2017 Comparative evaluation of bibliometric content networks by tomographic content analysis: An application to Parkinson's disease
abstract
To understand the current state of a discipline and to discover new knowledge of a certain theme, one builds bibliometric content networks based on the present knowledge entities. However, such networks can vary according to the collection of data sets relevant to the theme by querying knowledge entities. In this study we classify three different bibliometric content networks. The primary bibliometric network is based on knowledge entities relevant to a keyword of the theme, the secondary network is based on entities associated with the lower concepts of the keyword, and the tertiary network is based on entities influenced by the theme. To explore the content and properties of these networks, we propose a tomographic content analysis that takes a slice‐and‐dice approach to analyzing the networks. Our findings indicate that the primary network is best suited to understanding the current knowledge on a certain topic, whereas the secondary network is good at discovering new knowledge across fields associated with the topic, and the tertiary network is appropriate for outlining the current knowledge of the topic and relevant studies.
Keeheon Lee, Su Yeon Kim, Erin Hea-Jin Kim, Min Song 0001
J. Assoc. Inf. Sci. Technol.4
2017 Ensemble analysis of topical journal ranking in bioinformatics
abstract
Journal rankings, frequently determined by the journal impact factor or similar indices, are quantitative measures for evaluating a journal's performance in its discipline, which is presently a major research thrust in the bibliometrics field. Recently, text mining was adopted to augment journal ranking‐based evaluation with the content analysis of a discipline taking a time‐variant factor into consideration. However, previous studies focused mainly on a silo analysis of a discipline using either citation‐or content‐oriented approaches, and no attempt was made to analyze topical journal ranking and its change over time in a seamless and integrated manner. To address this issue, we propose a journal‐time‐topic model, an extension of Dirichlet multinomial regression, which we applied to the field of bioinformatics to understand journal contribution to topics in a field and the shift of topic trends. The journal‐time‐topic model allows us to identify which journals are the major leaders in what topics and the manner in which their topical focus. It also helps reveal an interesting distinct pattern in the journal impact factor of high‐ and low‐ranked journals. The study results shed a new light for understanding topic specific journal rankings and shifts in journals' concentration on a subject.
Min Song 0001, Su Yeon Kim, Keeheon Lee
J. Assoc. Inf. Sci. Technol.1
2016 Author credit-assignment schemas: A comparison and analysis
abstract
Credit assignment to multiple authors of a publication is a challenging task owing to the conventions followed within different areas of research. In this study, we present a review of different author credit‐assignment schemas, which are designed mainly based on author position and the total number of coauthors on the publication. We implemented, tested, and classified 15 author credit‐assignment schemas into 3 types: linear, curve, and “other” assignment schemas. Further investigation and analysis revealed that most of the methods provide reasonable credit‐assignment results, even though the credit‐assignment distribution approaches are quite different among different types. The evaluation of each schema based on P ub M ed articles published in 2013 shows that there exist positive correlations among different schemas and that the similarity of credit‐assignment distributions can be derived from the similar design principles that stress the number of coauthors or the author position, or consider both. We provide a summary about the features of each credit‐assignment schema to facilitate the selection of the appropriate one, depending on the different conditions required to meet diverse needs.
Jian Xu 0003, Ying Ding 0001, Min Song 0001, Tamy Chambers
J. Assoc. Inf. Sci. Technol.3
2015 DTMBIO 2015: International Workshop on Data and Text Mining in Biomedical Informatics
abstract
Held each year in conjunction with one of the largest data management conferences, CIKM, the Ninth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO'15) is organized to bring together researchers interested in development and application of cutting-edge data management and analysis methods with a specific focus on applications in biology and medicine. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO'15 will help scientists understand emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research.
Min Song 0001, Doheon Lee, Karin Verspoor
CIKM1
2015 Identifying the topology of the K-pop video community on YouTube: A combined Co-comment analysis approach
abstract
YouTube is a successful social network that people use to upload, watch, and comment on videos. We believe comments left on these videos can provide insight into user interests, but to this point have not been used to map out a specific video community. Our study investigates whether and how user commenting behavior impacts the topology of the K‐pop video community through analysis of co‐commenting behavior on these videos. We apply a traditional author cocitation analysis to this behavior, in a process we refer to as co‐comment analysis, to detect the topology of this community. This involves: a) an analysis of user co‐comments to elicit the inclination of user homophily within the community; b) an analysis of user co‐comments, weighted frequency of co‐comments, to detect user interests in the community; and c) an analysis of user co‐comments, weighted sentiment scores, to capture user opinions by polarity. The results indicate that users who comment on specific K‐pop videos also tend to comment on topically similar YouTube videos. We also find that the number of comments made by users correlates with the degree of positivity of their comments. Conversely, users who comment negatively on K‐pop videos are not inclined to form specific user groups, but rather present only their opinions individually.
Min Song 0001, Yoo Kyung Jeong, Ha Jin Kim
J. Assoc. Inf. Sci. Technol.1
2015 PKDE4J: Entity and relation extraction for public knowledge discovery
Min Song 0001, Won Chul Kim, Dahee Lee, Go Eun Heo, Keun Young Kang
J. Biomed. Informatics1
2014 DTMBIO 2014: International Workshop on Data and Text Mining in Biomedical Informatics
abstract
Held each year in conjunction with one of the largest data management conferences, CIKM, the Eighth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 14) is organized to bring together researchers interested in development and application of cutting-edge biomedical and healthcare technology. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 14 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research.
Luonan Chen, Doheon Lee, Hua Xu 0001, Min Song 0001
CIKM4
2014 Mapping biological entities using the Longest Approximately Common Prefix method
abstract
BACKGROUND: The significant growth in the volume of electronic biomedical data in recent decades has pointed to the need for approximate string matching algorithms that can expedite tasks such as named entity recognition, duplicate detection, terminology integration, and spelling correction. The task of source integration in the Unified Medical Language System (UMLS) requires considerable expert effort despite the presence of various computational tools. This problem warrants the search for a new method for approximate string matching and its UMLS-based evaluation. RESULTS: This paper introduces the Longest Approximately Common Prefix (LACP) method as an algorithm for approximate string matching that runs in linear time. We compare the LACP method for performance, precision and speed to nine other well-known string matching algorithms. As test data, we use two multiple-source samples from the Unified Medical Language System (UMLS) and two SNOMED Clinical Terms-based samples. In addition, we present a spell checker based on the LACP method. CONCLUSIONS: The Longest Approximately Common Prefix method completes its string similarity evaluations in less time than all nine string similarity methods used for comparison. The Longest Approximately Common Prefix outperforms these nine approximate string matching methods in its Maximum F1 measure when evaluated on three out of the four datasets, and in its average precision on two of the four datasets.
Alex Rudniy, Min Song 0001, James Geller
BMC Bioinform.2
2014 Content-based citation analysis: The next generation of citation analysis
abstract
Traditional citation analysis has been widely applied to detect patterns of scientific collaboration, map the landscapes of scholarly disciplines, assess the impact of research outputs, and observe knowledge transfer across domains. It is, however, limited, as it assumes all citations are of similar value and weights each equally. Content‐based citation analysis ( CCA ) addresses a citation's value by interpreting each one based on its context at both the syntactic and semantic levels. This paper provides a comprehensive overview of CAA research in terms of its theoretical foundations, methodical approaches, and example applications. In addition, we highlight how increased computational capabilities and publicly available full‐text resources have opened this area of research to vast possibilities, which enable deeper citation analysis, more accurate citation prediction, and increased knowledge discovery.
Ying Ding 0001, Guo Zhang 0007, Tamy Chambers, Min Song 0001, Xiaolong Wang 0009, ChengXiang Zhai
J. Assoc. Inf. Sci. Technol.4
2014 Productivity and influence in bioinformatics: A bibliometric analysis using PubMed central
abstract
Bioinformatics is a fast‐growing field based on the optimal use of “big data” gathered in genomic, proteomics, and functional genomics research. In this paper, we conduct a comprehensive and in‐depth bibliometric analysis of the field of bioinformatics by extracting citation data from PubMed Central full‐text. Citation data for the period 2000 to 2011, comprising 20,869 papers with 546,245 citations, was used to evaluate the productivity and influence of this emerging field. Four measures were used to identify productivity; most productive authors, most productive countries, most productive organizations, and most popular subject terms. Research impact was analyzed based on the measures of most cited papers, most cited authors, emerging stars, and leading organizations. Results show the overall trends between the periods 2000 to 2003 and 2004 to 2007 were dissimilar, while trends between the periods 2004 to 2007 and 2008 to 2011 were similar. In addition, the field of bioinformatics has undergone a significant shift, co‐evolving with other biomedical disciplines.
Min Song 0001, Su Yeon Kim, Guo Zhang 0007, Ying Ding 0001, Tamy Chambers
J. Assoc. Inf. Sci. Technol.1
2013 DTMBIO 2013: international workshop on data and text mining in biomedical informatics
abstract
The organizers of ACM Seventh International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 13) are pleased to announce that the seventh DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 13 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research.
Atul J. Butte, Doheon Lee, Hua Xu 0001, Min Song 0001
CIKM4
2013 Workshop summary for the 2013 international workshop on mining unstructured big data using natural language processing
abstract
No abstract available.
Xiaozhong Liu 0001, Ying Ding 0001, Min Song 0001
CIKM4
2013 Seeking beyond with IntegraL: A user study of sense-making enabled by anchor-based virtual integration of library systems
abstract
This article presents a user study showing the effectiveness of a linked‐based, virtual integration infrastructure that gives users access to relevant online resources, empowering them to design an information‐seeking path that is specifically relevant to their context. IntegraL provides a lightweight approach to improve and augment search functionality by dynamically generating context‐focused “anchors” for recognized elements of interest generated by library services. This article includes a description of how IntegraL's design supports users' information‐seeking behavior. A full user study with both objective and subjective measures of IntegraL and hypothesis testing regarding IntegraL's effectiveness of the user's information‐seeking experience are described along with data analysis, implications arising from this kind of virtual integration, and possible future directions.
Shuyuan Mary Ho, Michael Bieber, Min Song 0001, Xiangmin Zhang
J. Assoc. Inf. Sci. Technol.3
2013 Understanding the evolution of multiple scientific research domains using a content and network approach
abstract
Interdisciplinary research has been attracting more attention in recent decades. In this article, we compare the similarity between scientific research domains and quantifying the temporal similarities of domains. We narrowed our study to three research domains: information retrieval (IR), database (DB), and World Wide Web (W3), because the rapid development of the W3 domain substantially attracted research efforts from both IR and DB domains and introduced new research questions to these two areas. Most existing approaches either employed a content‐based technique or a cocitation or coauthorship network‐based technique to study the development trend of a research area. In this work, we proposed an effective way to quantify the similarities among different research domains by incorporating content similarity and coauthorship network similarity. Experimental results on DBLP (DataBase systems and Logic Programming) data related to IR, DB, and W3 domains showed that the W3 domain was getting closer to both IR and DB whereas the distance between IR and DB remained relatively constant. In addition, comparing to IR and W3 with the DB domain, the DB domain was more conservative and evolved relatively slower.
Xuning Tang, Christopher C. Yang, Min Song 0001
J. Assoc. Inf. Sci. Technol.3
2013 Text Categorization of Biomedical Data Sets Using Graph Kernels and a Controlled Vocabulary
abstract
Recently, graph representations of text have been showing improved performance over conventional bag-of-words representations in text categorization applications. In this paper, we present a graph-based representation for biomedical articles and use graph kernels to classify those articles into high-level categories. In our representation, common biomedical concepts and semantic relationships are identified with the help of an existing ontology and are used to build a rich graph structure that provides a consistent feature set and preserves additional semantic information that could improve a classifier's performance. We attempt to classify the graphs using both a set-based graph kernel that is capable of dealing with the disconnected nature of the graphs and a simple linear kernel. Finally, we report the results comparing the classification performance of the kernel classifiers to common text-based classifiers.
Said Bleik, Meenakshi Mishra, Jun Huan, Min Song 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 DTMBIO 2012: international workshop on data and text mining in biomedical informatics
abstract
The organizers of ACM Sixth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 12) are happy announce that the sixth DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 12 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research.
Min Song 0001, Doheon Lee, Hua Xu 0001, Sophia Ananiadou
CIKM1
2011 DTMBIO 2011: international workshop on data and textmining in biomedical informatics
abstract
ACM Fifth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 11) organizers are pleased to announce that the fifth DTMBIO will be held in conjunction with CIKM, one of the largest data and text mining conferences. While CIKM presents the state-of-the-art research in informatics with the primary focus on data and text mining, the main focus of DTMBIO is on biomedical informatics. DTMBIO delegates will bring forth interesting applications of up-to-date informatics in the context of biomedical research.
Sophia Ananiadou, Doheon Lee, Shamkant B. Navathe, Min Song 0001
CIKM4
2011 Novel Recommendation Based on Personal Popularity Tendency
abstract
Recently, novel recommender systems have attracted considerable attention in the research community. Recommending popular items may not always satisfy users. For example, although most users likely prefer popular items, such items are often not very surprising or novel because users may already know about the items. Also, such recommender systems hardly satisfy a group of users who prefer relatively obscure items. Existing novel recommender systems, however, still recommend mainly popular items or degrade the quality of recommendation. They do so because they do not consider the balance between novelty and preference-based recommendation. This paper proposes an efficient novel-recommendation method called Personal Popularity Tendency Matching (PPTM) which recommends novel items by considering an individual's Personal Popularity Tendency (or PPT). Considering PPT helps to diversify recommendations by reasonably penalizing popular items while improving the recommendation accuracy. We experimentally show that the proposed method, PPTM, is better than other methods in terms of both novelty and accuracy.
Jinoh Oh, Sun Park, Hwanjo Yu, Min Song 0001, Seung-Taek Park
ICDM4
2011 Building the process-drug-side effect network to discover the relationship between biological Processes and side effects
abstract
BACKGROUND: Side effects are unwanted responses to drug treatment and are important resources for human phenotype information. The recent development of a database on side effects, the side effect resource (SIDER), is a first step in documenting the relationship between drugs and their side effects. It is, however, insufficient to simply find the association of drugs with biological processes; that relationship is crucial because drugs that influence biological processes can have an impact on phenotype. Therefore, knowing which processes respond to drugs that influence the phenotype will enable more effective and systematic study of the effect of drugs on phenotype. To the best of our knowledge, the relationship between biological processes and side effects of drugs has not yet been systematically researched. METHODS: We propose 3 steps for systematically searching relationships between drugs and biological processes: enrichment scores (ES) calculations, t-score calculation, and threshold-based filtering. Subsequently, the side effect-related biological processes are found by merging the drug-biological process network and the drug-side effect network. Evaluation is conducted in 2 ways: first, by discerning the number of biological processes discovered by our method that co-occur with Gene Ontology (GO) terms in relation to effects extracted from PubMed records using a text-mining technique and second, determining whether there is improvement in performance by limiting response processes by drugs sharing the same side effect to frequent ones alone. RESULTS: The multi-level network (the process-drug-side effect network) was built by merging the drug-biological process network and the drug-side effect network. We generated a network of 74 drugs-168 side effects-2209 biological process relation resources. The preliminary results showed that the process-drug-side effect network was able to find meaningful relationships between biological processes and side effects in an efficient manner. CONCLUSIONS: We propose a novel process-drug-side effect network for discovering the relationship between biological processes and side effects. By exploring the relationship between drugs and phenotypes through a multi-level network, the mechanisms underlying the effect of specific drugs on the human body may be understood.
Sejoon Lee, Kwang Hyung Lee, Min Song 0001, Doheon Lee
BMC Bioinform.3
2011 Combining active learning and semi-supervised learning techniques to extract protein interaction sentences
abstract
BACKGROUND: Protein-protein interaction (PPI) extraction has been a focal point of many biomedical research and database curation tools. Both Active Learning and Semi-supervised SVMs have recently been applied to extract PPI automatically. In this paper, we explore combining the AL with the SSL to improve the performance of the PPI task. METHODS: We propose a novel PPI extraction technique called PPISpotter by combining Deterministic Annealing-based SSL and an AL technique to extract protein-protein interaction. In addition, we extract a comprehensive set of features from MEDLINE records by Natural Language Processing (NLP) techniques, which further improve the SVM classifiers. In our feature selection technique, syntactic, semantic, and lexical properties of text are incorporated into feature selection that boosts the system performance significantly. RESULTS: By conducting experiments with three different PPI corpuses, we show that PPISpotter is superior to the other techniques incorporated into semi-supervised SVMs such as Random Sampling, Clustering, and Transductive SVMs by precision, recall, and F-measure. CONCLUSIONS: Our system is a novel, state-of-the-art technique for efficiently extracting protein-protein interaction pairs.
Min Song 0001, Hwanjo Yu, Wook-Shin Han
BMC Bioinform.1
2010 Biomedical concept extraction using concept graphs and ontology-based mapping
abstract
Assigning keywords to articles can be extremely costly. In this paper we propose a new approach to biomedical concept extraction using semantic features of concept graphs to help in automatic labeling of scientific publications. The proposed system extracts key concepts similar to author-provided keywords. We represent full-text documents by graphs and map biomedical terms to predefined ontology concepts. In addition to occurrence frequency weights, we use concept relation weights to rank potential key concepts. We compare our technique to that of KEA's, a state-of-the-art keyphrase extraction software. The results show that using the relations weight significantly improves the performance of concept extraction. The results also highlight the subjectivity of the concept extraction procedure as well as of its evaluation.
Said Bleik, Yiran Wang 0002, Min Song 0001
BIBM4
2010 DTMBIO workshop summary
abstract
No abstract available.
Hagit Shatkay, Doheon Lee, Min Song 0001, Shamkant B. Navathe
CIKM3
2010 MKEM: a Multi-level Knowledge Emergence Model for mining undiscovered public knowledge
abstract
BACKGROUND: Since Swanson proposed the Undiscovered Public Knowledge (UPK) model, there have been many approaches to uncover UPK by mining the biomedical literature. These earlier works, however, required substantial manual intervention to reduce the number of possible connections and are mainly applied to disease-effect relation. With the advancement in biomedical science, it has become imperative to extract and combine information from multiple disjoint researches, studies and articles to infer new hypotheses and expand knowledge. METHODS: We propose MKEM, a Multi-level Knowledge Emergence Model, to discover implicit relationships using Natural Language Processing techniques such as Link Grammar and Ontologies such as Unified Medical Language System (UMLS) MetaMap. The contribution of MKEM is as follows: First, we propose a flexible knowledge emergence model to extract implicit relationships across different levels such as molecular level for gene and protein and Phenomic level for disease and treatment. Second, we employ MetaMap for tagging biological concepts. Third, we provide an empirical and systematic approach to discover novel relationships. RESULTS: We applied our system on 5000 abstracts downloaded from PubMed database. We performed the performance evaluation as a gold standard is not yet available. Our system performed with a good precision and recall and we generated 24 hypotheses. CONCLUSIONS: Our experiments show that MKEM is a powerful tool to discover hidden relationships residing in extracted entities that were represented by our Substance-Effect-Process-Disease-Body Part (SEPDB) model.
Ali Zeeshan Ijaz, Min Song 0001, Doheon Lee
BMC Bioinform.2
2010 Detecting duplicate biological entities using Markov random field-based edit distance
Min Song 0001, Alex Rudniy
Knowl. Inf. Syst.1
2009 Fast max-margin clustering for unsupervised word sense disambiguation in biomedical texts
abstract
BACKGROUND: We aim to solve the problem of determining word senses for ambiguous biomedical terms with minimal human effort. METHODS: We build a fully automated system for Word Sense Disambiguation by designing a system that does not require manually-constructed external resources or manually-labeled training examples except for a single ambiguous word. The system uses a novel and efficient graph-based algorithm to cluster words into groups that have the same meaning. Our algorithm follows the principle of finding a maximum margin between clusters, determining a split of the data that maximizes the minimum distance between pairs of data points belonging to two different clusters. RESULTS: On a test set of 21 ambiguous keywords from PubMed abstracts, our system has an average accuracy of 78%, outperforming a state-of-the-art unsupervised system by 2% and a baseline technique by 23%. On a standard data set from the National Library of Medicine, our system outperforms the baseline by 6% and comes within 5% of the accuracy of a supervised system. CONCLUSION: Our system is a novel, state-of-the-art technique for efficiently finding word sense clusters, and does not require training data or human effort for each new word to be disambiguated.
Weisi Duan, Min Song 0001, Alexander Yates
BMC Bioinform.2
2008 Detecting Duplicate Biological Entities Using Markov Random Field-Based Edit Distance
abstract
Duplicate entities detection in biological data became a demanded research task. In this paper, we propose a novel context-sensitive Markov random field-based edit distance. We apply the Markov random field theory to Needleman-Wunsch distance and combine MRFED with TFIDF, a token-based distance algorithm (SoftMRFED). We evaluate SoftMRFED and other distance algorithms (Levenstein, SoftTFIDF, and MongeElkan) at biological entity matching and synonym matching. The experiment results show SoftMRFED significantly outperforms other distance algorithms and its performance is superior to token-based distance algorithms in two matching tasks.
Min Song 0001, Alex Rudniy
BIBM1
2008 Document Clustering by Semantic Smoothing and Dynamic Growing Cell Structure (DynGCS) for Biomedical Literature
Min Song 0001, Xiaohua Hu 0001, Illhoi Yoo, Eric Koppel
DaWaK1
2007 A Hybrid Abbreviation Extraction Technique for Biomedical Literature
abstract
In this paper, we propose a novel technique to extract abbreviation combining natural language processing techniques and the Support Vector Machine (SVM) in biomedical literature. The proposed technique gives us the comparative advantages over others in the following aspects: 1) It incorporates lexical analysis techniques to supervised learning for extracting abbreviations. 2) It makes use of text chunking techniques to identify long forms of abbreviations. 3) It significantly improves Recall compared to other techniques. The experimental results show that our approach outperforms the leading abbreviation algorithms, Extract Abbrev, ALICE, and Acrophile, at least by 6% 13.9%, and 13.2% respectively, in both Precision and Recall on the Gold Standard Development corpus.
Min Song 0001, Illhoi Yoo
BIBM1
2007 A comparison study on algorithms of detecting long forms for short forms in biomedical text
abstract
MOTIVATION: With more and more research dedicated to literature mining in the biomedical domain, more and more systems are available for people to choose from when building literature mining applications. In this study, we focus on one specific kind of literature mining task, i.e., detecting definitions of acronyms, abbreviations, and symbols in biomedical text. We denote acronyms, abbreviations, and symbols as short forms (SFs) and their corresponding definitions as long forms (LFs). The study was designed to answer the following questions; i) how well a system performs in detecting LFs from novel text, ii) what the coverage is for various terminological knowledge bases in including SFs as synonyms of their LFs, and iii) how to combine results from various SF knowledge bases. METHOD: We evaluated the following three publicly available detection systems in detecting LFs for SFs: i) a handcrafted pattern/rule based system by Ao and Takagi, ALICE, ii) a machine learning system by Chang et al., and iii) a simple alignment-based program by Schwartz and Hearst. In addition, we investigated the conceptual coverage of two terminological knowledge bases: i) the UMLS (the Unified Medical Language System), and ii) the BioThesaurus (a thesaurus of names for all UniProt protein records). We also implemented a web interface that provides a virtual integration of various SF knowledge bases. RESULTS: We found that detection systems agree with each other on most cases, and the existing terminological knowledge bases have a good coverage of synonymous relationship for frequently defined LFs. The web interface allows people to detect SF definitions from text and to search several SF knowledge bases. AVAILABILITY: The web site is http://gauss.dbb.georgetown.edu/liblab/SFThesaurus.
Manabu Torii, Zhang-Zhi Hu, Min Song 0001, Cathy H. Wu
BMC Bioinform.3
2007 Integration of association rules and ontologies for semantic query expansion
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
Data Knowl. Eng.1
2006 Automatic Extraction for Creating a Lexical Repository of Abbreviations in the Biomedical Literature
Min Song 0001, Il-Yeol Song, Ki Jung Lee
DaWaK1
2006 A semi-supervised efficient learning approach to extract biological relationships from web-based biomedical digital library
Xiaohua Hu 0001, Tsau Young Lin, Il-Yeol Song, Illhoi Yoo, Min Song 0001
Web Intell. Agent Syst.5
2005 Mining undiscovered public knowledge from complementary and non-interactive biomedical literature through semantic pruning
abstract
Two complementary and non-interactive literature sets of articles, when they are considered together, can reveal useful information of scientific interest not apparent in either of the two document sets. Swanson called the existence of such knowledge, undiscovered public knowledge (UDPK). This paper proposes a semantic-based mining model for UDPK. Our method replaces manual ad-hoc pruning with using semantic knowledge from the biomedical ontologies. Using the semantic types and semantic relationships of the biomedical concepts, our prototype system can identify the relevant concepts collected from Medline and generate the novel hypothesis between these concepts. The system successfully replicates Swanson's two famous discoveries: Raynaud disease/fish oils and migraine/magnesium. Compared with previous approaches, our methods generate much fewer but more relevant novel hypotheses, and require much less human intervention in the discovery procedure.
Xiaohua Hu 0001, Illhoi Yoo, Min Song 0001, Yan-Qing Zhang 0001, Il-Yeol Song
CIKM3
2005 Semantic Query Expansion Combining Association Rules with Ontologies and Information Retrieval Techniques
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
DaWaK1
2005 An Automatic Unsupervised Querying Algorithm for Efficient Information Extraction in Biomedical Domain
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen
PAKDD1
2004 Extracting and mining protein-protein interaction network from biomedical literature
abstract
We present a biomedical literature data mining system SPIE-DM (Scalable and Portable Information Extraction and Data Mining) to extract and mine the protein-protein interaction network from biomedical literature such as MedLine. SPIE-DM consists of two phases: in phase 1, we develop a scalable and portable ie method (SPIE) to extract the protein-protein interaction from the biomedical literature. These extracted protein-protein interactions form a scale-free network graph. In phase 2, we apply a novel clustering method SFCluster to mine the protein-protein interaction network. The clusters in the network graph represent some potential protein complexes, which are very important for biologist to study the protein functionality. The clustering algorithm considers the characteristics of the scale-free network graphs and is based on the local density of the vertex and its neighborhood functions that can be used to find more meaningful clusters at different density levels. The experiments of SPIE-DM on around 1600 chromatin proteins indicate that our system is very promising for extracting and mining from biomedical literature databases.
Xiaohua Hu 0001, Illhoi Yoo, Il-Yeol Song, Min Song 0001, Jianchao Han, Mark Lechner
CIBCB4
2004 Ontology-Based Scalable and Portable Information Extraction System to Extract Biological Knowledge from Huge Collection of Biomedical Web Documents
abstract
Automated discovery and extraction of biological knowledge from biomedical web documents has become essential because of the enormous amount of biomedical literature published each year. In this paper we present an ontology-based scalable and portable information extraction system to automatically extract biological knowledge from huge collection of online biomedical web documents. Our method integrates ontology-based semantic tagging, information extraction and data mining together, automatically learns the patterns based on a few user seed tuples, and then extract new tuples from the biomedical web documents based on the discovered patterns. A novel system SPIE (Scalable and Portable Information Extraction) is implemented and tested on the PuBMed to find the chromatin protein-protein interaction and the experimental results indicate our approach is very effective in extracting biological knowledge from huge collection of biomedical web documents.
Xiaohua Hu 0001, Tsau Young Lin, Il-Yeol Song, Xia Lin, Illhoi Yoo, Mark Lechner, Min Song 0001
Web Intelligence7