EDBT 2026 Demo / reviewers in the wild / expert
Erjia Yan
dblp:47/7380
· DBLP profile ↗
32ranked-venue papers in the field
13as first author
4since 2021 · last 2024
0000-0002-0365-9340ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 27 (13 first)Big Data, Cloud & Distributed Data Systems · 3Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | To move or to be promoted: Examining the effect of promotions and academic mobility on professors' productivity and impactabstractAbstract Promotions and academic mobility are trajectory‐altering events in a researcher's career. This paper compiles a unique large data set and investigates publication and citation differences between two groups of researchers: the ones who are mobile and their counterparts who stay at a university with a promotion. This paper finds that mobile researchers often have a lesser productivity increase than their post‐promotion counterparts. The difference is largely driven by male professors in physical science and clinical health fields moving from more research‐intensive to less research‐intensive institutions. In contrast, the citational impact differences between the two groups are largely minimal. Chaojiang Wu, Erjia Yan, Chaoqun Ni, Jiangen He |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2023 | Gender and country biases in Wikipedia citations to scholarly publicationsabstractAbstract Ensuring Wikipedia cites scholarly publications based on quality and relevancy without biases is critical to credible and fair knowledge dissemination. We investigate gender‐ and country‐based biases in Wikipedia citation practices using linked data from the Web of Science and a Wikipedia citation dataset. Using coarsened exact matching, we show that publications by women are cited less by Wikipedia than expected, and publications by women are less likely to be cited than those by men. Scholarly publications by authors affiliated with non‐Anglosphere countries are also disadvantaged in getting cited by Wikipedia, compared with those by authors affiliated with Anglosphere countries. The level of gender‐ or country‐based inequalities varies by research field, and the gender‐country intersectional bias is prominent in math‐intensive STEM fields. To ensure the credibility and equality of knowledge presentation, Wikipedia should consider strategies and guidelines to cite scholarly publications independent of the gender and country of authors. Jiajing Chen, Erjia Yan, Chaoqun Ni |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2021 | Citation cascade and the evolution of topic relevanceabstractAbstract Citation analysis, as a tool for quantitative studies of science, has long emphasized direct citation relations, leaving indirect or high‐order citations overlooked. However, a series of early and recent studies demonstrate the existence of indirect and continuous citation impact across generations. Adding to the literature on high‐order citations, we introduce the concept of a citation cascade: the constitution of a series of subsequent citing events initiated by a certain publication. We investigate this citation structure by analyzing more than 450,000 articles and over 6 million citation relations. We show that citation impact exists not only within the three generations documented in prior research but also in much further generations. Still, our experimental results indicate that two to four generations are generally adequate to trace a work's scientific impact. We also explore specific structural properties—such as depth, width, structural virality, and size—which account for differences among individual citation cascades. Finally, we find evidence that it is more important for a scientific work to inspire trans‐domain (or indirectly related domain) works than to receive only intradomain recognition in order to achieve high impact. Our methods and findings can serve as a new tool for scientific evaluation and the modeling of scientific history. Erjia Yan, Yi Bu 0001 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2021 | Gender imbalance in the productivity of funded projects: A study of the outputs of National Institutes of Health R01 grantsabstractAbstract This study examines the relationship between team's gender composition and outputs of funded projects using a large data set of National Institutes of Health (NIH) R01 grants and their associated publications between 1990 and 2017. This study finds that while the women investigators' presence in NIH grants is generally low, higher women investigator presence is on average related to slightly lower number of publications. This study finds empirically that women investigators elect to work in fields in which fewer publications per million‐dollar funding is the norm. For fields where women investigators are relatively well represented, they are as productive as men. The overall lower productivity of women investigators may be attributed to the low representation of women in high productivity fields dominated by men investigators. The findings shed light on possible reasons for gender disparity in grant productivity. Chaojiang Wu, Erjia Yan, Yongjun Zhu 0001, Kai Li 0010 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | Authors' status and the perceived quality of their work: Measuring citation sentiment change in nobel articlesabstractPrior research in status ordering has used numeric indicators to examine the impact of a status change on the perception of a scientist's work. This study measures the perception change directly as reflected in citation sentiment, with the attainment of a Nobel Prize in Chemistry or a Nobel Prize in Physiology or Medicine considered the status change. The article identifies 12,393 citances to 25 Nobel articles in PubMed Central and includes a control article set of 75 articles with 30,851 citances. The results show a moderate increase in citation sentiment toward Nobel articles postaward. Dynamically, for Nobel articles there is a steady sentiment increase, and a Nobel Prize seems to co‐occur with this trend. This trend, however, is not evident in the control article set. Erjia Yan, Zheng Chen 0010, Kai Li 0010 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2018 | Correlated Anomaly Detection from Large Streaming DataabstractCorrelated anomaly detection (CAD) from streaming data is a type of group anomaly detection and an essential task in useful real-time data mining applications like botnet detection, financial event detection, industrial process monitor, etc. The primary approach for this type of detection in previous researches is based on principal score (PS) of divided batches or sliding windows by computing top eigenvalues of the correlation matrix, e.g. the Lanczos algorithm. However, this paper brings up the phenomenon of principal score degeneration for large data set, and then mathematically and practically prove current PS-based methods are likely to fail for CAD on large-scale streaming data even if the number of correlated anomalies grows with the data size at a reasonable rate; in reality, anomalies tend to be the minority of the data, and this issue can be more serious. We propose a framework with two novel randomized algorithms rPS and gPS for better detection of correlated anomalies from large streaming data of various correlation strength. The experiment shows high and balanced recall and estimated accuracy of our framework for anomaly detection from a large server log data set and a U.S. stock daily price data set in comparison to direct principal score evaluation and some other recent group anomaly detection algorithms. Moreover, our techniques significantly improve the computation efficiency and scalability for principal score calculation. Zheng Chen 0010, Xinli Yu 0002, Yuan Ling, Xiaohua Hu 0001, Erjia Yan |
IEEE BigData | 7 |
| 2018 | Which domains do open-access journals do best in? A 5-year longitudinal studyabstractAlthough researchers have begun to investigate the difference in scientific impact between closed‐access and open‐access journals, studies that focus specifically on dynamic and disciplinary differences remain scarce. This study serves to fill this gap by using a large longitudinal dataset to examine these differences. Using CiteScore as a proxy for journal scientific impact, we employ a series of statistical tests to identify the quartile categories and disciplinary areas in which impact trends differ notably between closed‐ and open‐access journals. We find that closed‐access journals have a noticeable advantage in social sciences (for example, business and economics), whereas open‐access journals perform well in medical and healthcare domains (for example, health profession and nursing). Moreover, we find that after controlling for a journal's rank and disciplinary differences, there are statistically more closed‐access journals in the top 10%, Quartile 1, and Quartile 2 categories as measured by CiteScore; in contrast, more open‐access journals in Quartile 4 gained scientific impact from 2011 to 2015. Considering dynamic and disciplinary trends in tandem, we find that more closed‐access journals in Social Sciences gained in impact, whereas in biochemistry and medicine, more open‐access journals experienced such gains. Erjia Yan, Kai Li 0010 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2018 | Data set mentions and citations: A content analysis of full-text publicationsabstractThis study provides evidence of data set mentions and citations in multiple disciplines based on a content analysis of 600 publications in PLoS One. We find that data set mentions and citations varied greatly among disciplines in terms of how data sets were collected, referenced, and curated. While a majority of articles provided free access to data, formal ways of data attribution such as DOIs and data citations were used in a limited number of articles. In addition, data reuse took place in less than 30% of the publications that used data, suggesting that researchers are still inclined to create and use their own data sets, rather than reusing previously curated data. This paper provides a comprehensive understanding of how data sets are used in science and helps institutions and publishers make useful data policies. Erjia Yan, Kai Li 0010 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Fast botnet detection from streaming logs using online lanczos methodabstractBotnet, a group of coordinated bots, is becoming the main platform of malicious Internet activities like DDOS, click fraud, web scraping, spam/rumor distribution, etc. This paper focuses on design and experiment of a new approach for botnet detection from streaming web server logs, motivated by its wide applicability, real-time protection capability, ease of use and better security of sensitive data. Our algorithm is inspired by a Principal Component Analysis (PCA) to capture correlation in data, and we are first to recognize and adapt Lanczos method to improve the time complexity of PCA-based botnet detection from cubic to sub-cubic, which enables us to more accurately and sensitively detect botnets with sliding time windows rather than fixed time windows. We contribute a generalized online correlation matrix update formula, and a new termination condition for Lanczos iteration for our purpose based on error bound and non-decreasing eigenvalues of symmetric matrices. On our dataset of an ecommerce website logs, experiments show the time cost of Lanczos method with different time windows are consistently only 20% to 25% of PCA. Zheng Chen 0010, Xinli Yu 0002, Cui Lin, Jianliang Gao, Xiaohua Hu 0001, Wei-Shih Yang, Erjia Yan |
IEEE BigData | 10 |
| 2017 | Large-scale joint topic, sentiment & user preference analysis for online reviewsabstractThis paper presents a non-trivial reconstruction of a previous joint topic-sentiment-preference review model TSPRA with stick-breaking representation under the framework of variational inference (VI) and stochastic variational inference (SVI). TSPRA is a Gibbs Sampling based model that solves topics, word sentiments and user preferences altogether and has been shown to achieve good performance, but for large dataset it can only learn from a relatively small sample. We develop the variational models vTSPRA and svTSPRA to improve the time use, and our new approach is capable of processing millions of reviews. We rebuild the generative process, improve the rating regression, solve and present the coordinate-ascent updates of variational parameters, and show the time complexity of each iteration is theoretically linear to the corpus size, and the experiments on Amazon datasets show it converges faster than TSPRA and attains better results given the same amount of time. In addition, we tune svTSPRA into an online algorithm ovTSPRA that can monitor oscillations of sentiment and preference overtime. Some interesting fluctuations are captured and possible explanations are provided. The results give strong visual evidence that user preference is better treated as an independent factor from sentiment. Xinli Yu 0002, Zheng Chen 0010, Wei-Shih Yang, Xiaohua Hu 0001, Erjia Yan, Guangrong Li |
IEEE BigData | 5 |
| 2017 | A natural language interface to a graph-based bibliographic information retrieval system
Yongjun Zhu 0001, Erjia Yan, Il-Yeol Song |
Data Knowl. Eng. | 2 |
| 2017 | Adding the dimension of knowledge trading to source impact assessment: Approaches, indicators, and implicationsabstractThe objective of this paper is to systematically assess sources' (e.g., journals and proceedings) impact in knowledge trading. While there have been efforts at evaluating different aspects of journal impact, the dimension of knowledge trading is largely absent. To fill the gap, this study employed a set of trading‐based indicators, including weighted degree centrality, Shannon entropy, and weighted betweenness centrality, to assess sources' trading impact. These indicators were applied to several time‐sliced source‐to‐source citation networks that comprise 33,634 sources indexed in the Scopus database. The results show that several interdisciplinary sources, such as Nature, PLoS One, Proceedings of the National Academy of Sciences, and Science, and several specialty sources, such as Lancet, Lecture Notes in Computer Science, Journal of the American Chemical Society, Journal of Biological Chemistry, and New England Journal of Medicine, have demonstrated their marked importance in knowledge trading. Furthermore, this study also reveals that, overall, sources have established more trading partners, increased their trading volumes, broadened their trading areas, and diversified their trading contents over the past 15 years from 1997 to 2011. These results inform the understanding of source‐level impact assessment and knowledge diffusion. Erjia Yan, Yongjun Zhu 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | The use of a graph-based system to improve bibliographic information retrieval: System design, implementation, and evaluationabstractIn this article, we propose a graph‐based interactive bibliographic information retrieval system—GIBIR. GIBIR provides an effective way to retrieve bibliographic information. The system represents bibliographic information as networks and provides a form‐based query interface. Users can develop their queries interactively by referencing the system‐generated graph queries. Complex queries such as “papers on information retrieval, which were cited by John's papers that had been presented in SIGIR” can be effectively answered by the system. We evaluate the proposed system by developing another relational database‐based bibliographic information retrieval system with the same interface and functions. Experiment results show that the proposed system executes the same queries much faster than the relational database‐based system, and on average, our system reduced the execution time by 72% (for 3‐node query), 89% (for 4‐node query), and 99% (for 5‐node query). Yongjun Zhu 0001, Erjia Yan, Il-Yeol Song |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Science communication and dissemination in different cultures: An analysis of the audience for TED videos in China and abroadabstractDisseminated across the world in more than 100 languages and viewed over 1 billion times, TED Talks is a successful example of web‐based science communication. This study investigates the impact of TED Talks videos on YouKu, a Chinese video portal, and YouTube using 6 measures of impact: number of views; likes; dislikes; comments; bookmarks; and shares. In particular, we study the relationship between the topicality and impact of these videos. Findings demonstrate that topics vary greatly in terms of their impact: Topics on entertainment and psychology/philosophy receive more views and likes, whereas design/art and astronomy/biology/oceanography attract fewer comments and bookmarks. Moreover, we identify several topical differences between YouKu and YouTube users. Topics on global issues and technology are more popular on YouKu, whereas topics on entertainment and psychology/philosophy are more popular on YouTube. By analyzing the popularity distribution of videos and the audience characteristics of YouKu, we find that women are more interested in topics on education and psychology/philosophy, whereas men favor topics on technology and astronomy/biology/oceanography. Xuelian Pan, Erjia Yan, Weina Hua |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Disciplinary knowledge production and diffusion in scienceabstractThis study examines patterns of dynamic disciplinary knowledge production and diffusion. It uses a citation data set of Scopus‐indexed journals and proceedings. The journal‐level citation data set is aggregated into 27 subject areas and these subjects are selected as the unit of analysis. A 3‐step approach is employed: the first step examines disciplines' citation characteristics through scientific trading dimensions; the second step analyzes citation flows between pairs of disciplines; and the third step uses egocentric citation networks to assess individual disciplines' citation flow diversity through S hannon entropy. The results show that measured by scientific impact, the subjects of C hemical E ngineering, E nergy, and E nvironmental S cience have the fastest growth. Furthermore, most subjects are carrying out more diversified knowledge trading practices by importing higher volumes of knowledge from a greater number of subjects. The study also finds that the growth rates of disciplinary citations align with the growth rates of global research and development ( R & D ) expenditures, thus providing evidence to support the impact of R & D expenditures on knowledge production. Erjia Yan |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Using path-based approaches to examine the dynamic structure of discipline-level citation networks: 1997-2011abstractThe objective of this paper is to identify the dynamic structure of several time‐dependent, discipline‐level citation networks through a path‐based method. A network data set is prepared that comprises 27 subjects and their citations aggregated from more than 27,000 journals and proceedings indexed in the Scopus database. A maximum spanning tree method is employed to extract paths in the weighted, directed, and cyclic networks. This paper finds that subjects such as Medicine, Biochemistry, Chemistry, Materials Science, Physics, and Social Sciences are the ones with multiple branches in the spanning tree. This paper also finds that most paths connect science, technology, engineering, and mathematics (STEM) fields; 2 critical paths connecting STEM and non‐STEM fields are the one from Mathematics to Decision Sciences and the one from Medicine to Social Sciences. Erjia Yan, Qi Yu 0005 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | A lead-lag analysis of the topic evolution patterns for preprints and publicationsabstractThis study applied LDA (latent D irichlet allocation) and regression analysis to conduct a lead‐lag analysis to identify different topic evolution patterns between preprints and papers from arXiv and the W eb of S cience ( WoS ) in astrophysics over the last 20 years (1992–2011). Fifty topics in arXiv and WoS were generated using an LDA algorithm and then regression models were used to explain 4 types of topic growth patterns. Based on the slopes of the fitted equation curves, the paper redefines the topic trends and popularity. Results show that arXiv and WoS share similar topics in a given domain, but differ in evolution trends. Topics in WoS lose their popularity much earlier and their durations of popularity are shorter than those in arXiv . This work demonstrates that open access preprints have stronger growth tendency as compared to traditional printed publications. Beibei Hu, Xianlei Dong, Timothy D. Bowman, Ying Ding 0001, Stasa Milojevic, Chaoqun Ni, Erjia Yan, Vincent Larivière |
J. Assoc. Inf. Sci. Technol. | 8 |
| 2015 | Research dynamics, impact, and dissemination: A topic-level analysisabstractIn informetrics, journals have been used as a standard unit to analyze research impact, productivity, and scholarship. The increasing practice of interdisciplinary research challenges the effectiveness of journal‐based assessments. The aim of this article is to highlight topics as a valuable unit of analysis. A set of topic‐based approaches is applied to a data set on library and information science publications. Results show that topic‐based approaches are capable of revealing the research dynamics, impact, and dissemination of the selected data set. The article also identifies a nonsignificant relationship between topic popularity and impact and argues for the need to use both variables in describing topic characteristics. Additionally, a flow map illustrates critical topic‐level knowledge dissemination channels. Erjia Yan |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | Finding knowledge paths among scientific disciplinesabstractThis paper uncovers patterns of knowledge dissemination among scientific disciplines. Although the transfer of knowledge is largely unobservable, citations from one discipline to another have been proven to be an effective proxy to study disciplinary knowledge flow. This study constructs a knowledge‐flow network in which a node represents a Journal Citation Reports subject category and a link denotes the citations from one subject category to another. Using the concept of shortest path, several quantitative measurements are proposed and applied to a knowledge‐flow network. Based on an examination of subject categories in Journal Citation Reports, this study indicates that social science domains tend to be more self‐contained, so it is more difficult for knowledge from other domains to flow into them; at the same time, knowledge from science domains, such as biomedicine‐, chemistry‐, and physics‐related domains, can access and be accessed by other domains more easily. This study also shows that social science domains are more disunified than science domains, because three fifths of the knowledge paths from one social science domain to another require at least one science domain to serve as an intermediate. This work contributes to discussions on disciplinarity and interdisciplinarity by providing empirical analysis. Erjia Yan |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Scholarly network similarities: How bibliographic coupling networks, citation networks, cocitation networks, topical networks, coauthorship networks, and coword networks relate to each otherabstractThis study explores the similarity among six types of scholarly networks aggregated at the institution level, including bibliographic coupling networks, citation networks, cocitation networks, topical networks, coauthorship networks, and coword networks. Cosine distance is chosen to measure the similarities among the six networks. The authors found that topical networks and coauthorship networks have the lowest similarity; cocitation networks and citation networks have high similarity; bibliographic coupling networks and cocitation networks have high similarity; and coword networks and topical networks have high similarity. In addition, through multidimensional scaling, two dimensions can be identified among the six networks: Dimension 1 can be interpreted as citation‐based versus noncitation‐based, and Dimension 2 can be interpreted as social versus cognitive. The authors recommend the use of hybrid or heterogeneous networks to study research interaction and scholarly communications. Erjia Yan, Ying Ding 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | Discovering author impact: A PageRank perspective
Erjia Yan, Ying Ding 0001 |
Inf. Process. Manag. | 1 |
| 2011 | Modeling topic and community structure in social tagging: The TTR-LDA-Community modelabstractThe presence of social networks in complex systems has made networks and community structure a focal point of study in many domains. Previous studies have focused on the structural emergence and growth of communities and on the topics displayed within the network. However, few scholars have closely examined the relationship between the thematic and structural properties of networks. Therefore, this article proposes the Tagger Tag Resource-Latent Dirichlet Allocation-Community model (TTR-LDA-Community model), which combines the Latent Dirichlet Allocation (LDA) model with the Girvan-Newman community detection algorithm through an inference mechanism. Using social tagging data from Delicious, this article demonstrates the clustering of active taggers into communities, the topic distributions within communities, and the ranking of taggers, tags, and resources within these communities. The data analysis evaluates patterns in community structure and topical affiliations diachronically. The article evaluates the effectiveness of community detection and the inference mechanism embedded in the model and finds that the TTR-LDA-Community model outperforms other traditional models in tag prediction. This has implications for scholars in domains interested in community detection, profiling, and recommender systems. Daifeng Li, Ying Ding 0001, Cassidy R. Sugimoto, Bing He 0003, Jie Tang 0001, Erjia Yan, Tianxi Dong |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2011 | The cognitive structure of Library and Information Science: Analysis of article title wordsabstractAbstract This study comprises a suite of analyses of words in article titles in order to reveal the cognitive structure of Library and Information Science (LIS). The use of title words to elucidate the cognitive structure of LIS has been relatively neglected. The present study addresses this gap by performing (a) co‐word analysis and hierarchical clustering, (b) multidimensional scaling, and (c) determination of trends in usage of terms. The study is based on 10,344 articles published between 1988 and 2007 in 16 LIS journals. Methodologically, novel aspects of this study are: (a) its large scale, (b) removal of non‐specific title words based on the “word concentration” measure (c) identification of the most frequent terms that include both single words and phrases, and (d) presentation of the relative frequencies of terms using “heatmaps”. Conceptually, our analysis reveals that LIS consists of three main branches: the traditionally recognized library‐related and information‐related branches, plus an equally distinct bibliometrics/scientometrics branch. The three branches focus on: libraries, information, and science, respectively. In addition, our study identifies substructures within each branch. We also tentatively identify “information seeking behavior” as a branch that is establishing itself separate from the three main branches. Furthermore, we find that cognitive concepts in LIS evolve continuously, with no stasis since 1992. The most rapid development occurred between 1998 and 2001, influenced by the increased focus on the Internet. The change in the cognitive landscape is found to be driven by the emergence of new information technologies, and the retirement of old ones. Stasa Milojevic, Cassidy R. Sugimoto, Erjia Yan, Ying Ding 0001 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2011 | P-Rank: An indicator measuring prestige in heterogeneous scholarly networksabstractAbstract Ranking scientific productivity and prestige are often limited to homogeneous networks. These networks are unable to account for the multiple factors that constitute the scholarly communication and reward system. This study proposes a new informetric indicator, P‐Rank, for measuring prestige in heterogeneous scholarly networks containing articles, authors, and journals. P‐Rank differentiates the weight of each citation based on its citing papers, citing journals, and citing authors. Articles from 16 representative library and information science journals are selected as the dataset. Principle Component Analysis is conducted to examine the relationship between P‐Rank and other bibliometric indicators. We also compare the correlation and rank variances between citation counts and P‐Rank scores. This work provides a new approach to examining prestige in scholarly communication networks in a more comprehensive and nuanced way. Erjia Yan, Ying Ding 0001, Cassidy R. Sugimoto |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | Institutional interactions: Exploring social, cognitive, and geographic relationships between institutions as demonstrated through citation networksabstractAbstract The objective of this research is to examine the interaction of institutions, based on their citation and collaboration networks. The domain of library and information science is examined, using data from 1965–2010. A linear model is formulated to explore the factors that are associated with institutional citation behaviors, using the number of citations as the dependent variable, and the number of collaborations, physical distance, and topical distance as independent variables. It is found that institutional citation behaviors are associated with social, topical, and geographical factors. Dynamically, the number of citations is becoming more associated with collaboration intensity and less dependent on the country boundary and/or physical distance. This research is informative for scientometricians and policy makers. Erjia Yan, Cassidy R. Sugimoto |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Dynamic Features of Social Tagging Vocabulary: Delicious, Flickr and YouTubeabstractThis article investigates the dynamic features of social tagging vocabularies in Delicious, Flickr and YouTube from 2003 to 2008. It analyzes the evolution of the usage of the most popular tags in each of these three social networks. We find that for different tagging systems, the dynamic features reflect different cognitive processes. At the macro level, the tag growth obeys power-law distribution for all three tagging systems with exponents lower than one. At the micro level, the tag growth of popular resources in all three tagging systems follows a similar power-law distribution. Moreover, we find that the exponents of tag growth varied in different evolving stages of popular individual resources. Daifeng Li, Ying Ding 0001, Stasa Milojevic, Bing He 0003, Erjia Yan, Tianxi Dong |
ASONAM | 6 |
| 2010 | Community-based topic modeling for social taggingabstractExploring community is fundamental for uncovering the connections between structure and function of complex networks and for practical applications in many disciplines such as biology and sociology. In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism. The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008. Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time. We evaluate the effectiveness of our model and show that the TTR-LDA-Community model is meaningful for understanding communities and outperforms TTR-LDA and LDA models in tag prediction. Daifeng Li, Bing He 0003, Ying Ding 0001, Jie Tang 0001, Cassidy R. Sugimoto, Erjia Yan, Juan-Zi Li, Tianxi Dong |
CIKM | 7 |
| 2010 | Upper tag ontology for integrating social tagging dataabstractAbstract Data integration and mediation have become central concerns of information technology over the past few decades. With the advent of the Web and the rapid increases in the amount of data and the number of Web documents and users, researchers have focused on enhancing the interoperability of data through the development of metadata schemes. Other researchers have looked to the wealth of metadata generated by bookmarking sites on the Social Web. While several existing ontologies have capitalized on the semantics of metadata created by tagging activities, the Upper Tag Ontology (UTO) emphasizes the structure of tagging activities to facilitate modeling of tagging data and the integration of data from different bookmarking sites as well as the alignment of tagging ontologies. UTO is described and its utility in modeling, harvesting, integrating, searching, and analyzing data is demonstrated with metadata harvested from three major social tagging systems (Delicious, Flickr, and YouTube). Ying Ding 0001, Elin K. Jacob, Michael A. H. Fried, Ioan Toma, Erjia Yan, Schubert Foo, Stasa Milojevic |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2010 | Weighted citation: An indicator of an article's prestigeabstractAbstract The authors propose using the technique of weighted citation to measure an article's prestige. The technique allocates a different weight to each reference by taking into account the impact of citing journals and citation time intervals. Weightedcitation captures prestige, whereas citation counts capture popularity. They compare the value variances for popularity and prestige for articles published in the Journal of the American Society for Information Science and Technology from 1998 to 2007, and find that the majority have comparable status. Erjia Yan, Ying Ding 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Perspectives on social taggingabstractAbstract Social tagging is one of the major phenomena transforming the World Wide Web from a static platform into an actively shared information space. This paper addresses various aspects of social tagging, including different views on the nature of social tagging, how to make use of social tags, and how to bridge social tagging with other Web functionalities; it discusses the use of facets to facilitate browsing and searching of tagging data; and it presents an analogy between bibliometrics and tagometrics, arguing that established bibliometric methodologies can be applied to analyze tagging behavior on the Web. Based on the Upper Tag Ontology (UTO), a Web crawler was built to harvest tag data from Delicious, Flickr, and YouTube in September 2007. In total, 1.8 million objects, including bookmarks, photos, and videos, 3.1 million taggers, and 12.1 million tags were collected and analyzed. Some tagging patterns and variations are identified and discussed. Ying Ding 0001, Elin K. Jacob, Schubert Foo, Erjia Yan, Nicolas L. George, Lijiang Guo |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2009 | PageRank for ranking authors in co-citation networksabstractAbstract This paper studies how varied damping factors in the PageRank algorithm influence the ranking of authors and proposes weighted PageRank algorithms. We selected the 108 most highly cited authors in the information retrieval (IR) area from the 1970s to 2008 to form the author co‐citation network. We calculated the ranks of these 108 authors based on PageRank with the damping factor ranging from 0.05 to 0.95. In order to test the relationship between different measures, we compared PageRank and weighted PageRank results with the citation ranking, h‐index, and centrality measures. We found that in our author co‐citation network, citation rank is highly correlated with PageRank with different damping factors and also with different weighted PageRank algorithms; citation rank and PageRank are not significantly correlated with centrality measures; and h‐index rank does not significantly correlate with centrality measures but does significantly correlate with other measures. The key factors that have impact on the PageRank of authors in the author co‐citation network are being co‐cited with important authors. Ying Ding 0001, Erjia Yan, Arthur R. Frazho, James Caverlee |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Applying centrality measures to impact analysis: A coauthorship network analysisabstractAbstract Many studies on coauthorship networks focus on network topology and network statistical mechanics. This article takes a different approach by studying micro‐level network properties with the aim of applying centrality measures to impact analysis. Using coauthorship data from 16 journals in the field of library and information science (LIS) with a time span of 20 years (1988–2007), we construct an evolving coauthorship network and calculate four centrality measures (closeness centrality, betweenness centrality, degree centrality, and PageRank) for authors in this network. We find that the four centrality measures are significantly correlated with citation counts. We also discuss the usability of centrality measures in author ranking and suggest that centrality measures can be useful indicators for impact analysis. Erjia Yan, Ying Ding 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |