VLDB 2026 Research / reviewers in the wild / expert
Mike Thelwall
dblp:t/MikeThelwall · also Michael Thelwall
· DBLP profile ↗
111ranked-venue papers in the field
61as first author
12since 2021 · last 2026
0000-0001-6065-205XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 110 (61 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is generative AI reshaping academic practices worldwide? A survey of adoption, benefits, and concernsabstract• Generative AI tools are highly adopted in academia, with significant differences across fields, genders, and countries. • PhD students and early-career academics are the highest adopters for research purposes. • AI applications are widely used for translation, proofreading, and literature review, but less for data analysis in research activities. • Content creation, learning support, and assignment design are key reasons for using AI in teaching, with different patterns based on academic positions and countries. • Inaccurate information, plagiarism, and reduced critical thinking skills are the top concerns of AI use in academia. Although generative AI is transforming academic research and education, little is known about the role, gender, international, and disciplinary variations in uptake and use. This 20-country survey of publishing academics shows the widespread awareness and adoption of generative AI tools in academia, but with substantial international and disciplinary differences, and some role and gender differences. In particular, females were 10 % less likely to use Gen AI frequently (daily or weekly) for research, which may exacerbate gender inequalities. Perhaps surprisingly, the highest adoption rates occurred in some non-Western nations, possibly because of a greater need for translation services. The highest awareness is in the social sciences, perhaps because of the greater need for text analysis. Across all groups, these tools were mainly used for academic writing rather than data analysis and support for critical thinking. Despite this, personalized instruction and problem-solving are among generative AI's most generally claimed benefits. However, participants in all groups were skeptical about the creativity, accuracy, and consistency of AI-generated content in academic contexts. The most significant concerns about using generative AI in academia were inaccuracy, plagiarism, discouraging critical thinking, a lack of transparency and explainability, intellectual property rights violations, and data privacy risks. For policymakers, the findings point to fields and countries that may need action to prevent falling behind, as well as the ongoing need to investigate and monitor the impacts of generative AI on research practices. Ehsan Mohammadi, Mike Thelwall, Yizhou Cai, Taylor Collier, Iman Tahamtan, Azar Eftekhar |
Inf. Process. Manag. | 2 |
| 2025 | Estimating the quality of published medical research with ChatGPT
Mike Thelwall, Xiaorui Jiang, Peter A. Bath |
Inf. Process. Manag. | 1 |
| 2025 | Assessing the societal influence of academic research with ChatGPT: Impact case study evaluationsabstractAbstract Academics and departments are sometimes judged by how their research has benefited society. For example, the UK's Research Excellence Framework (REF) assesses Impact Case Studies (ICSs), which are five‐page evidence‐based claims of societal impacts. This article investigates whether ChatGPT can evaluate societal impact claims and therefore potentially support expert human assessors. For this, various parts of 6220 public ICSs from REF2021 were fed to ChatGPT 4o‐mini along with the REF2021 evaluation guidelines, comparing ChatGPT's predictions with published departmental average ICS scores. The results suggest that the optimal strategy for high correlations with expert scores is to input the title and summary of an ICS but not the remaining text and to modify the original REF guidelines to encourage a stricter evaluation. The scores generated by this approach correlated positively with departmental average scores in all 34 Units of Assessment (UoAs), with values between 0.18 (Economics and Econometrics) and 0.56 (Psychology, Psychiatry and Neuroscience). At the departmental level, the corresponding correlations were higher, reaching 0.71 for Sport and Exercise Sciences, Leisure and Tourism. Thus, ChatGPT‐based ICS evaluations are simple and viable to support or cross‐check expert judgments, although their value varies substantially between fields. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2025 | ChatGPT for complex text evaluation tasksabstractAbstract ChatGPT and other large language models (LLMs) have been successful at natural and computer language processing tasks with varying degrees of complexity. This brief communication summarizes the lessons learned from a series of investigations into its use for the complex text analysis task of research quality evaluation. In summary, ChatGPT is very good at understanding and carrying out complex text processing tasks in the sense of producing plausible responses with minimum input from the researcher. Nevertheless, its outputs require systematic testing to assess their value because they can be misleading. In contrast to simple tasks, the outputs from complex tasks are highly varied and better results can be obtained by repeating the prompts multiple times in different sessions and averaging the ChatGPT outputs. Varying ChatGPT's configuration parameters from their defaults does not seem to be useful, except for the length of the output requested. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2025 | Is OpenAlex suitable for research quality evaluation and which citation indicator is best?abstractAbstract This article compares (1) citation analysis with OpenAlex and Scopus, testing their citation counts, document type/coverage, and subject classifications and (2) three citation‐based indicators: raw counts, (field and year) Normalized Citation Scores (NCS), and Normalized Log‐transformed Citation Scores (NLCS). Methods (1&2): The indicators calculated from 28.6 million articles were compared through 8704 correlations on two gold standards for 97,816 UK Research Excellence Framework (REF) 2021 articles. The primary gold standard is ChatGPT scores, and the secondary is the average REF2021 expert review score for the department submitting the article. Results: (1) OpenAlex provides better citation counts than Scopus, and its inclusive document classification/scope does not seem to cause substantial field normalization problems. The broadest OpenAlex classification scheme provides the best indicators. (2) Counterintuitively, raw citation counts are at least as good as nearly all field normalized indicators and better for single years, and NCS is better than NLCS. (1&2) There are substantial field differences. Thus, (1) OpenAlex is suitable for citation analysis in most fields and (2) the major citation‐based indicators seem to work counterintuitively compared to quality judgments. Field normalization seems ineffective because more cited fields tend to produce higher quality work, affecting interdisciplinary research or within‐field topic differences. Mike Thelwall, Xiaorui Jiang |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2024 | Factors associating with or predicting more cited or higher quality journal articles: An Annual Review of Information Science and Technology (ARIST) paperabstractAbstract Identifying factors that associate with more cited or higher quality research may be useful to improve science or to support research evaluation. This article reviews evidence for the existence of such factors in article text and metadata. It also reviews studies attempting to estimate article quality or predict long‐term citation counts using statistical regression or machine learning for journal articles or conference papers. Although the primary focus is on document‐level evidence, the related task of estimating the average quality scores of entire departments from bibliometric information is also considered. The review lists a huge range of factors that associate with higher quality or more cited research in some contexts (fields, years, journals) but the strength and direction of association often depends on the set of papers examined, with little systematic pattern and rarely any cause‐and‐effect evidence. The strongest patterns found include the near universal usefulness of journal citation rates, author numbers, reference properties, and international collaboration in predicting (or associating with) higher citation counts, and the greater usefulness of citation‐related information for predicting article quality in the medical, health and physical sciences than in engineering, social sciences, arts, and humanities. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2024 | Which international co-authorships produce higher quality journal articles?abstractAbstract International collaboration is sometimes encouraged in the belief that it generates higher quality research or is more capable of addressing societal problems. Nevertheless, while there is evidence that the journal articles of international teams tend to be more cited than average, perhaps from increased international audiences, there is no science‐wide direct academic evidence of a connection between international collaboration and research quality. This article empirically investigates the connection between international collaboration and research quality for the first time, with 148,977 UK‐based journal articles with post publication expert review scores from the 2021 Research Excellence Framework (REF). Using an ordinal regression model controlling for collaboration, international partners increased the odds of higher quality scores in 27 out of 34 Units of Assessment (UoAs) and all Main Panels. The results therefore give the first large scale evidence of the fields in which international co‐authorship for articles is usually apparently beneficial. At the country level, the results suggests that UK collaboration with other high research‐expenditure economies generates higher quality research, even when the countries produce lower citation impact journal articles than the United Kingdom. Worryingly, collaborations with lower research‐expenditure economies tend to be judged lower quality, possibly through misunderstanding Global South research goals. Mike Thelwall, Kayvan Kousha, Mahshid Abdoli, Emma Stuart, Meiko Makita, Paul Wilson 0001, Jonathan M. Levitt |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2023 | Do altmetric scores reflect article quality? Evidence from the UK Research Excellence Framework 2021abstractAbstract Altmetrics are web‐based quantitative impact or attention indicators for academic articles that have been proposed to supplement citation counts. This article reports the first assessment of the extent to which mature altmetrics from Altmetric.com and Mendeley associate with individual article quality scores. It exploits expert norm‐referenced peer review scores from the UK Research Excellence Framework 2021 for 67,030+ journal articles in all fields 2014–2017/2018, split into 34 broadly field‐based Units of Assessment (UoAs). Altmetrics correlated more strongly with research quality than previously found, although less strongly than raw and field normalized Scopus citation counts. Surprisingly, field normalizing citation counts can reduce their strength as a quality indicator for articles in a single field. For most UoAs, Mendeley reader counts are the best altmetric (e.g., three Spearman correlations with quality scores above 0.5), tweet counts are also a moderate strength indicator in eight UoAs (Spearman correlations with quality scores above 0.3), ahead of news (eight correlations above 0.3, but generally weaker), blogs (five correlations above 0.3), and Facebook (three correlations above 0.3) citations, at least in the United Kingdom. In general, altmetrics are the strongest indicators of research quality in the health and physical sciences and weakest in the arts and humanities. Mike Thelwall, Kayvan Kousha, Mahshid Abdoli, Emma Stuart, Meiko Makita, Paul Wilson 0001, Jonathan M. Levitt |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2023 | Why are coauthored academic articles more cited: Higher quality or larger audience?abstractAbstract Collaboration is encouraged because it is believed to improve academic research, supported by indirect evidence in the form of more coauthored articles being more cited. Nevertheless, this might not reflect quality but increased self‐citations or the “audience effect”: citations from increased awareness through multiple author networks. We address this with the first science wide investigation into whether author numbers associate with journal article quality, using expert peer quality judgments for 122,331 articles from the 2014–20 UK national assessment. Spearman correlations between author numbers and quality scores show moderately strong positive associations (0.2–0.4) in the health, life, and physical sciences, but weak or no positive associations in engineering and social sciences, with weak negative/positive or no associations in various arts and humanities, and a possible negative association for decision sciences. This gives the first systematic evidence that greater numbers of authors associates with higher quality journal articles in the majority of academia outside the arts and humanities, at least for the UK. Positive associations between team size and citation counts in areas with little association between team size and quality also show that audience effects or other nonquality factors account for the higher citation rates of coauthored articles in some fields. Mike Thelwall, Kayvan Kousha, Mahshid Abdoli, Emma Stuart, Meiko Makita, Paul Wilson 0001, Jonathan M. Levitt |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2023 | In which fields are citations indicators of research quality?abstractAbstract Citation counts are widely used as indicators of research quality to support or replace human peer review and for lists of top cited papers, researchers, and institutions. Nevertheless, the relationship between citations and research quality is poorly evidenced. We report the first large‐scale science‐wide academic evaluation of the relationship between research quality and citations (field normalized citation counts), correlating them for 87,739 journal articles in 34 field‐based UK Units of Assessment (UoA). The two correlate positively in all academic fields, from very weak (0.1) to strong (0.5), reflecting broadly linear relationships in all fields. We give the first evidence that the correlations are positive even across the arts and humanities. The patterns are similar for the field classification schemes of Scopus and Dimensions.ai, although varying for some individual subjects and therefore more uncertain for these. We also show for the first time that no field has a citation threshold beyond which all articles are excellent quality, so lists of top cited articles are not pure collections of excellence, and neither is any top citation percentile indicator. Thus, while appropriately field normalized citations associate positively with research quality in all fields, they never perfectly reflect it, even at high values. Mike Thelwall, Kayvan Kousha, Emma Stuart, Meiko Makita, Mahshid Abdoli, Paul Wilson 0001, Jonathan M. Levitt |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2021 | Male or female gender-polarized YouTube videos are less viewedabstractAbstract As one of the world's most visited websites, YouTube is potentially influential for learning gendered attitudes. Nevertheless, despite evidence of gender influences within the site for some topics, the extent to which YouTube reflects or promotes male/female or other gender divides is unknown. This article analyses 10,211 YouTube videos published in 12 months from 2014 to 2015 using commenter‐portrayed genders (inferred from usernames) and view counts from the end of 2019. Nonbinary genders are omitted for methodological reasons. Although there were highly male and female topics or themes (e.g., vehicles or beauty) and male or female gendering is the norm, videos with topics attracting both males and females tended to have more viewers (after approximately 5 years) than videos in male or female gendered topics. Similarly, within each topic, videos with gender balanced sets of commenters tend to attract more viewers. Thus, YouTube does not seem to be driving male–female gender differences. Mike Thelwall, David Foster |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2021 | Do new research issues attract more citations? A comparison between 25 Scopus subject categoriesabstractAbstract Finding new ways to help researchers and administrators understand academic fields is an important task for information scientists. Given the importance of interdisciplinary research, it is essential to be aware of disciplinary differences in aspects of scholarship, such as the significance of recent changes in a field. This paper identifies potential changes in 25 subject categories through a term comparison of words in article titles, keywords and abstracts in 1 year compared to the previous 4 years. The scholarly influence of new research issues is indirectly assessed with a citation analysis of articles matching each trending term. While topic‐related words dominate the top terms, style, national focus, and language changes are also evident. Thus, as reflected in Scopus, fields evolve along multiple dimensions. Moreover, while articles exploiting new issues are usually more cited in some fields, such as Organic Chemistry, they are usually less cited in others, including History. The possible causes of new issues being less cited include externally driven temporary factors, such as disease outbreaks, and internally driven temporary decisions, such as a deliberate emphasis on a single topic (e.g., through a journal special issue). Mike Thelwall, Pardeep Sud |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2020 | Which health and biomedical topics generate the most Facebook interest and the strongest citation relationships?
Ehsan Mohammadi, Karl B. Gregory, Mike Thelwall, Nilofar Barahmand |
Inf. Process. Manag. | 3 |
| 2020 | Female citation impact superiority 1996-2018 in six out of seven English-speaking nationsabstractAbstract Efforts to combat continuing gender inequalities in academia need to be informed by evidence about where differences occur. Citations are relevant as potential evidence in appointment and promotion decisions, but it is unclear whether there have been historical gender differences in average citation impact that might explain the current shortfall of senior female academics. This study investigates the evolution of gender differences in citation impact 1996–2018 for six million articles from seven large English‐speaking nations: Australia, Canada, Ireland, Jamaica, New Zealand, UK, and the USA. The results show that a small female citation advantage has been the norm over time for all these countries except the USA, where there has been no practical difference. The female citation advantage is largest, and statistically significant in most years, for Australia and the UK. This suggests that any academic bias against citing female‐authored research cannot explain current employment inequalities. Nevertheless, comparisons using recent citation data, or avoiding it altogether, during appointments or promotion may disadvantage females in some countries by underestimating the likely greater impact of their work, especially in the long term. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2020 | Academic collaboration rates and citation associations vary substantially between countries and fieldsabstractAbstract Research collaboration is promoted by governments and research funders, but if the relative prevalence and merits of collaboration vary internationally then different national and disciplinary strategies may be needed to promote it. This study compares the team size and field normalized citation impact of research across all 27 Scopus broad fields in the 10 countries with the most journal articles indexed in Scopus 2008–2012. The results show that team size varies substantially by discipline and country, with Japan (4.2) having two‐thirds more authors per article than the United Kingdom (2.5). Solo authorship is rare in China (4%) but common in the United Kingdom (27%). While increasing team size associates with higher citation impact in almost all countries and fields, this association is much weaker in China than elsewhere. There are also field differences in the association between citation impact and collaboration. For example, larger team sizes in the Business, Management & Accounting category do not seem to associate with greater research impact, and for China and India, solo authorship associates with higher citation impact in this field. Overall, there are substantial international and field differences in the extent to which researchers collaborate and the extent to which collaboration associates with higher citation impact. Mike Thelwall, Nabeil Maflahi |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2019 | She's Reddit: A source of statistically significant gendered interest information?
Mike Thelwall, Emma Stuart |
Inf. Process. Manag. | 1 |
| 2018 | Co-saved, co-tweeted, and co-cited networksabstractCounts of tweets and Mendeley user libraries have been proposed as altmetric alternatives to citation counts for the impact assessment of articles. Although both have been investigated to discover whether they correlate with article citations, it is not known whether users tend to tweet or save (in Mendeley) the same kinds of articles that they cite. In response, this article compares pairs of articles that are tweeted, saved to a Mendeley library, or cited by the same user, but possibly a different user for each source. The study analyzes 1,131,318 articles published in 2012, with minimum tweeted (10), saved to Mendeley (100), and cited (10) thresholds. The results show surprisingly minor overall overlaps between the three phenomena. The importance of journals for Twitter and the presence of many bots at different levels of activity suggest that this site has little value for impact altmetrics. The moderate differences between patterns of saving and citation suggest that Mendeley can be used for some types of impact assessments, but sensitivity is needed for underlying differences. Fereshteh Didegah, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | How quickly do publications get read? The evolution of mendeley reader counts for new articlesabstractWithin science, citation counts are widely used to estimate research impact but publication delays mean that they are not useful for recent research. This gap can be filled by Mendeley reader counts, which are valuable early impact indicators for academic articles because they appear before citations and correlate strongly with them. Nevertheless, it is not known how Mendeley readership counts accumulate within the year of publication, and so it is unclear how soon they can be used. In response, this paper reports a longitudinal weekly study of the Mendeley readers of articles in 6 library and information science journals from 2016. The results suggest that Mendeley readers accrue from when articles are first available online and continue to steadily build. For journals with large publication delays, articles can already have substantial numbers of readers by their publication date. Thus, Mendeley reader counts may even be useful as early impact indicators for articles before they have been officially published in a journal issue. If field normalized indicators are needed, then these can be generated when journal issues are published using the online first date. Nabeil Maflahi, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | National scientific performance evolution patterns: Retrenchment, successful expansion, or overextensionabstractNational governments would like to preside over an expanding and increasingly high‐impact science system but are these two goals largely independent or closely linked? This article investigates the relationship between changes in the share of the world's scientific output and changes in relative citation impact for 2.6 million articles from 26 fields in the 25 countries with the most Scopus‐indexed journal articles from 1996 to 2015. There is a negative correlation between expansion and relative citation impact, but their relationship varies. China, Spain, Australia, and Poland were successful overall across the 26 fields, expanding both their share of the world's output and its relative citation impact, whereas Japan, France, Sweden, and Israel had decreased shares and relative citation impact. In contrast, the USA, UK, Germany, Italy, Russia, The Netherlands, Switzerland, Finland, and Denmark all enjoyed increased relative citation impact despite a declining share of publications. Finally, India, South Korea, Brazil, Taiwan, and Turkey all experienced sustained expansion but a recent fall in relative citation impact. These results may partly reflect changes in the coverage of Scopus and the selection of fields. Mike Thelwall, Jonathan M. Levitt |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | TensiStrength: Stress and relaxation magnitude detection for social media texts
Mike Thelwall |
Inf. Process. Manag. | 1 |
| 2017 | Patent citation analysis with GoogleabstractCitations from patents to scientific publications provide useful evidence about the commercial impact of academic research, but automatically searchable databases are needed to exploit this connection for large‐scale patent citation evaluations. Google covers multiple different international patent office databases but does not index patent citations or allow automatic searches. In response, this article introduces a semiautomatic indirect method via Bing to extract and filter patent citations from Google to academic papers with an overall precision of 98%. The method was evaluated with 322,192 science and engineering Scopus articles from every second year for the period 1996–2012. Although manual Google Patent searches give more results, especially for articles with many patent citations, the difference is not large enough to be a major problem. Within Biomedical Engineering, Biotechnology, and Pharmacology & Pharmaceutics, 7% to 10% of Scopus articles had at least one patent citation but other fields had far fewer, so patent citation analysis is only relevant for a minority of publications. Low but positive correlations between Google Patent citations and Scopus citations across all fields suggest that traditional citation counts cannot substitute for patent citations when evaluating research. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Are wikipedia citations important evidence of the impact of scholarly articles and books?abstractIndividual academics and research evaluators often need to assess the value of published research. Although citation counts are a recognized indicator of scholarly impact, alternative data is needed to provide evidence of other types of impact, including within education and wider society. Wikipedia is a logical choice for both of these because the role of a general encyclopaedia is to be an understandable repository of facts about a diverse array of topics and hence it may cite research to support its claims. To test whether Wikipedia could provide new evidence about the impact of scholarly research, this article counted citations to 302,328 articles and 18,735 monographs in English indexed by Scopus in the period 2005 to 2012. The results show that citations from Wikipedia to articles are too rare for most research evaluation purposes, with only 5% of articles being cited in all fields. In contrast, a third of monographs have at least one citation from Wikipedia, with the most in the arts and humanities. Hence, Wikipedia citations can provide extra impact evidence for academic monographs. Nevertheless, the results may be relatively easily manipulated and so Wikipedia is not recommended for evaluations affecting stakeholder interests. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | News stories as evidence for research? BBC citations from articles, Books, and WikipediaabstractAlthough news stories target the general public and are sometimes inaccurate, they can serve as sources of real‐world information for researchers. This article investigates the extent to which academics exploit journalism using content and citation analyses of online BBC News stories cited by Scopus articles. A total of 27,234 Scopus‐indexed publications have cited at least one BBC News story, with a steady annual increase. Citations from the arts and humanities (2.8% of publications in 2015) and social sciences (1.5%) were more likely than citations from medicine (0.1%) and science (<0.1%). Surprisingly, half of the sampled Scopus‐cited science and technology (53%) and medicine and health (47%) stories were based on academic research, rather than otherwise unpublished information, suggesting that researchers have chosen a lower‐quality secondary source for their citations. Nevertheless, the BBC News stories that were most frequently cited by Scopus, Google Books, and Wikipedia introduced new information from many different topics, including politics, business, economics, statistics, and reports about events. Thus, news stories are mediating real‐world knowledge into the academic domain, a potential cause for concern. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Goodreads reviews to assess the wider impacts of booksabstractAlthough peer‐review and citation counts are commonly used to help assess the scholarly impact of published research, informal reader feedback might also be exploited to help assess the wider impacts of books, such as their educational or cultural value. The social website Goodreads seems to be a reasonable source for this purpose because it includes a large number of book reviews and ratings by many users inside and outside of academia. To check this, Goodreads book metrics were compared with different book‐based impact indicators for 15,928 academic books across broad fields. Goodreads engagements were numerous enough in the arts (85% of books had at least one), humanities (80%), and social sciences (67%) for use as a source of impact evidence. Low and moderate correlations between Goodreads book metrics and scholarly or non‐scholarly indicators suggest that reader feedback in Goodreads reflects the many purposes of books rather than a single type of impact. Although Goodreads book metrics can be manipulated, they could be used guardedly by academics, authors, and publishers in evaluations. Kayvan Kousha, Mike Thelwall, Mahshid Abdoli |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Web citations in patents: Evidence of technological impact?abstractPatents sometimes cite webpages either as general background to the problem being addressed or to identify prior publications that limit the scope of the patent granted. Counts of the number of patents citing an organization's website may therefore provide an indicator of its technological capacity or relevance. This article introduces methods to extract URL citations from patents and evaluates the usefulness of counts of patent web citations as a technology indicator. An analysis of patents citing 200 US universities or 177 UK universities found computer science and engineering departments to be frequently cited, as well as research‐related webpages, such as Wikipedia, YouTube, or the Internet Archive. Overall, however, patent URL citations seem to be frequent enough to be useful for ranking major US and the top few UK universities if popular hosted subdomains are filtered out, but the hit count estimates on the first search engine results page should not be relied upon for accuracy. Enrique Orduña-Malea, Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Book genre and author gender: Romance>Paranormal-Romance to Autobiography>MemoirabstractAlthough gender differences are known to exist in the publishing industry and in reader preferences, there is little public systematic data about them. This article uses evidence from the book‐based social website Goodreads to provide a large scale analysis of 50 major English book genres based on author genders. The results show gender differences in authorship in almost all categories and gender differences the level of interest in, and ratings of, books in a minority of categories. Perhaps surprisingly in this context, there is not a clear gender‐based relationship between the success of an author and their prevalence within a genre. The unexpected almost universal authorship gender differences should give new impetus to investigations of the importance of gender in fiction and the success of minority genders in some genres should encourage publishers and librarians to take their work seriously, except perhaps for most male‐authored chick‐lit. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | ResearchGate articles: Age, discipline, audience size, and impactabstractThe large multidisciplinary academic social website ResearchGate aims to help academics to connect with each other and to publicize their work. Despite its popularity, little is known about the age and discipline of the articles uploaded and viewed in the site and whether publication statistics from the site could be useful impact indicators. In response, this article assesses samples of ResearchGate articles uploaded at specific dates, comparing their views in the site to their Mendeley readers and Scopus‐indexed citations. This analysis shows that ResearchGate is dominated by recent articles, which attract about three times as many views as older articles. ResearchGate has uneven coverage of scholarship, with the arts and humanities, health professions, and decision sciences poorly represented and some fields receiving twice as many views per article as others. View counts for uploaded articles have low to moderate positive correlations with both Scopus citations and Mendeley readers, which is consistent with them tending to reflect a wider audience than Scopus‐publishing scholars. Hence, for articles uploaded to the site, view counts may give a genuinely new audience indicator. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | Goodreads: A social network site for book readersabstractGoodreads is an Amazon‐owned book‐based social web site for members to share books, read, review books, rate books, and connect with other readers. Goodreads has tens of millions of book reviews, recommendations, and ratings that may help librarians and readers to select relevant books. This article describes a first investigation of the properties of Goodreads users, using a random sample of 50,000 members. The results suggest that about three quarters of members with a public profile are female, and that there is little difference between male and female users in patterns of behavior, except for females registering more books and rating them less positively. Goodreads librarians and super‐users engage extensively with most features of the site. The absence of strong correlations between book‐based and social usage statistics (e.g., numbers of friends, followers, books, reviews, and ratings) suggests that members choose their own individual balance of social and book activities and rarely ignore one at the expense of the other. Goodreads is therefore neither primarily a book‐based website nor primarily a social network site but is a genuine hybrid, social navigation site. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | SlideShare presentations, citations, users, and trends: A professional site with academic and educational usesabstractSlideShare is a free social website that aims to help users distribute and find presentations. Owned by LinkedIn since 2012, it targets a professional audience but may give value to scholarship through creating a long‐term record of the content of talks. This article tests this hypothesis by analyzing sets of general and scholarly related SlideShare documents using content and citation analysis and popularity statistics reported on the site. The results suggest that academics, students, and teachers are a minority of SlideShare uploaders, especially since 2010, with most documents not being directly related to scholarship or teaching. About two thirds of uploaded SlideShare documents are presentation slides, with the remainder often being files associated with presentations or video recordings of talks. SlideShare is therefore a presentation‐centered site with a predominantly professional user base. Although a minority of the uploaded SlideShare documents are cited by, or cite, academic publications, probably too few articles are cited by SlideShare to consider extracting SlideShare citations for research evaluation. Nevertheless, scholars should consider SlideShare to be a potential source of academic and nonacademic information, particularly in library and information science, education, and business. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Can Amazon.com reviews help to assess the wider impacts of books?abstractAlthough citation counts are often used to evaluate the research impact of academic publications, they are problematic for books that aim for educational or cultural impact. To fill this gap, this article assesses whether a number of simple metrics derived from A mazon.com reviews of academic books could provide evidence of their impact. Based on a set of 2,739 academic monographs from 2008 and a set of 1,305 best‐selling books in 15 A mazon.com academic subject categories, the existence of significant but low or moderate correlations between citations and numbers of reviews, combined with other evidence, suggests that online book reviews tend to reflect the wider popularity of a book rather than its academic impact, although there are substantial disciplinary differences. Metrics based on online reviews are therefore recommended for the evaluation of books that aim at a wide audience inside or outside academia when it is important to capture the broader impacts of educational or cultural activities and when they cannot be manipulated in advance of the evaluation. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | An automatic method for assessing the teaching impact of books from online academic syllabiabstractScholars writing books that are widely used to support teaching in higher education may be undervalued because of a lack of evidence of teaching value. Although sales data may give credible evidence for textbooks, these data may poorly reflect educational uses of other types of books. As an alternative, this article proposes a method to search automatically for mentions of books in online academic course syllabi based on Bing searches for syllabi mentioning a given book, filtering out false matches through an extensive set of rules. The method had an accuracy of over 90% based on manual checks of a sample of 2,600 results from the initial Bing searches. Over one third of about 14,000 monographs checked had one or more academic syllabus mention, with more in the arts and humanities (56%) and social sciences (52%). Low but significant correlations between syllabus mentions and citations across most fields, except the social sciences, suggest that books tend to have different levels of impact for teaching and research. In conclusion, the automatic syllabus search method gives a new way to estimate the educational utility of books in a way that sales data and citation counts cannot. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | When are readership counts as useful as citation counts? Scopus versus Mendeley for LIS journalsabstractIn theory, articles can attract readers on the social reference sharing site M endeley before they can attract citations, so M endeley altmetrics could provide early indications of article impact. This article investigates the influence of time on the number of M endeley readers of an article through a theoretical discussion and an investigation into the relationship between counts of readers of, and citations to, 4 general library and information science ( LIS ) journals. For this discipline, it takes about 7 years for articles to attract as many S copus citations as M endeley readers, and after this the S pearman correlation between readers and citers is stable at about 0.6 for all years. This suggests that M endeley readership counts may be useful impact indicators for both newer and older articles. The lack of dates for individual M endeley article readers and an unknown bias toward more recent articles mean that readership data should be normalized individually by year, however, before making any comparisons between articles published in different years. Nabeil Maflahi, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Can Mendeley bookmarks reflect readership? A survey of user motivationsabstractAlthough Mendeley bookmarking counts appear to correlate moderately with conventional citation metrics, it is not known whether academic publications are bookmarked in Mendeley in order to be read or not. Without this information, it is not possible to give a confident interpretation of altmetrics derived from Mendeley. In response, a survey of 860 Mendeley users shows that it is reasonable to use Mendeley bookmarking counts as an indication of readership because most (55%) users with a Mendeley library had read or intended to read at least half of their bookmarked publications. This was true across all broad areas of scholarship except for the arts and humanities (42%). About 85% of the respondents also declared that they bookmarked articles in Mendeley to cite them in their publications, but some also bookmark articles for use in professional (50%), teaching (25%), and educational activities (13%). Of course, it is likely that most readers do not record articles in Mendeley and so these data do not represent all readers. In conclusion, Mendeley bookmark counts seem to be indicators of readership leading to a combination of scholarly impact and wider professional impact. Ehsan Mohammadi, Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Not all international collaboration is beneficial: The Mendeley readership and citation impact of biochemical research collaborationabstractBiochemistry is a highly funded research area that is typified by large research teams and is important for many areas of the life sciences. This article investigates the citation impact and Mendeley readership impact of biochemistry research from 2011 in the Web of Science according to the type of collaboration involved. Negative binomial regression models are used that incorporate, for the first time, the inclusion of specific countries within a team. The results show that, holding other factors constant, larger teams robustly associate with higher impact research, but including additional departments has no effect and adding extra institutions tends to reduce the impact of research. Although international collaboration is apparently not advantageous in general, collaboration with the United States, and perhaps also with some other countries, seems to increase impact. In contrast, collaborations with some other nations seems to decrease impact, although both findings could be due to factors such as differing national proportions of excellent researchers. As a methodological implication, simpler statistical models would find international collaboration to be generally beneficial and so it is important to take into account specific countries when examining collaboration. Pardeep Sud, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Chatting through pictures? A classification of images tweeted in one week in the UK and USAabstractTwitter is used by a substantial minority of the populations of many countries to share short messages, sometimes including images. Nevertheless, despite some research into specific images, such as selfies, and a few news stories about specific tweeted photographs, little is known about the types of images that are routinely shared. In response, this article reports a content analysis of random samples of 800 images tweeted from the UK or USA during a week at the end of 2014. Although most images were photographs, a substantial minority were hybrid or layered image forms: phone screenshots, collages, captioned pictures, and pictures of text messages. About half were primarily of one or more people, including 10% that were selfies, but a wide variety of other things were also pictured. Some of the images were for advertising or to share a joke but in most cases the purpose of the tweet seemed to be to share the minutiae of daily lives, performing the function of chat or gossip, sometimes in innovative ways. Mike Thelwall, Olga Goriunova, Farida Vis, Simon Faulkner, Anne Burns, Jim Aulich, Amalia Más-Bleda, Emma Stuart, Francesco D'Orazio |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Guideline references and academic citations as evidence of the clinical value of health researchabstractThis article introduces a new source of evidence of the value of medical‐related research: citations from clinical guidelines. These give evidence that research findings have been used to inform the day‐to‐day practice of medical staff. To identify whether citations from guidelines can give different information from that of traditional citation counts, this article assesses the extent to which references in clinical guidelines tend to be highly cited in the academic literature and highly read in Mendeley. Using evidence from the United Kingdom, references associated with the UK's National Institute of Health and Clinical Excellence (NICE) guidelines tended to be substantially more cited than comparable articles, unless they had been published in the most recent 3 years. Citation counts also seemed to be stronger indicators than Mendeley readership altmetrics. Hence, although presence in guidelines may be particularly useful to highlight the contributions of recently published articles, for older articles citation counts may already be sufficient to recognize their contributions to health in society. Mike Thelwall, Nabeil Maflahi |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Mendeley readership counts: An investigation of temporal and disciplinary differencesabstractScientists and managers using citation‐based indicators to help evaluate research cannot evaluate recent articles because of the time needed for citations to accrue. Reading occurs before citing, however, and so it makes sense to count readers rather than citations for recent publications. To assess this, Mendeley readers and citations were obtained for articles from 2004 to late 2014 in five broad categories (agriculture, business, decision science, pharmacy, and the social sciences) and 50 subcategories. In these areas, citation counts tended to increase with every extra year since publication, and readership counts tended to increase faster initially but then stabilize after about 5 years. The correlation between citations and readers was also higher for longer time periods, stabilizing after about 5 years. Although there were substantial differences between broad fields and smaller differences between subfields, the results confirm the value of Mendeley reader counts as early scientific impact indicators. Mike Thelwall, Pardeep Sud |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Does research with statistics have more impact? The citation rank advantage of structural equation modelingabstractStatistics are essential to many areas of research and individual statistical techniques may change the ways in which problems are addressed as well as the types of problems that can be tackled. Hence, specific techniques may tend to generate high‐impact findings within science. This article estimates the citation advantage of a technique by calculating the average citation rank of articles using it in the issue of the journal in which they were published. Applied to structural equation modeling (SEM) and four related techniques in 3 broad fields, the results show citation advantages that vary by technique and broad field. For example, SEM seems to be more influential in all broad fields than the 4 simpler methods, with one exception, and hence seems to be particularly worth adding to statistical curricula. In contrast, Pearson correlation apparently has the highest average impact in medicine but the least in psychology. In conclusion, the results suggest that the importance of a statistical technique may vary by discipline and that even simple techniques can help to generate high‐impact research in some contexts. Mike Thelwall, Paul Wilson 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Mendeley readership altmetrics for medical articles: An analysis of 45 fieldsabstractMedical research is highly funded and often expensive and so is particularly important to evaluate effectively. Nevertheless, citation counts may accrue too slowly for use in some formal and informal evaluations. It is therefore important to investigate whether alternative metrics could be used as substitutes. This article assesses whether one such altmetric, Mendeley readership counts, correlates strongly with citation counts across all medical fields, whether the relationship is stronger if student readers are excluded, and whether they are distributed similarly to citation counts. Based on a sample of 332,975 articles from 2009 in 45 medical fields in Scopus, citation counts correlated strongly (about 0.7; 78% of articles had at least one reader) with Mendeley readership counts (from the new version 1 applications programming interface [API]) in almost all fields, with one minor exception, and the correlations tended to decrease slightly when student readers were excluded. Readership followed either a lognormal or a hooked power law distribution, whereas citations always followed a hooked power law, showing that the two may have underlying differences. Mike Thelwall, Paul Wilson 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | An automatic method for extracting citations from Google BooksabstractRecent studies have shown that counting citations from books can help scholarly impact assessment and that Google Books (GB) is a useful source of such citation counts, despite its lack of a public citation index. Searching GB for citations produces approximate matches, however, and so its raw results need time‐consuming human filtering. In response, this article introduces a method to automatically remove false and irrelevant matches from GB citation searches in addition to introducing refinements to a previous GB manual citation extraction method. The method was evaluated by manual checking of sampled GB results and comparing citations to about 14,500 monographs in the Thomson Reuters Book Citation Index (BKCI) against automatically extracted citations from GB across 24 subject areas. GB citations were 103% to 137% as numerous as BKCI citations in the humanities, except for tourism (72%) and linguistics (91%), 46% to 85% in social sciences, but only 8% to 53% in the sciences. In all cases, however, GB had substantially more citing books than did BKCI, with BKCI's results coming predominantly from journal articles. Moderate correlations between the GB and BKCI citation counts in social sciences and humanities, with most BKCI results coming from journal articles rather than books, suggests that they could measure the different aspects of impact, however. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2015 | Who reads research articles? An altmetrics analysis of Mendeley user categoriesabstractLittle detailed information is known about who reads research articles and the contexts in which research articles are read. Using data about people who register in M endeley as readers of articles, this article explores different types of users of C linical M edicine, E ngineering and T echnology, S ocial S cience, P hysics, and C hemistry articles inside and outside academia. The majority of readers for all disciplines were PhD students, postgraduates, and postdocs but other types of academics were also represented. In addition, many C linical M edicine articles were read by medical professionals. The highest correlations between citations and M endeley readership counts were found for types of users who often authored academic articles, except for associate professors in some sub‐disciplines. This suggests that M endeley readership can reflect usage similar to traditional citation impact if the data are restricted to readers who are also authors without the delay of impact measured by citation counts. At the same time, M endeley statistics can also reveal the hidden impact of some research articles, such as educational value for nonauthor users inside academia or the impact of research articles on practice for readers outside academia. Ehsan Mohammadi, Mike Thelwall, Stefanie Haustein, Vincent Larivière |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2015 | How is research blogged? A content analysis approachabstractBlogs that cite academic articles have emerged as a potential source of alternative impact metrics for the visibility of the blogged articles. Nevertheless, to evaluate more fully the value of blog citations, it is necessary to investigate whether research blogs focus on particular types of articles or give new perspectives on scientific discourse. Therefore, we studied the characteristics of peer‐reviewed references in blogs and the typical content of blog posts to gain insight into bloggers' motivations. The sample consisted of 391 blog posts from 2010 to 2012 in Researchblogging.org's health category. The bloggers mostly cited recent research articles or reviews from top multidisciplinary and general medical journals. Using content analysis methods, we created a general classification scheme for blog post content with 10 major topic categories, each with several subcategories. The results suggest that health research bloggers rarely self‐cite and that the vast majority of their blog posts (90%) include a general discussion of the issue covered in the article, with more than one quarter providing health‐related advice based on the article(s) covered. These factors suggest a genuine attempt to engage with a wider, nonacademic audience. Nevertheless, almost 30% of the posts included some criticism of the issues being discussed. Hadas Shema, Judit Bar-Ilan, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | ResearchGate: Disseminating, communicating, and measuring Scholarship?abstractResearchGate is a social network site for academics to create their own profiles, list their publications, and interact with each other. Like Academia.edu, it provides a new way for scholars to disseminate their work and hence potentially changes the dynamics of informal scholarly communication. This article assesses whether ResearchGate usage and publication data broadly reflect existing academic hierarchies and whether individual countries are set to benefit or lose out from the site. The results show that rankings based on ResearchGate statistics correlate moderately well with other rankings of academic institutions, suggesting that ResearchGate use broadly reflects the traditional distribution of academic capital. Moreover, while Brazil, India, and some other countries seem to be disproportionately taking advantage of ResearchGate, academics in China, South Korea, and Russia may be missing opportunities to use ResearchGate to maximize the academic impact of their publications. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | Are scholarly articles disproportionately read in their own country? An analysis of mendeley readersabstractInternational collaboration tends to result in more highly cited research and, partly as a result of this, many research funding schemes are specifically international in scope. Nevertheless, it is not clear whether this citation advantage is the result of higher quality research or due to other factors, such as a larger audience for the publications. To test whether the apparent advantage of internationally collaborative research may be due to additional interest in articles from the countries of the authors, this article assesses the extent to which the national affiliations of the authors of articles affect the national affiliations of their Mendeley readers. Based on E nglish‐language Web of Science articles in 10 fields from science, medicine, social science, and the humanities, the results of statistical models comparing author and reader affiliations suggest that, in most fields, Mendeley users are disproportionately readers of articles authored from within their own country. In addition, there are several cases in which Mendeley users from certain countries tend to ignore articles from specific other countries, although it is not clear whether this reflects national biases or different national specialisms within a field. In conclusion, research funders should not incentivize international collaboration on the basis that it is, in general, higher quality because its higher impact may be primarily due to its larger audience. Moreover, authors should guard against national biases in their reading to select only the best and most relevant publications to inform their research. Mike Thelwall, Nabeil Maflahi |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | Can the impact of non-Western academic books be measured? An investigation of Google Books and Google Scholar for MalaysiaabstractCitation indicators are increasingly used in book‐based disciplines to support peer review in the evaluation of authors and to gauge the prestige of publishers. However, because global citation databases seem to offer weak coverage of books outside the West, it is not clear whether the influence of non‐Western books can be assessed with citations. To investigate this, citations were extracted from Google Books and Google Scholar to 1,357 arts, humanities and social sciences (AHSS) books published by 5 university presses during 1961–2012 in 1 non‐Western nation, Malaysia. A significant minority of the books (23% in Google Books and 37% in Google Scholar, 45% in total) had been cited, with a higher proportion cited if they were older or in English. The combination of Google Books and Google Scholar is therefore recommended, with some provisos, for non‐Western countries seeking to differentiate between books with some impact and books with no impact, to identify the highly‐cited works or to develop an indicator of academic publisher prestige. Abdullah Abrizah, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Tweeting biomedicine: An analysis of tweets and citations in the biomedical literatureabstractData collected by social media platforms have been introduced as new sources for indicators to help measure the impact of scholarly research in ways that are complementary to traditional citation analysis. Data generated from social media activities can be used to reflect broad types of impact. This article aims to provide systematic evidence about how often Twitter is used to disseminate information about journal articles in the biomedical sciences. The analysis is based on 1.4 million documents covered by both PubMed and Web of Science and published between 2010 and 2012. The number of tweets containing links to these documents was analyzed and compared to citations to evaluate the degree to which certain journals, disciplines, and specialties were represented on Twitter and how far tweets correlate with citation impact. With less than 10% of PubMed articles mentioned on Twitter, its uptake is low in general but differs between journals and specialties. Correlations between tweets and citations are low, implying that impact metrics based on tweets are different from those based on citations. A framework using the coverage of articles and the correlation between Twitter mentions and citations is proposed to facilitate the evaluation of novel social‐media‐based metrics. Stefanie Haustein, Isabella Peters, Cassidy R. Sugimoto, Mike Thelwall, Vincent Larivière |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2014 | Disseminating research with web CV hyperlinksabstractSome curricula vitae (web CVs) of academics on the web, including homepages and publication lists, link to open‐access (OA) articles, resources, abstracts in publishers' websites, or academic discussions, helping to disseminate research. To assess how common such practices are and whether they vary by discipline, gender, and country, the authors conducted a large‐scale e‐mail survey of astronomy and astrophysics, public health, environmental engineering, and philosophy across 15 European countries and analyzed hyperlinks from web CVs of academics. About 60% of the 2,154 survey responses reported having a web CV or something similar, and there were differences between disciplines, genders, and countries. A follow‐up outlink analysis of 2,700 web CVs found that a third had at least one outlink to an OA target, typically a public eprint archive or an individual self‐archived file. This proportion was considerably higher in astronomy (48%) and philosophy (37%) than in environmental engineering (29%) and public health (21%). There were also differences in linking to publishers' websites, resources, and discussions. Perhaps most important, however, the amount of linking to OA publications seems to be much lower than allowed by publishers and journals, suggesting that many opportunities for disseminating full‐text research online are being missed, especially in disciplines without established repositories. Moreover, few academics seem to be exploiting their CVs to link to discussions, resources, or article abstracts, which seems to be another missed opportunity for publicizing research. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | arXiv E-prints and the journal of record: An analysis of roles and relationshipsabstractSince its creation in 1991, arXiv has become central to the diffusion of research in a number of fields. Combining data from the entirety of arXiv and the Web of Science ( WoS ), this article investigates (a) the proportion of papers across all disciplines that are on arXiv and the proportion of arXiv papers that are in the WoS , (b) the elapsed time between arXiv submission and journal publication, and (c) the aging characteristics and scientific impact of arXiv e‐prints and their published version. It shows that the proportion of WoS papers found on arXiv varies across the specialties of physics and mathematics, and that only a few specialties make extensive use of the repository. Elapsed time between arXiv submission and journal publication has shortened but remains longer in mathematics than in physics. In physics, mathematics, as well as in astronomy and astrophysics, arXiv versions are cited more promptly and decay faster than WoS papers. The arXiv versions of papers—both published and unpublished—have lower citation rates than published papers, although there is almost no difference in the impact of the arXiv versions of published and unpublished papers. Vincent Larivière, Cassidy R. Sugimoto, Benoit Macaluso, Stasa Milojevic, Blaise Cronin, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2014 | Mendeley readership altmetrics for the social sciences and humanities: Research evaluation and knowledge flowsabstractAlthough there is evidence that counting the readers of an article in the social reference site, Mendeley, may help to capture its research impact, the extent to which this is true for different scientific fields is unknown. In this study, we compare Mendeley readership counts with citations for different social sciences and humanities disciplines. The overall correlation between Mendeley readership counts and citations for the social sciences was higher than for the humanities. Low and medium correlations between Mendeley bookmarks and citation counts in all the investigated disciplines suggest that these measures reflect different aspects of research impact. Mendeley data were also used to discover patterns of information flow between scientific fields. Comparing information flows based on Mendeley bookmarking data and cross‐disciplinary citation analysis for the disciplines revealed substantial similarities and some differences. Thus, the evidence from this study suggests that Mendeley readership data could be used to help capture knowledge transfer across scientific disciplines, especially for people that read but do not author articles, as well as giving impact evidence at an earlier stage than is possible with citation counts. Ehsan Mohammadi, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Do blog citations correlate with a higher number of future citations? Research blogs as a potential source for alternative metricsabstractJournal‐based citations are an important source of data for impact indices. However, the impact of journal articles extends beyond formal scholarly discourse. Measuring online scholarly impact calls for new indices, complementary to the older ones. This article examines a possible alternative metric source, blog posts aggregated at ResearchBlogging.org, which discuss peer‐reviewed articles and provide full bibliographic references. Articles reviewed in these blogs therefore receive “blog citations.” We hypothesized that articles receiving blog citations close to their publication time receive more journal citations later than the articles in the same journal published in the same year that did not receive such blog citations. Statistically significant evidence for articles published in 2009 and 2010 support this hypothesis for seven of 12 journals (58%) in 2009 and 13 of 19 journals (68%) in 2010. We suggest, based on these results, that blog citations can be used as an alternative metric source. Hadas Shema, Judit Bar-Ilan, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2014 | Academia.edu: Social network or Academic Network?abstractAcademic social network sites Academia.edu and ResearchGate, and reference sharing sites Mendeley, Bibsonomy, Zotero, and CiteULike, give scholars the ability to publicize their research outputs and connect with each other. With millions of users, these are a significant addition to the scholarly communication and academic information‐seeking eco‐structure. There is thus a need to understand the role that they play and the changes, if any, that they can make to the dynamics of academic careers. This article investigates attributes of philosophy scholars on Academia.edu, introducing a median‐based, time‐normalizing method to adjust for time delays in joining the site. In comparison to students, faculty tend to attract more profile views but female philosophers did not attract more profile views than did males, suggesting that academic capital drives philosophy uses of the site more than does friendship and networking. Secondary analyses of law, history, and computer science confirmed the faculty advantage (in terms of higher profile views) except for females in law and females in computer science. There was also a female advantage for both faculty and students in law and computer science as well as for history students. Hence, Academia.edu overall seems to reflect a hybrid of scholarly norms (the faculty advantage) and a female advantage that is suggestive of general social networking norms. Finally, traditional bibliometric measures did not correlate with any Academia.edu metrics for philosophers, perhaps because more senior academics use the site less extensively or because of the range informal scholarly activities that cannot be measured by bibliometric methods. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2013 | Determinants of research citation impact in nanoscience and nanotechnologyabstractThis study investigates a range of metrics available when a nanoscience and nanotechnology article is published to see which metrics correlate more with the number of citations to the article. It also introduces the degree of internationality of journals and references as new metrics for this purpose. The journal impact factor; the impact of references; the internationality of authors, journals, and references; and the number of authors, institutions, and references were all calculated for papers published in nanoscience and nanotechnology journals in the Web of Science from 2007 to 2009. Using a zero‐inflated negative binomial regression model on the data set, the impact factor of the publishing journal and the citation impact of the cited references were found to be the most effective determinants of citation counts in all four time periods. In the entire 2007 to 2009 period, apart from journal internationality and author numbers and internationality, all other predictor variables had significant effects on citation counts. Fereshteh Didegah, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2013 | Scholars on soap boxes: Science communication and dissemination in TED videosabstractOnline videos provide a novel, and often interactive, platform for the popularization of science. One successful collection is hosted on theTED(Technology,Entertainment,Design) website. This study uses a range of bibliometric (citation) and webometric (usage and bookmarking) indicators to examineTEDvideos in order to provide insights into the type and scope of their impact. The results suggest thatTEDTalks impact primarily the public sphere, with about three‐quarters of a billion total views, rather than the academic realm. Differences were found among broad disciplinary areas, with art and design videos having generally lower levels of impact but science and technology videos generating otherwise average impact forTED. Many of the metrics were only loosely related, but there was a general consensus about the most popular videos as measured through views or comments onYouTube and theTEDsite. Moreover, most videos were found in at least one online syllabus and videos in online syllabi tended to be more viewed, discussed, and blogged. Less‐liked videos generated more discussion, although this may be because they are more controversial. Science and technology videos presented by academics were more liked than those by nonacademics, showing that academics are not disadvantaged in this new media environment. Cassidy R. Sugimoto, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2013 | Topic-based sentiment analysis for the social web: The role of mood and issue-related wordsabstractGeneral sentiment analysis for the social web has become increasingly useful for shedding light on the role of emotion in online communication and offline events in both academic research and data journalism. Nevertheless, existing general‐purpose social web sentiment analysis algorithms may not be optimal for texts focussed around specific topics. This article introduces 2 new methods, mood setting and lexicon extension, to improve the accuracy of topic‐specific lexical sentiment strength detection for the social web. Mood setting allows the topic mood to determine the default polarity for ostensibly neutral expressive text. Topic‐specific lexicon extension involves adding topic‐specific words to the default general sentiment lexicon. Experiments with 8 data sets show that both methods can improve sentiment analysis performance in corpora and are recommended when the topic focus is tightest. Mike Thelwall, Kevan Buckley |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | The role of online videos in research communication: A content analysis of YouTube videos cited in academic publicationsabstractAlthough there is some evidence that online videos are increasingly used by academics for informal scholarly communication and teaching, the extent to which they are used in published academic research is unknown. This article explores the extent to which YouTube videos are cited in academic publications and whether there are significant broad disciplinary differences in this practice. To investigate, we extracted the URL citations to YouTube videos from academic publications indexed by Scopus. A total of 1,808 Scopus publications cited at least one YouTube video, and there was a steady upward growth in citing online videos within scholarly publications from 2006 to 2011, with YouTube citations being most common within arts and humanities (0.3%) and the social sciences (0.2%). A content analysis of 551 YouTube videos cited by research articles indicated that in science (78%) and in medicine and health sciences (77%), over three fourths of the cited videos had either direct scientific (e.g., laboratory experiments) or scientific‐related contents (e.g., academic lectures or education) whereas in the arts and humanities, about 80% of the YouTube videos had art, culture, or history themes, and in the social sciences, about 63% of the videos were related to news, politics, advertisements, and documentaries. This shows both the disciplinary differences and the wide variety of innovative research communication uses found for videos within the different subject areas. Kayvan Kousha, Mike Thelwall, Mahshid Abdoli |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2012 | Sentiment strength detection for the social webabstractAbstract Sentiment analysis is concerned with the automatic extraction of sentiment‐related information from text. Although most sentiment analysis addresses commercial tasks, such as extracting opinions from product reviews, there is increasing interest in the affective dimension of the social web, and Twitter in particular. Most sentiment analysis algorithms are not ideally suited to this task because they exploit indirect indicators of sentiment that can reflect genre or topic instead. Hence, such algorithms used to process social web texts can identify spurious sentiment patterns caused by topics rather than affective phenomena. This article assesses an improved version of the algorithm SentiStrength for sentiment strength detection across the social web that primarily uses direct indications of sentiment. The results from six diverse social web data sets (MySpace, Twitter, YouTube, Digg, Runners World, BBC Forums) indicate that SentiStrength 2 is successful in the sense of performing better than a baseline approach for all data sets in both supervised and unsupervised cases. SentiStrength is not always better than machine‐learning approaches that exploit indirect indicators of sentiment, however, and is particularly weaker for positive sentiment in news‐related discussions. Overall, the results suggest that, even unsupervised, SentiStrength is robust enough to be applied to a wide variety of different social web contexts. Mike Thelwall, Kevan Buckley, Georgios Paltoglou |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Commenting on YouTube videos: From guatemalan rock to El Big BangabstractAbstract YouTube is one of the world's most popular websites and hosts numerous amateur and professional videos. Comments on these videos might be researched to give insights into audience reactions to important issues or particular videos. Yet, little is known about YouTube discussions in general: how frequent they are, who typically participates, and the role of sentiment. This article fills this gap through an analysis of large samples of text comments on YouTube videos. The results identify patterns and give some benchmarks against which future YouTube research into individual videos can be compared. For instance, the typical YouTube comment was mildly positive, was posted by a 29‐year‐old male, and contained 58 characters. About 23% of comments in the complete comment sets were replies to previous comments. There was no typical density of discussion on YouTube videos in the sense of the proportion of replies to other comments: videos with both few and many replies were common. The YouTube audience engaged with each other disproportionately when making negative comments, however; positive comments elicited few replies. The biggest trigger of discussion seemed to be religion, whereas the videos attracting the least discussion were predominantly from the Music, Comedy, and How to & Style categories. This suggests different audience uses for YouTube, from passive entertainment to active debating. Mike Thelwall, Pardeep Sud, Farida Vis |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Link and co-inlink network diagrams with URL citations or title mentionsabstractWebometric network analyses have been used to map the connectivity of groups of websites to identify clusters, important sites or overall structure. Such analyses have mainly been based upon hyperlink counts, the number of hyperlinks between a pair of websites, although some have used title mentions or URL citations instead. The ability to automatically gather hyperlink counts from Yahoo! ceased in April 2011 and the ability to manually gather such counts was due to cease by early 2012, creating a need for alternatives. This article assesses URL citations and title mentions as possible replacements for hyperlinks in both binary and weighted direct link and co‐inlink network diagrams. It also assesses three different types of data for the network connections: hit count estimates, counts of matching URLs, and filtered counts of matching URLs. Results from analyses of U.S. library and information science departments and U.K. universities give evidence that metrics based upon URLs or titles can be appropriate replacements for metrics based upon hyperlinks for both binary and weighted networks, although filtered counts of matching URLs are necessary to give the best results for co‐title mention and co‐URL citation network diagrams. Mike Thelwall, Pardeep Sud, David Wilkinson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Trending Twitter topics in English: An international comparisonabstractThe worldwide span of the microblogging service Twitter provides an opportunity to make international comparisons of trending topics of interest, such as news stories. Previous international comparisons of news interests have tended to use surveys and may bypass topics not well covered in the mainstream media. This study uses 9 months of English‐language Tweets from the United Kingdom, United States, India, South Africa, New Zealand, and Australia. Based upon the top 50 trending keywords in each country from the 0.5 billion Tweets collected, festivals or religious events are the most common, followed by media events, politics, human interest, and sports. U.S. trending topics have the most interest in the other countries and Indian trending topics the least. Conversely, India is the most interested in other countries’ trending topics and the United States the least. This gives evidence of an international hierarchy of perceived importance or relevance with some issues, such as the international interest in U.S. Thanksgiving celebrations, apparently not being directly driven by the media. This hierarchy echoes, and may be caused by, similar news coverage trends. Although the current imbalanced international news coverage does not seem to be out of step with public news interests, the political implication is that the Twitter‐using public reflects, and hence seems to implicitly accept, international imbalances in news media agenda setting rather than combating them. This is an issue for those believing that these imbalances make the media too powerful. David Wilkinson, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2012 | Twitter, MySpace, Digg: Unsupervised Sentiment Analysis in Social MediaabstractSentiment analysis is a growing area of research with significant applications in both industry and academia. Most of the proposed solutions are centered around supervised, machine learning approaches and review-oriented datasets. In this article, we focus on the more common informal textual communication on the Web, such as online discussions, tweets and social network comments and propose an intuitive, less domain-specific, unsupervised, lexicon-based approach that estimates the level of emotional intensity contained in text in order to make a prediction. Our approach can be applied to, and is tested in, two different but complementary contexts: subjectivity detection and polarity classification. Extensive experiments were carried on three real-world datasets, extracted from online social Web sites and annotated by human evaluators, against state-of-the-art supervised approaches. The results demonstrate that the proposed algorithm, even though unsupervised, outperforms machine learning solutions in the majority of cases, overall presenting a very robust and reliable solution for sentiment analysis of informal communication on the Web. Georgios Paltoglou, Mike Thelwall |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | A combined bibliometric indicator to predict article impact
Jonathan M. Levitt, Mike Thelwall |
Inf. Process. Manag. | 2 |
| 2011 | Assessing the citation impact of books: The role of Google Books, Google Scholar, and ScopusabstractAbstract Citation indictors are increasingly used in some subject areas to support peer review in the evaluation of researchers and departments. Nevertheless, traditional journal‐based citation indexes may be inadequate for the citation impact assessment of book‐based disciplines. This article examines whether online citations from Google Books and Google Scholar can provide alternative sources of citation evidence. To investigate this, we compared the citation counts to 1,000 books submitted to the 2008 U.K. Research Assessment Exercise (RAE) from Google Books and Google Scholar with Scopus citations across seven book‐based disciplines (archaeology; law; politics and international studies; philosophy; sociology; history; and communication, cultural, and media studies). Google Books and Google Scholar citations to books were 1.4 and 3.2 times more common than were Scopus citations, and their medians were more than twice and three times as high as were Scopus median citations, respectively. This large number of citations is evidence that in book‐oriented disciplines in the social sciences, arts, and humanities, online book citations may be sufficiently numerous to support peer review for research evaluation, at least in the United Kingdom. Kayvan Kousha, Mike Thelwall, Somayeh Rezaie |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2011 | Variations between subjects in the extent to which the social sciences have become more interdisciplinaryabstractIncreasing interdisciplinarity has been a policy objective since the 1990s, promoted by many governments and funding agencies, but the question is: How deeply has this affected the social sciences? Although numerous articles have suggested that research has become more interdisciplinary, yet no study has compared the extent to which the interdisciplinarity of different social science subjects has changed. To address this gap, changes in the level of interdisciplinarity since 1980 are investigated for subjects with many articles in the Social Sciences Citation Index (SSCI), using the percentage of cross-disciplinary citing documents (PCDCD) to evaluate interdisciplinarity. For the 14 SSCI subjects investigated, the median level of interdisciplinarity, as measured using cross-disciplinary citations, declined from 1980 to 1990, but rose sharply between 1990 and 2000, confirming previous research. This increase was not fully matched by an increase in the percentage of articles that were assigned to more than one subject category. Nevertheless, although on average the social sciences have recently become more interdisciplinary, the extent of this change varies substantially from subject to subject. The SSCI subject with the largest increase in interdisciplinarity between 1990 and 2000 was Information Science & Library Science (IS&LS) but there is evidence that the level of interdisciplinarity of IS&LS increased less quickly during the first decade of this century. Jonathan M. Levitt, Mike Thelwall, Charles Oppenheim |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2011 | Sentiment in Twitter eventsabstractThe microblogging site Twitter generates a constant stream of communication, some of which concerns events of general interest. An analysis of Twitter may, therefore, give insights into why particular events resonate with the population. This article reports a study of a month of English Twitter posts, assessing whether popular events are typically associated with increases in sentiment strength, as seems intuitively likely. Using the top 30 events, determined by a measure of relative increase in (general) term usage, the results give strong evidence that popular events are normally associated with increases in negative sentiment strength and some evidence that peaks of interest in events have stronger positive sentiment than the time before the peak. It seems that many positive events, such as the Oscars, are capable of generating increased negative sentiment in reaction to them. Nevertheless, the surprisingly small average change in sentiment associated with popular events (typically 1% and only 6% for Tiger Woods' confessions) is consistent with events affording posters opportunities to satisfy pre-existing personal goals more often than eliciting instinctive reactions. Mike Thelwall, Kevan Buckley, Georgios Paltoglou |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | A comparison of methods for collecting web citation data for academic organizationsabstractAbstract The primary webometric method for estimating the online impact of an organization is to count links to its website. Link counts have been available from commercial search engines for over a decade but this was set to end by early 2012 and so a replacement is needed. This article compares link counts to two alternative methods: URL citations and organization title mentions. New variations of these methods are also introduced. The three methods are compared against each other using Yahoo!. Two of the three methods (URL citations and organization title mentions) are also compared against each other using Bing. Evidence from a case study of 131 UK universities and 49 US Library and Information Science (LIS) departments suggests that Bing's Hit Count Estimates (HCEs) for popular title searches are not useful for webometric research but that Yahoo!'s HCEs for all three types of search and Bing's URL citation HCEs seem to be consistent. For exact URL counts the results of all three methods in Yahoo! and both methods in Bing are also consistent. Four types of accuracy factors are also introduced and defined: search engine coverage, search engine retrieval variation, search engine retrieval anomalies, and query polysemy. Mike Thelwall, Pardeep Sud |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Can the impact of scholarly images be assessed online? An exploratory study using image identification technologyabstractAbstract The web contains a huge number of digital pictures. For scholars publishing such images it is important to know how well used their images are, but no method seems to have been developed for monitoring the value of academic images. In particular, can the impact of scientific or artistic images be assessed through identifying images copied or reused on the Internet? This article explores a case study of 260 NASA images to investigate whether the TinEye search engine could theoretically help to provide this information. The results show that the selected pictures had a median of 11 online copies each. However, a classification of 210 of these copies reveals that only 1.4% were explicitly used in academic publications, reflecting research impact, and the majority of the NASA pictures were used for informal scholarly (or educational) communication (37%). Additional analyses of world famous paintings and scientific images about pathology and molecular structures suggest that image contents are important for the type and extent of image use. Although it is reasonable to use statistics derived from TinEye for assessing image reuse value, the extent of its image indexing is not known. Kayvan Kousha, Mike Thelwall, Somayeh Rezaie |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2010 | Sentiment in short strength detection informal textabstractAbstract A huge number of informal messages are posted every day in social network sites, blogs, and discussion forums. Emotions seem to be frequently important in these texts for expressing friendship, showing social support or as part of online arguments. Algorithms to identify sentiment and sentiment strength are needed to help understand the role of emotion in this informal communication and also to identify inappropriate or anomalous affective utterances, potentially associated with threatening behavior to the self or others. Nevertheless, existing sentiment detection algorithms tend to be commercially oriented, designed to identify opinions about products rather than user behaviors. This article partly fills this gap with a new algorithm, SentiStrength, to extract sentiment strength from informal English text, using new methods to exploit the de facto grammars and spelling styles of cyberspace. Applied to MySpace comments and with a lookup table of term sentiment strengths optimized by machine learning, SentiStrength is able to predict positive emotion with 60.6% accuracy and negative emotion with 72.8% accuracy, both based upon strength scales of 1–5. The former, but not the latter, is better than baseline and a wide range of general machine learning approaches. Mike Thelwall, Kevan Buckley, Georgios Paltoglou, Di Cai, Arvid Kappas |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Policy-relevant Webometrics for individual scientific fieldsabstractAbstract Despite over 10 years of research there is no agreement on the most suitable roles for Webometric indicators in support of research policy and almost no field‐based Webometrics. This article partly fills these gaps by analyzing the potential of policy‐relevant Webometrics for individual scientific fields with the help of 4 case studies. Although Webometrics cannot provide robust indicators of knowledge flows or research impact, it can provide some evidence of networking and mutual awareness. The scope of Webometrics is also relatively wide, including not only research organizations and firms but also intermediary groups like professional associations, Web portals, and government agencies. Webometrics can, therefore, provide evidence about the research process to compliment peer review, bibliometric, and patent indicators: tracking the early, mainly prepublication development of new fields and research funding initiatives, assessing the role and impact of intermediary organizations and the need for new ones, and monitoring the extent of mutual awareness in particular research areas. Mike Thelwall, Antje Klitkou, Arnold Verbeek, David Stuart, Celine Vincent |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Public dialogs in social network sites: What is their purpose?abstractAbstract Social network sites (SNSs) such as MySpace and Facebook are important venues for interpersonal communication, especially among youth. One way in which members can communicate is to write public messages on each other's profile, but how is this unusual means of communication used in practice? An analysis of 2,293 public comment exchanges extracted from large samples of U.S. and U.K. MySpace members found them to be relatively rapid, but rarely used for prolonged exchanges. They seem to fulfill two purposes: making initial contact and keeping in touch occasionally such as at birthdays and other important dates. Although about half of the dialogs seem to exchange some gossip, the dialogs seem typically too short to play the role of gossip‐based “social grooming” for typical pairs of Friends, but close Friends may still communicate extensively in SNSs with other methods. Mike Thelwall, David Wilkinson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Data mining emotion in social network communication: Gender differences in MySpaceabstractAbstract Despite the rapid growth in social network sites and in data mining for emotion (sentiment analysis), little research has tied the two together, and none has had social science goals. This article examines the extent to which emotion is present in MySpace comments, using a combination of data mining and content analysis, and exploring age and gender. A random sample of 819 public comments to or from U.S. users was manually classified for strength of positive and negative emotion. Two thirds of the comments expressed positive emotion, but a minority (20%) contained negative emotion, confirming that MySpace is an extraordinarily emotion‐rich environment. Females are likely to give and receive more positive comments than are males, but there is no difference for negative comments. It is thus possible that females are more successful social network site users partly because of their greater ability to textually harness positive affect. Mike Thelwall, David Wilkinson, Sukhvinder Uppal |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Social network site changes over time: The case of MySpaceabstractAbstract The uptake of social network sites (SNSs) has been highly trend‐driven, with Friendster, MySpace, and Facebook being successively the most popular. Given that teens are often early adopters of communication technologies, it seems reasonable to assume that the typical user of any particular SNS would change over time, probably becoming older and covering different segments of the population. This article analyzes changes in MySpace self‐reported member demographics and behavior from 2007 to 2010 using four large samples of members and focusing on the United States. The results indicate that despite its take‐up rate declining, with only about 1 in 10 members being active a year after joining, the dominant (modal) age for active U.S. members remains midadolescence, but has shifted by about 2 years from 15 to 17, and the U.S. dominance of MySpace is shrinking. There also has been a dramatic increase in the median number of Friends for new U.S. members, from 12 to 96—probably due to MySpace's automated Friend Finder. Some factors show little change, however, including the female majority, the 5% minority gay membership, and the approximately 50% private profiles. In addition, there has been an increase in the proportion of Latino/Hispanic U.S. members, suggesting a shifting ethnic profile. Overall, MySpace has surprisingly stable membership demographics and is apparently maintaining its primary youth appeal, perhaps because of its music orientation. David Wilkinson, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Google book search: Citation analysis for social science and the humanitiesabstractAbstract In both the social sciences and the humanities, books and monographs play significant roles in research communication. The absence of citations from most books and monographs from the Thomson Reuters/Institute for Scientific Information databases (ISI) has been criticized, but attempts to include citations from or to books in the research evaluation of the social sciences and humanities have not led to widespread adoption. This article assesses whether Google Book Search (GBS) can partially fill this gap by comparing citations from books with citations from journal articles to journal articles in 10 science, social science, and humanities disciplines. Book citations were 31% to 212% of ISI citations and, hence, numerous enough to supplement ISI citations in the social sciences and humanities covered, but not in the sciences (3%–5%), except for computing (46%), due to numerous published conference proceedings. A case study was also made of all 1,923 articles in the 51 information science and library science ISI‐indexed journals published in 2003. Within this set, highly book‐cited articles tended to receive many ISI citations, indicating a significant relationship between the two types of citation data, but with important exceptions that point to the additional information provided by book citations. In summary, GBS is clearly a valuable new source of citation data for the social sciences and humanities. One practical implication is that book‐oriented scholars should consult it for additional citations to their work when applying for promotion and tenure. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Citation levels and collaboration within library and information scienceabstractAbstract Collaboration is a major research policy objective, but does it deliver higher quality research? This study uses citation analysis to examine the Web of Science (WoS) Information Science & Library Science subject category (IS&LS) to ascertain whether, in general, more highly cited articles are more highly collaborative than other articles. It consists of two investigations. The first investigation is a longitudinal comparison of the degree and proportion of collaboration in five strata of citation; it found that collaboration in the highest four citation strata (all in the most highly cited 22%) increased in unison over time, whereas collaboration in the lowest citation strata (un‐cited articles) remained low and stable. Given that over 40% of the articles were un‐cited, it seems important to take into account the differences found between un‐cited articles and relatively highly cited articles when investigating collaboration in IS&LS. The second investigation compares collaboration for 35 influential information scientists; it found that their more highly cited articles on average were not more highly collaborative than their less highly cited articles. In summary, although collaborative research is conducive to high citation in general, collaboration has apparently not tended to be essential to the success of current and former elite information scientists. Jonathan M. Levitt, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Assessing global diffusion with Web memetics: The spread and evolution of a popular jokeabstractAbstract Memes are small units of culture, analogous to genes, which flow from person to person by copying or imitation. More than any previous medium, the Internet has the technical capabilities for global meme diffusion. Yet, to spread globally, memes need to negotiate their way through cultural and linguistic borders. This article introduces a new broad method,Web memetics, comprising extensive Web searches and combined quantitative and qualitative analyses, to identify and assess: (a) the different versions of a meme, (b) its evolution online, and (c) its Web presence and translation into common Internet languages. This method is demonstrated through one extensively circulated joke about men, women, and computers. The results show that the joke has mutated into several different versions and is widely translated, and that translations incorporate small, local adaptations while retaining the English versions' fundamental components. In conclusion, Web memetics has demonstrated its ability to identify and track the evolution and spread of memes online, with interesting results, albeit for only one case study. Limor Shifman, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Homophily in MySpaceabstractAbstract Social network sites like MySpace are increasingly important environments for expressing and maintaining interpersonal connections, but does online communication exacerbate or ameliorate the known tendency for offline friendships to form between similar people (homophily)? This article reports an exploratory study of the similarity between the reported attributes of pairs of active MySpace Friends based upon a systematic sample of 2,567 members joining on June 18, 2007 and Friends who commented on their profile. The results showed no evidence of gender homophily but significant evidence of homophily for ethnicity, religion, age, country, marital status, attitude towards children, sexual orientation, and reason for joining MySpace. There were also some imbalances: women and the young were disproportionately commenters, and commenters tended to have more Friends than commentees. Overall, it seems that although traditional sources of homophily are thriving in MySpace networks of active public connections, gender homophily has completely disappeared. Finally, the method used has wide potential for investigating and partially tracking homophily in society, providing early warning of socially divisive trends. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Handbook of Research on Web Log Analysis
Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | A statistical analysis of the web presences of European life sciences research teamsabstractAbstract Web links have been used for around ten years to explore the online impact of academic information and information producers. Nevertheless, few studies have attempted to relate link counts to relevant offline attributes of the owners of the targeted Web sites, with the exception of research productivity. This article reports the results of a study to relate site inlink counts to relevant owner characteristics for over 400 European life‐science research group Web sites. The analysis confirmed that research‐group size and Web‐presence size were important for attracting Web links, although research productivity was not. Little evidence was found for significant influence of any of an array of factors, including research‐group leader gender and industry connections. In addition, the choice of search engine for link data created a surprising international difference in the results, with Google perhaps giving unreliable results. Overall, the data collection, statistical analysis and results interpretation were all complex and it seems that we still need to know more about search engines, hyperlinks, and their function in science before we can draw conclusions on their usefulness and role in the canon of science and technology indicators. Franz Barjak, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Assessing the impact of disciplinary research on teaching: An automatic analysis of online syllabusesabstractAbstract The impact of published academic research in the sciences and social sciences, when measured, is commonly estimated by counting citations from journal articles. The Web has now introduced new potential sources of quantitative data online that could be used to measure aspects of research impact. In this article we assess the extent to which citations from online syllabuses could be a valuable source of evidence about the educational utility of research. An analysis of online syllabus citations to 70,700 articles published in 2003 in the journals of 12 subjects indicates that online syllabus citations were sufficiently numerous to be a useful impact indictor in some social sciences, including political science and information and library science, but not in others, nor in any sciences. This result was consistent with current social science research having, in general, more educational value than current science research. Moreover, articles frequently cited in online syllabuses were not necessarily highly cited by other articles. Hence it seems that online syllabus citations provide a valuable additional source of evidence about the impact of journals, scholars, and research articles in some social sciences. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Is multidisciplinary research more highly cited? A macrolevel studyabstractAbstract Interdisciplinary collaboration is a major goal in research policy. This study uses citation analysis to examine diverse subjects in the Web of Science and Scopus to ascertain whether, in general, research published in journals classified in more than one subject is more highly cited than research published in journals classified in a single subject. For each subject, the study divides the journals into two disjoint sets called Multi and Mono . Multi consists of all journals in the subject and at least one other subject whereas Mono consists of all journals in the subject and in no other subject. The main findings are: (a) For social science subject categories in both the Web of Science and Scopus, the average citation levels of articles in Mono and Multi are very similar; and (b) for Scopus subject categories within life sciences, health sciences, and physical sciences, the average citation level of Mono articles is roughly twice that of Multi articles. Hence, one cannot assume that in general, multidisciplinary research will be more highly cited, and the converse is probably true for many areas of science. A policy implication is that, at least in the sciences, multidisciplinary researchers should not be evaluated by citations on the same basis as monodisciplinary researchers. Jonathan M. Levitt, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Extracting accurate and complete results from search engines: Case study windows liveabstractAbstract Although designed for general Web searching, Webometrics and related research commercial search engines are also used to produce estimated hit counts or lists of URLs matching a query. Unfortunately, however, they do not return all matching URLs for a search and their hit count estimates are unreliable. In this article, we assess whether it is possible to obtain complete lists of matching URLs from Windows Live, and whether any of its hit count estimates are robust. As part of this, we introduce two new methods to extract extra URLs from search engines: automated query splitting and automated domain and TLD searching. Both methods successfully identify additional matching URLs but the findings suggest that there is no way to get complete lists of matching URLs or accurate hit counts from Windows Live, although some estimating suggestions are provided. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Social networks, gender, and friending: An analysis of MySpace member profilesabstractAbstract In 2007, the social networking Web site MySpace apparently overthrew Google as the most visited Web site for U.S. Web users. If this heralds a new era of widespread online social networking, then it is important to investigate user behaviour and attributes. Although there has been some research into social networking already, basic demographic data is essential to set previous results in a wider context and to give insights to researchers, marketers and developers. In this article, the demographics of MySpace members are explored through data extracted from two samples of 15,043 and 7,627 member profiles. The median declared age of users was surprisingly high at 21, with a small majority of females. The analysis confirmed some previously reported findings and conjectures about social networking, for example, that female members tend to be more interested in friendship and males more interested in dating. In addition, there was some evidence of three different friending dynamics, oriented towards close friends, acquaintances, or strangers. Perhaps unsurprisingly, female and younger members had more friends than others, and females were more likely to maintain private profiles, but both males and females seemed to prefer female friends, with this tendency more marked in females for their closest friend. The typical MySpace user is apparently female, 21, single, with a public profile, interested in online friendship and logging on weekly to engage with a mixed list of mainly female “friends” who are predominantly acquaintances. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Quantitative comparisons of search engine resultsabstractAbstract Search engines are normally used to find information or Web sites, but Webometric investigations use them for quantitative data such as the number of pages matching a query and the international spread of those pages. For this type of application, the accuracy of the hit count estimates and range of URLs in the full results are important. Here, we compare the applications programming interfaces of Google, Yahoo!, and Live Search for 1,587 single word searches. The hit count estimates were broadly consistent but with Yahoo! and Google, reporting 5–6 times more hits than Live Search. Yahoo! tended to return slightly more matching URLs than Google, with Live Search returning significantly fewer. Yahoo!'s result URLs included a significantly wider range of domains and sites than the other two, and there was little consistency between the three engines in the number of different domains. In contrast, the three engines were reasonably consistent in the number of different top‐level domains represented in the result URLs, although Yahoo! tended to return the most. In conclusion, quantitative results from the three search engines are mostly consistent but with unexpected types of inconsistency that users should be aware of. Google is recommended for hit count estimates but Yahoo! is recommended for all other Webometric purposes. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Online presentations as a source of scientific impact? An analysis of PowerPoint files citing academic journalsabstractAbstract Open‐access online publication has made available an increasingly wide range of document types for scientometric analysis. In this article, we focus on citations in online presentations, seeking evidence of their value as nontraditional indicators of research impact. For this purpose, we searched for online PowerPoint files mentioning any one of 1,807 ISI‐indexed journals in ten science and ten social science disciplines. We also manually classified 1,378 online PowerPoint citations to journals in eight additional science and social science disciplines. The results showed that very few journals were cited frequently enough in online PowerPoint files to make impact assessment worthwhile, with the main exceptions being popular magazines likeScientific AmericanandHarvard Business Review. Surprisingly, however, there was little difference overall in the number of PowerPoint citations to science and to the social sciences, and also in the proportion representing traditional impact (about 60%) and wider impact (about 15%). It seems that the main scientometric value for online presentations may be in tracking the popularization of research, or for comparing the impact of whole journals rather than individual articles. Mike Thelwall, Kayvan Kousha |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Information-centered research for large-scale analyses of new information sourcesabstractAbstract New mass publishing genres, such as blogs and personal home pages provide a rich source of social data that is yet to be fully exploited by the social sciences and humanities. Information‐centered research (ICR) not only provides a genuinely new and useful information science research model for this type of data, but can also contribute to the emerging e‐research infrastructure. Nevertheless, ICR should not be conducted on a purely abstract level, but should relate to potentially relevant problems. Mike Thelwall, Paul Wouters, Jenny Fry |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | Which factors explain the Web impact of scientists' personal homepages?abstractAbstract In recent years, a considerable body of Webometric research has used hyperlinks to generate indicators for the impact of Web documents and the organizations that created them. The relationship between this Web impact and other, offline impact indicators has been explored for entire universities, departments, countries, and scientific journals, but not yet for individual scientists—an important omission. The present research closes this gap by investigating factors that may influence the Web impact (i.e., inlink counts) of scientists' personal homepages. Data concerning 456 scientists from five scientific disciplines in six European countries were analyzed, showing that both homepage content and personal and institutional characteristics of the homepage owners had significant relationships with inlink counts. A multivariate statistical analysis confirmed that full‐text articles are the most linked‐to content in homepages. At the individual homepage level, hyperlinks are related to several offline characteristics. Notable differences regarding total inlinks to scientists' homepages exist between the scientific disciplines and the countries in the sample. There also are both gender and age effects: fewer external inlinks (i.e., links from other Web domains) to the homepages of female and of older scientists. There is only a weak relationship between a scientist's recognition and homepage inlinks and, surprisingly, no relationship between research productivity and inlink counts. Contrary to expectations, the size of collaboration networks is negatively related to hyperlink counts. Some of the relationships between hyperlinks to homepages and the properties of their owners can be explained by the content that the homepage owners put on their homepage and their level of Internet use; however, the findings about productivity and collaborations do not seem to have a simple, intuitive explanation. Overall, the results emphasize the complexity of the phenomenon of Web linking, when analyzed at the level of individual pages. Franz Barjak, Xuemei Li 0002, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2007 | Google Scholar citations and Google Web/URL citations: A multi-discipline exploratory analysisabstractAbstract We use a new data gathering method, “Web/URL citation,” Web/URL and Google Scholar to compare traditional and Web‐based citation patterns across multiple disciplines (biology, chemistry, physics, computing, sociology, economics, psychology, and education) based upon a sample of 1,650 articles from 108 open access (OA) journals published in 2001. A Web/URL citation of an online journal article is a Web mention of its title, URL, or both. For each discipline, except psychology, we found significant correlations between Thomson Scientific (formerly Thomson ISI, here: ISI) citations and both Google Scholar and Google Web/URL citations. Google Scholar citations correlated more highly with ISI citations than did Google Web/URL citations, indicating that the Web/URL method measures a broader type of citation phenomenon. Google Scholar citations were more numerous than ISI citations in computer science and the four social science disciplines, suggesting that Google Scholar is more comprehensive for social sciences and perhaps also when conference articles are valued and published online. We also found large disciplinary differences in the percentage overlap between ISI and Google Scholar citation sources. Finally, although we found many significant trends, there were also numerous exceptions, suggesting that replacing traditional citation sources with the Web or Google Scholar for research impact calculations would be problematic. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | How is science cited on the Web? A classification of google unique Web citationsabstractAbstract Although the analysis of citations in the scholarly literature is now an established and relatively well understood part of information science, not enough is known about citations that can be found on the Web. In particular, are there new Web types, and if so, are these trivial or potentially useful for studying or evaluating research communication? We sought evidence based upon a sample of 1,577 Web citations of the URLs or titles of research articles in 64 open‐access journals from biology, physics, chemistry, and computing. Only 25% represented intellectual impact, from references of Web documents (23%) and other informal scholarly sources (2%). Many of the Web/URL citations were created for general or subject‐specific navigation (45%) or for self‐publicity (22%). Additional analyses revealed significant disciplinary differences in the types of Google unique Web/URL citations as well as some characteristics of scientific open‐access publishing on the Web. We conclude that the Web provides access to a new and different type of citation information, one that may therefore enable us to measure different aspects of research, and the research process in particular; but to obtain good information, the different types should be separated. Kayvan Kousha, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | Identifying and characterizing public science-related fears from RSS feedsabstractAbstract A feature of modern democracies is public mistrust of scientists and the politicization of science policy, e.g., concerning stem cell research and genetically modified food. While the extent of this mistrust is debatable, its political influence is tangible. Hence, science policy researchers and science policy makers need early warning of issues that resonate with a wide public so that they can make timely and informed decisions. In this article, a semi‐automatic method for identifying significant public science‐related concerns from a corpus of Internet‐based RSS (Really Simple Syndication) feeds is described and shown to be an improvement on a previous similar system because of the introduction of feed‐based aggregation. In addition, both the RSS corpus and the concept of public science‐related fears are deconstructed, revealing hidden complexity. This article also provides evidence that genetically modified organisms and stem cell research were the two major policy‐relevant science concern issues, although mobile phone radiation and software security also generated significant interest. Mike Thelwall, Rudy Prabowo |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | A comparison of feature selection methods for an evolving RSS feed corpus
Rudy Prabowo, Mike Thelwall |
Inf. Process. Manag. | 2 |
| 2006 | Amanda Spink, Bernhard J. Jansen, Web Search: Public Searching of the Web, Information and Knowledge Management Series, Kluwer Academic Publishers, Dordrecht, The Netherlands, ISBN: 1402022689
Mike Thelwall |
Inf. Process. Manag. | 1 |
| 2006 | Automated Web issue analysis: A nurse prescribing case study
Mike Thelwall, Saheeda Thelwall, Ruth Fairclough |
Inf. Process. Manag. | 1 |
| 2006 | Interpreting social science link analysis research: A theoretical frameworkabstractAbstract Link analysis in various forms is now an established technique in many different subjects, reflecting the perceived importance of links and of the Web. A critical but very difficult issue is how to interpret the results of social science link analyses. It is argued that the dynamic nature of the Web, its lack of quality control, and the online proliferation of copying and imitation mean that methodologies operating within a highly positivist, quantitative framework are ineffective. Conversely, the sheer variety of the Web makes application of qualitative methodologies and pure reason very problematic to large‐scale studies. Methodology triangulation is consequently advocated, in combination with a warning that the Web is incapable of giving definitive answers to large‐scale link analysis research questions concerning social factors underlying link creation. Finally, it is claimed that although theoretical frameworks are appropriate for guiding research, a Theory of Link Analysis is not possible. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Language evolution and the spread of ideas on the Web: A procedure for identifying emergent hybrid word family membersabstractAbstract Word usage is of interest to linguists for its own sake as well as to social scientists and others who seek to track the spread of ideas, for example, in public debates over political decisions. The historical evolution of language can be analyzed with the tools of corpus linguistics through evolving corpora and the Web. But word usage statistics can only be gathered for known words. In this article, techniques are described and tested for identifying new words from the Web, focusing on the case when the words are related to a topic and have a hybrid form with a common sequence of letters. The results highlight the need to employ a combination of search techniques and show the wide potential of hybrid word family investigations in linguistics and social science. Mike Thelwall, Liz Price |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Are raw RSS feeds suitable for broad issue scanning? A science concern case studyabstractAbstract Broad issue scanning is the task of identifying important public debates arising in a given broad issue; really simple syndication (RSS) feeds are a natural information source for investigating broad issues. RSS, as originally conceived, is a method for publishing timely and concise information on the Internet, for example, about the main stories in a news site or the latest postings in a blog. RSS feeds are potentially a nonintrusive source of high‐quality data about public opinion: Monitoring a large number may allow quantitative methods to extract information relevant to a given need. In this article we describe an RSS feed‐based coword frequency method to identify bursts of discussion relevant to a given broad issue. A case study of public science concerns is used to demonstrate the method and assess the suitability of raw RSS feeds for broad issue scanning (i.e., without data cleansing). An attempt to identify genuine science concern debates from the corpus through investigating the top 1,000 “burst” words found only two genuine debates, however. The low success rate was mainly caused by a few pathological feeds that dominated the results and obscured any significant debates. The results point to the need to develop effective data cleansing procedures for RSS feeds, particularly if there is not a large quantity of discussion about the broad issue, and a range of potential techniques is suggested. Finally, the analysis confirmed that the time series information generated by real‐time monitoring of RSS feeds could usefully illustrate the evolution of new debates relevant to a broad issue. Mike Thelwall, Rudy Prabowo, Ruth Fairclough |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Web crawling ethics revisited: Cost, privacy, and denial of serviceabstractAbstract Ethical aspects of the employment of Web crawlers for information science research and other contexts are reviewed. The difference between legal and ethical uses of communications technologies is emphasized as well as the changing boundary between ethical and unethical conduct. A review of the potential impacts on Web site owners is used to underpin a new framework for ethical crawling, and it is argued that delicate human judgment is required for each individual case, with verdicts likely to change over time. Decisions can be based upon an approximate cost‐benefit analysis, but it is crucial that crawler owners find out about the technological issues affecting the owners of the sites being crawled in order to produce an informed assessment. Mike Thelwall, David Stuart |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Web issue analysis: An integrated water resource management case studyabstractAbstract In this article Web issue analysis is introduced as a new technique to investigate an issue as reflected on the Web. The issue chosen, integrated water resource management (IWRM), is a United Nations–initiated paradigm for managing water resources in an international context, particularly in developing nations. As with many international governmental initiatives, there is a considerable body of online information about it: 41,381 hypertext markup language (HTML) pages and 28,735 PDF documents mentioning the issue were downloaded. A page uniform resource locator (URL) and link analysis revealed the international and sectoral spread of IWRM. A noun and noun phrase occurrence analysis was used to identify the issues most commonly discussed, revealing some unexpected topics such as private sector and economic growth. Although the complexity of the methods required to produce meaningful statistics from the data is disadvantageous to easy interpretation, it was still possible to produce data that could be subject to a reasonably intuitive interpretation. Hence Web issue analysis is claimed to be a useful new technique for information science. Mike Thelwall, Katie Vann, Ruth Fairclough |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2005 | Mathematical models for academic webs: Linear relationship or non-linear power law?
Nigel Payne, Mike Thelwall |
Inf. Process. Manag. | 2 |
| 2005 | A modeling approach to uncover hyperlink patterns: the case of Canadian universities
Liwen Vaughan, Mike Thelwall |
Inf. Process. Manag. | 2 |
| 2005 | The clustering power of low frequency words in academic WebsabstractAbstract The value of low frequency words for subject‐based academic Web site clustering is assessed. A new technique is introduced to compare the relative clustering power of different vocabularies. The technique is designed for word frequency tests in large document clustering exercises. Results for the Australian and New Zealand academic Web spaces indicate that low frequency words are useful for clustering academic Web sites along subject lines; removing low frequency words results in sites becoming, on average, less dissimilar to sites from other subjects. Liz Price, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2005 | Text characteristics of English language university Web sitesabstractAbstract The nature of the contents of academic Web sites is of direct relevance to the new field of scientific Web intelligence, and for search engine and topic‐specific crawler designers. We analyze word frequencies in national academic Webs using the Web sites of three English‐speaking nations: Australia, New Zealand, and the United Kingdom. Strong regularities were found in page size and word frequency distributions, but with significant anomalies. At least 26% of pages contain no words. High frequency words include university names and acronyms, Internet terminology, and computing product names: not always words in common usage away from the Web. A minority of low frequency words are spelling mistakes, with other common types including nonwords, proper names, foreign language terms or computer science variable names. Based upon these findings, recommendations for data cleansing and filtering are made, particularly for clustering applications. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2004 | Methods for reporting on the targets of links from national systems of university Web sites
Mike Thelwall |
Inf. Process. Manag. | 1 |
| 2004 | Finding similar academic Web sites with links, bibliometric couplings and colinks
Mike Thelwall, David Wilkinson |
Inf. Process. Manag. | 1 |
| 2004 | Search engine coverage bias: evidence and possible causes
Liwen Vaughan, Mike Thelwall |
Inf. Process. Manag. | 2 |
| 2004 | Do the Web sites of higher rated scholars have significantly more online impact?abstractAbstract The quality and impact of academic Web sites is of interest to many audiences, including the scholars who use them and Web educators who need to identify best practice. Several large‐scale European Union research projects have been funded to build new indicators for online scientific activity, reflecting recognition of the importance of the Web for scholarly communication. In this paper we address the key question of whether higher rated scholars produce higher impact Web sites, using the United Kingdom as a case study and measuring scholars' quality in terms of university‐wide average research ratings. Methodological issues concerning the measurement of the online impact are discussed, leading to the adoption of counts of links to a university's constituent single domain Web sites from an aggregated counting metric. The findings suggest that universities with higher rated scholars produce significantly more Web content but with a similar average online impact. Higher rated scholars therefore attract more total links from their peers, but only by being more prolific, refuting earlier suggestions. It can be surmised that general Web publications are very different from scholarly journal articles and conference papers, for which scholarly quality does associate with citation impact. This has important implications for the construction of new Web indicators, for example that online impact should not be used to assess the quality of small groups of scholars, even within a single discipline. Mike Thelwall, Gareth Harries |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2004 | Webometrics: An introduction to the special issueabstractAbstract Webometrics, the quantitative study of Web phenomena, is a field encompassing contributions from information science, computer science, and statistical physics. Its methodology draws especially from bibliometrics. This special issue presents contributions that both push forward the field and illustrate a wide range of webometric approaches. Mike Thelwall, Liwen Vaughan |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | The connection between the research of a university and counts of links to its web pages: An investigation based upon a classification of the relationships of pages to the research of the host universityabstractAbstract Results from recent advances in link metrics have demonstrated that the hyperlink structure of national university systems can be strongly related to the research productivity of the individual institutions. This paper uses a page categorization to show that restricting the metrics to subsets more closely related to the research of the host university can produce even stronger associations. A partial overlap was also found between the effects of applying advanced document models and separating page types, but the best results were achieved through a combination of the two. Mike Thelwall, Gareth Harries |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Three target document range metrics for university web sitesabstractAbstract Three new metrics are introduced that measure the range of use of a university Web site by its peers through different heuristics for counting links targeted at its pages. All three give results that correlate significantly with the research productivity of the target institution. The directory range model, which is based upon summing the number of distinct directories targeted by each other university, produces the most promising results of any link metric yet. Based upon an analysis of changes between models, it is suggested that range models measure essentially the same quantity as their predecessors but are less susceptible to spurious causes of multiple links and are therefore more robust. Mike Thelwall, David Wilkinson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Graph structure in three national academic Webs: Power laws with anomaliesabstractAbstract The graph structures of three national university publicly indexable Webs from Australia, New Zealand, and the UK were analyzed. Strong scale‐free regularities for page indegrees, outdegrees, and connected component sizes were in evidence, resulting in power laws similar to those previously identified for individual university Web sites and for the AltaVista‐indexed Web. Anomalies were also discovered in most distributions and were tracked down to root causes. As a result, resource driven Web sites and automatically generated pages were identified as representing a significant break from the assumptions of previous power law models. It follows that attempts to track average Web linking behavior would benefit from using techniques to minimize or eliminate the impact of such anomalies. Mike Thelwall, David Wilkinson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Scholarly Use of the Web: What Are the Key Inducers of Links to Journal Web SitesabstractAbstract Web links have been studied by information scientists for at least six years but it is only in the past two that clear evidence has emerged to show that counts of links to scholarly Web spaces (universities and departments) can correlate significantly with research measures, giving some credence to their use for the investigation of scholarly communication. This paper reports on a study to investigate the factors that influence the creation of links to journal Web sites. An empirical approach is used: collecting data and testing for significant patterns. The specific questions addressed are whether site age and site content are inducers of links to a journal's Web site as measured by the ratio of link counts to Journal Impact Factors, two variables previously discovered to be related. A new methodology for data collection is also introduced that uses the Internet Archive to obtain an earliest known creation date for Web sites. The results show that both site age and site content are significant factors for the disciplines studied: library and information science, and law. Comparisons between the two fields also show disciplinary differences in Web site characteristics. Scholars and publishers should be particularly aware that richer content on a journal's Web site tends to generate links and thus the traffic to the site. Liwen Vaughan, Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2002 | Conceptualizing documentation on the Web: An evaluation of different heuristic-based models for counting links between university Web sitesabstractAbstract All known previous Web link studies have used the Web page as the primary indivisible source document for counting purposes. Arguments are presented to explain why this is not necessarily optimal and why other alternatives have the potential to produce better results. This is despite the fact that individual Web files are often the only choice if search engines are used for raw data and are the easiest basic Web unit to identify. The central issue is of defining the Web “document”: that which should comprise the single indissoluble unit of coherent material. Three alternative heuristics are defined for the educational arena based upon the directory, the domain and the whole university site. These are then compared by implementing them on a set of 108 UK university institutional Web sites under the assumption that a more effective heuristic will tend to produce results that correlate more highly with institutional research productivity. It was discovered that the domain and directory models were able to successfully reduce the impact of anomalous linking behavior between pairs of Web sites, with the latter being the method of choice. Reasons are then given as to why a document model on its own cannot eliminate all anomalies in Web linking behavior. Finally, the results from all models give a clear confirmation of the very strong association between the research productivity of a UK university and the number of incoming links from its peers' Web sites. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2001 | Extracting macroscopic information from Web linksabstractAbstract Much has been written about the potential and pitfalls of macroscopic Web‐based link analysis, yet there have been no studies that have provided clear statistical evidence that any of the proposed calculations can produce results over large areas of the Web that correlate with phenomena external to the Internet. This article attempts to provide such evidence through an evaluation of Ingwersen's ( 1998 ) proposed external Web Impact Factor (WIF) for the original use of the Web: the interlinking of academic research. In particular, it studies the case of the relationship between academic hyperlinks and research activity for universities in Britain, a country chosen for its variety of institutions and the existence of an official government rating exercise for research. After reviewing the numerous reasons why link counts may be unreliable, it demonstrates that four different WIFs do, in fact, correlate with the conventional academic research measures. The WIF delivering the greatest correlation with research rankings was the ratio of Web pages with links pointing at research‐based pages to faculty numbers. The scarcity of links to electronic academic papers in the data set suggests that, in contrast to citation analysis, this WIF is measuring the reputations of universities and their scholars, rather than the quality of their publications. Mike Thelwall |
J. Assoc. Inf. Sci. Technol. | 1 |