Jim Jansen

dblp:00/4818 · also Bernard J. Jansen, Bernard Jim Jansen · DBLP profile ↗
← Back
136ranked-venue papers
44as first author
45since 2021 · last 2026
0000-0002-6468-6609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 66 · 33 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 57 · 5 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 7 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 "Do You Need a Little Help?": A Mixed Methods Analysis of 961 Nudges for Blue-Collar and White-Collar Participants During a User Study of Two Digital Information Services
abstract
Usability studies often overlook valuable insights from moderator-to-participant interventions. This study proposes that these interventions, moderator-provided “nudges”, are a rich source of insights on user needs, cognitive load, and system design gaps. We analyzed transcripts from 86 sessions involving two digital library systems and 56 blue-collar (BC) and 30 white-collar (WC) participants, and identified 962 instances of moderator interventions (i.e., nudges). Findings show significant disparities in both the volume and composition of nudges across groups. BC participants required 21 times as many nudges (N=919) as WC participants (N=43). Of the BC participants’ nudges, 737 (80.2%) were system-related, and 182 (19.8%) were user-related. Treating nudges as usability insights for HCI research enables researchers to identify where system scaffolding, such as explicit language, progressive guidance, or simplified workflows, is necessary. The study contributes to inclusive design theory by demonstrating how occupational background influences interventions in user studies and provides design implications for making digital information services more inclusive across occupational groups.
Jinan Y. Azem, Leen F. Al Qadi, Anas Rustom, Joni Salminen, Jim Jansen
DIS5
2026 Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
abstract
As generative AI (GenAI) is increasingly applied in persona development to represent real users, understanding the implications and limitations of this technology is essential for establishing robust practices. This scoping review analyzes how 81 articles (2022-2025) use GenAI techniques for the creation, evaluation, and application of personas. The articles exhibited good level of reproducibility, with 61% of articles sharing resources (personas, code, or datasets). Furthermore, conversational persona interfaces are increasingly provided alongside traditional profiles. However, nearly half (45%) of the articles lack evaluation, and the majority (86%) use only GPT models. In some articles, GenAI use creates a risk of circularity, in which the same GenAI model both generates and evaluates outputs. Our findings also suggest that GenAI seems to reduce the role of human developers in the persona-creation process. To mitigate the associated risks, we propose actionable guidelines for the responsible integration of GenAI into persona development.
Danial Amin 0001, Joni Salminen, Farhan Ahmed, Sonja M. H. Tervola, Sankalp Sethi, Jim Jansen
CHI6
2026 "Pathways to the Metaverse": Exploring the User Experience Mechanisms Driving Technology Acceptance in Virtual Lab Visits with an LLM-powered Avatar
abstract
Metaverse environments combined with large language models (LLMs) enable guided interaction through LLM-powered avatars that function as embodied conversational agents. In our study, we examined how scholars interact with an LLM-powered avatar modeled after a real professor during a virtual reality (VR) tour of a research lab. As little is known about how metaverse characteristics shape the user experience (UX) mechanisms that drive acceptance of such technologies, we conducted a 2 (avatar realism: abstract vs. hyperrealistic) × 2 (immersion: desktop vs. headset-based VR) within-subjects study (N= 30), where academic participants engaged in a virtual lab tour guided by the professor avatar. We conducted path analyses on three conceptual models and, based on the results, proposed the Virtual Lab Acceptance Model (VLAM), which features an experiential path (where perceived immersion increases empathy towards the avatar and task enjoyment) and a rational path (where perceived realism increases avatar credibility and task confidence). Flow states amplify these pathways by strengthening task experiences. Task enjoyment is the strongest predictor of behavioral intention. These findings inform HCI research on metaverse characteristics to drive technology acceptance through UX mechanisms, yielding design implications for developing LLM-powered avatars for virtual labs.
Xinyi Tu 0001, Francesco Biondani, Danial Amin 0001, Angelica Fabillar, Trang Thi Thu Xuan, Huma Bano Adeel, Sonja M. H. Tervola, Carlo Berlingeri, Franco Fummi, Joni Salminen, Jim Jansen
IUI11
2026 Leveraging Participatory Personas for Reflexive Co-design of Personalization in Large Language Models
Kathleen W. Guan, Sarthak Giri, Mohammed Al Owayyed, Jim Jansen, Gayane Sedrakyan, João Fernando Ferreira Gonçalves, Mark de Reuver, Caroline A. Figueroa
UMAP4
2026 AI representing personas representing user groups: Applying the agency theory to examine interaction challenges of conversational personas as decision-making tools
abstract
The proliferation of artificial intelligence (AI) technologies has led to the rise of conversational decision-making support systems, such as dialogue persona systems that provide conversational access to various user segments. For example, product managers can ask personas about features before implementing them, politicians can learn about the needs of local communities through personas, and so on. Nascent research has looked at challenges when users interact with AI personas, but has not framed it as a principal–agent problem, in which the AI represents a persona that itself represents real people in the data. This setting exposes unique interaction challenges that decision makers face when engaging with AI-generated conversational personas, which we examine through a user study with 56 participants using AI-generated conversational personas. Our results indicate seven interaction challenges: (1) Hidden Information, (2) Hidden Personas, (3) Hidden UI, (4) Lack of AI Agency, (5) AI’s Selective Attention, (6) Confusing Distributional Information, and (7) Conversational Cold Start that we conceptually link with agency theory. We discuss how the interaction challenges could be alleviated and suggest directions for future work. • This article explores how decision makers interact with AI-generated conversational personas derived from real survey data. • It conducts a comparative study between conversational personas and traditional profile personas in decision-support contexts. • The study employs a think-aloud user study with 56 participants to capture interaction experiences and challenges. • It identifies seven specific interaction challenges unique to conversational personas that may hinder effective decision making. • The article provides insights and recommendations for designing conversational decision support systems using AI-generated personas.
Joni Salminen, Soon-Gyo Jung, Ilkka Kaate, Trang Thi Thu Xuan, Jinan Y. Azem, Kholoud Khalil Aldous, Danial Amin 0001, Jim Jansen
Decis. Support Syst.8
2026 AI-generated personas: Representing user needs with generative AI models
Joni Salminen, Lene Nielsen, Ali Farooq 0001, Jim Jansen
Int. J. Hum. Comput. Stud.4
2025 Are We Still Under-Serving the Underserved?: An Analysis of 56 Blue-Collar Workers Using 2 Online Information Services
abstract
We examined the accessibility of online information services (OISs) for underserved communities through a user study involving 56 blue collar participants interacting with a website and an app for five tasks. The blue collar participants were generally unsuccessful on both platforms, with 12.7% (n=7) unable to successfully complete any tasks; a hundred percent required at least minor assistance. Participants were also inefficient, taking 28.62 more steps than optimal (143.1%) on the website and 10.41 more steps (47.3%) on the app. Time inefficiency was also noteworthy, with 535.76 more seconds than optimal (248.0%) on the website and 266.55 more seconds (142.4%) on the app. Though still poor, the app yielded better outcomes with higher success rates and usability ratings. Digital proficiency correlated with success on both platforms, which is good news as this is addressable by OIS providers. Qualitative analysis revealed that many in this underserved population were unaware that these valuable OISs were available to them. Findings underscore the need for OIS providers to prioritize targeted outreach to inform underserved communities that OISs are open and welcoming. Designing OISs with accessibility and simplicity targeted for mobile devices is crucial for bridging the digital literacy gap and empowering underserved communities to engage effectively with OISs.
Jinan Y. Azem, Joni Salminen, Kholoud Khalil Aldous, Fatou Gueye, Jim Jansen
Conference on Designing Interactive Systems5
2025 When Personas Talk to You: Evaluating the Evolution of User Personas from Static Profiles to Conversational User Interfaces
abstract
The development of persona systems provides a possibility for end users to interact with different persona modalities. In a 54-participant randomized controlled experiment, we compare two persona interaction modalities, document and dialogue personas, both generated using AI approaches from survey data. Overall, dialogue personas appear to be perceived more favorably than document personas. However, document personas exhibit a wider range of perceptions, suggesting that experiences with document personas are more polarizing among users. The document personas had higher transparency and were perceived as more complete, but the task completion was perceived as more difficult, although the task success rate was higher. The dialogue personas were perceived as more usable, with a higher System Usability Scale score, and more enjoyable. Our findings provide critical insights into the increasingly important area of persona interaction modalities and the broad paradigm of human-persona interaction.
Ilkka Kaate, Joni Salminen, Soon-Gyo Jung, Trang Thi Thu Xuan, Jinan Y. Azem, João M. Santos 0001, Jim Jansen
Conference on Designing Interactive Systems7
2025 Representing Religious Practices via AI-Generated Personas: A Case Study of Ramadan Behaviors from Four Predominantly Muslim Countries
abstract
With nearly two billion Muslims worldwide, designing technology that serves their needs is a significant task, especially during Ramadan, the season of spiritual reflection and change in lifestyle. We present a data-driven approach to generating personas representing differing religious views of Ramadan, based on a large-scale survey conducted in four Muslim-majority countries: Egypt, Indonesia, the United Arab Emirates, and Saudi Arabia. Our approach furthers the representation of underrepresented groups by embedding religious and cultural factors in persona development, aligning with the principles of value-sensitive design. Via correlation analysis, we identified seven distinct groups that reflect diverse practices and values during the Islamic holy month. These findings informed persona creation, capturing variations in spiritual engagement, digital media consumption, and the planning of Ramadan. The resulting personas provide actionable insights for developers designing inclusive applications such as charitable platforms and ecommerce systems aligned with Muslim values. We discuss the implications of employing personas for culturally aware system design and demonstrate how AI-generated personas can support inclusive design in religiously diverse settings. Ultimately, this work contributes to the growing research on human-centered technologies for value-aligned and context-aware systems.
Leen F. Al Qadi, Soon-Gyo Jung, Danial Amin 0001, Amani Alabed, Joni Salminen, Jim Jansen
AICCSA6
2025 Using AI for User Representation: An Analysis of 83 Persona Prompts
abstract
We analyzed 83 persona prompts from 27 research articles that used large language models (LLMs) to generate user personas. Findings show that the prompts predominantly generate single personas. Several prompts express a desire for short or concise persona descriptions, which deviates from the tradition of creating rich, informative, and rounded persona profiles. Text is the most common format for generated persona attributes, followed by numbers. Text and numbers are often generated together, and demographic attributes are included in nearly all generated personas. Researchers use up to 12 prompts in a single study, though most research uses a small number of prompts. Comparison and testing multiple LLMs is rare. More than half of the prompts require the persona output in structured format, such as JSON, and $74 \%$ of the prompts insert data or dynamical variables. We discuss the implications of increased use of computational personas for user representation.
Joni Salminen, Danial Amin 0001, Jim Jansen
AICCSA3
2025 What is User Engagement?: A Systematic Review of 241 Research Articles in Human-Computer Interaction and Beyond
abstract
User engagement (UE) is widely discussed in HCI articles, but its definition, reliability, and application remain elusive. This research conducts a systematic literature review of 241 articles from 1993 to 2023 to analyze how UE is defined and measured within the domain of HCI. Our findings reveal significant definitional inconsistencies that hinder UE's practical application in HCI research and system design. Based on our findings, we recommend using UE as a categorical label rather than a unified construct until more systematic frameworks are established. We also highlight the need for divergent views of UE across HCI research communities as a valuable avenue to pursue. This divergent view approach can help HCI researchers focus on specific, measurable aspects of UE that align with specific community practices and norms. Our findings also suggest that until such a framework emerges, researchers should be aware of its limitations when using UE as a research construct.
Jim Jansen, Kathleen W. Guan, Joni Salminen, Kholoud Khalil Aldous, Soon-Gyo Jung
CHI1
2025 "You Always Get an Answer": Analyzing Users' Interaction with AI-Generated Personas Given Unanswerable Questions and Risk of Hallucination
abstract
We investigated the presence and acceptance of hallucinations (i.e., accidental misinformation) of an AI-generated persona system that leverages large language models for persona creation from survey data in a 54-user within-subjects experiment. After interacting with the personas, users were given a task to ask the personas a series of questions, including an unanswerable question, meaning the personas lacked the data to answer the question. The AI-generated persona system provided a plausible but incorrect answer half (52%) of the time, and more than half of the time (57%), the users accepted the incorrect answer, and the rest of the time, users answered the unanswerable question correctly (no answer). We found that when the AI-generated persona hallucinated, the user was significantly more likely to answer the unanswerable question incorrectly. Also, for genders separately, when the AI-generated persona hallucinated, it was significantly more likely for the female user and the male users to answer the unanswerable question incorrectly. We identified four themes in the AI-generated persona's answers and found that users perceive AI-generated persona's answers as long and unclear for the unanswerable question. Findings imply that personas leveraging LLMs require guardrails to ensure that personas clearly state the possibility of data restrictions and hallucinations when asked unanswerable questions.
Ilkka Kaate, Joni Salminen, Soon-Gyo Jung, Trang Thi Thu Xuan, Essi Häyhänen, Jinan Y. Azem, Jim Jansen
IUI7
2025 The 'fourth wall' and other usability issues in AI-generated personas: comparing chat-based and profile personas
abstract
Large Language Models (LLMs) are emerging as a powerful tool for AI-generated personas. This study evaluates the usability of AI-generated personas, comparing chat and profile formats. The findings indicate chat personas tend to be perceived more favourably, and profile personas exhibit greater variability in user perception. The increased difficulty and longer dwell time experienced by users with the profile persona, despite negative usability metrics, paradoxically resulted in better task performance. Usability issues indicate that many current limitations of AI, including verbosity, hallucinations, and empty rhetoric which was described as the persona having ‘no soul’, are inherited in AI-generated chat personas. However, there are also new issues. For one, the risk of information overload in an AI-generated profile persona implies that the AI does not consider human users’ cognitive limitations when designing the persona (but usability scores for profile personas increase with dwell time, implying that users get used to the longer format the more time they spend). Another is the ‘fourth wall’ effect of AI-generated chat personas in which the user feels they are talking to someone describing the persona rather than the persona itself. Future work could address the usability paradox and the fourth wall effect of using personas.CCS CONCEPTS Human-centered computing Human computer interaction (HCI)
Ilkka Kaate, Joni Salminen, Soon-Gyo Jung, João M. Santos 0001, Essi Häyhänen, Trang Xuan, Jinan Y. Azem, Jim Jansen
Behav. Inf. Technol.8
2025 Is Deepfake Diversity Real? Analyzing the Diversity of Deepfake Avatars
abstract
Deepfake technology is increasingly integrated into global mobile and web services when human representation is not feasible or cost-effective. Our analysis of 202 deepfake avatars from three deepfake providers reveals significant demographic disparities with 18 out of 48 possible demographic groups unrepresented. Deepfake avatars' gender distribution was nearly balanced (49.01% male, 50.99% female), but older age groups (Baby Boomers and Silent Generation) were substantially underrepresented by 64.36% and 76.24%, respectively, relative to the average number of all deepfake avatars. Differences in language representation were present in deepfake avatar providers with only 1.06% of global languages covered. The findings indicate that current deepfake technology lacks diversity, primarily favoring young white individuals, neglecting older demographics, Asians, and Middle Eastern populations, with underrepresentation of 40.59% and 52.48%, respectively, relative to the average number of all deepfake avatars. Only 15.27% of deepfake avatars portray any occupational characteristics. Addressing these diversity gaps is crucial for better serving varied user groups and warrants attention from deepfake providers and caution from those using deepfakes.
Ilkka Kaate, Joni Salminen, Reham Al Tamime, Soon-Gyo Jung, Jim Jansen
Expert Syst. Appl.5
2025 Generative AI personas considered harmful? Putting forth twenty challenges of algorithmic user representation in human-computer interaction
abstract
• Shows how GenAI fundamentally transforms existing persona development issues through evolutionary amplification rather than creating entirely new problems, with traditional biases becoming algorithmic discrimination and manual inconsistencies becoming convincing AI hallucinations. • Reveals how traditional limitations manifest differently in GenAI contexts across transparency, fairness, reliability, and control domains, with expert validation showing 60% of challenges are more problematic for GenAIPs than conventional approaches. • Documents how GenAI transforms not just technical challenges but harm distribution, with persona developers facing operational complexity while target user groups bear severe consequences through systematic misrepresentation and exclusion. • Provides evidence that while GenAIPs appear to solve traditional limitations, they transform existing challenges into more complex forms requiring novel validation approaches and human-AI collaboration frameworks for responsible implementation. Generative AI personas (GenAIPs) promise user-centred design efficiency, but their impact on different persona challenges remains unexplored. Inspired by Dijkstra’s classic essay on harmful programming constructs, we analyze twenty challenges in persona development using Human-Centered AI principles. Through literature review and expert survey (n=17), we find that GenAIPs transform rather than eliminate traditional persona challenges. Experts rated all challenges as problematic for GenAIPs (M > 4.0), with the highest concerns for hallucinations (M=5.94), over-sanitization (M=5.82), and lack of standardization (M=5.59). 12 out of 20 challenges are considered more problematic for GenAIPs than conventional personas, particularly bias amplification, validation challenges, and accessibility without expertise. We provide HCAI-grounded guidelines demonstrating that effective GenAIP implementation requires human-AI collaboration rather than automation and prioritizing user welfare over technical efficiency.
Danial Amin 0001, Joni Salminen, Jim Jansen, Joon Gi Shin, Daehyun Kim 0005
Int. J. Hum. Comput. Stud.3
2025 PersonaCraft: Leveraging language models for data-driven persona development
Soon-Gyo Jung, Joni Salminen, Kholoud Khalil Aldous, Jim Jansen
Int. J. Hum. Comput. Stud.4
2025 Demographics do not matter?: Exploring the impact of gender and ethnicity on users' identification with AI-generated personas
abstract
Demographics are considered foundational information in most persona profiles. However, the effect of persona ethnicity and gender on designers’ identification with the persona has limited evaluation in the human-computer interaction literature. We conducted a study with 64 professional designers from the United States, Indian, Korean, and Mexican nationalities to investigate the effects of AI-generated persona ethnicity and gender on persona identification. The personas were created using Generative AI in the persona narratives and the persona video creation. The contribution of this work is that, against assumptions, neither persona ethnicity nor gender play a major role in persona identification among designers with different ethnic backgrounds. While there were some insinuations of ethnicity and gender in the open-ended feedback from the designers, the emergent qualitative themes describing persona identification were overwhelmingly universal and applicable regardless of ethnicity or gender. This implies that professional designers can effectively use personas with different demographic backgrounds, and effects of demographic attributes in personas leading to stereotyping are less impactful than presumed.
Ilkka Kaate, Joni Salminen, Soon-Gyo Jung, João M. Santos 0001, Kholoud Khalil Aldous, Essi Häyhänen, Jinan Y. Azem, Jim Jansen
Int. J. Hum. Comput. Stud.8
2024 "There's Something About Noura": Exploring Think-Aloud Reasonings for Users' Persona Choice in a Design Task
abstract
Stakeholders like designers use personas to learn about users. After persona development, stakeholders are usually presented with a persona set. However, there is little research on how stakeholders select a persona from a persona set. A think-aloud analysis with 37 stakeholders who were asked to select a persona for a content design task reveals that persona selection is influenced by comparative, non-comparative, and subjective elements. Persona choice is often made with task compatibility in mind: interests, professions, and education were important contextual factors in our focal task. Storifying is commonly applied by stakeholders, reflecting personas’ narrative nature. The persona’s picture is often evoked, in addition to nationality and name, though demographics do not play a decisive role. Stakeholders refer to a host of persona attributes when explicating their persona choice. Overall, reasonings for persona choice are multifaceted and individualistic, as we might expect given the information-richness of personas.
Sercan Sengün, Joni Salminen, Soon-Gyo Jung, Kholoud Khalil Aldous, Jim Jansen
Conference on Designing Interactive Systems5
2024 Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
abstract
Large language models (LLMs) can generate personas based on prompts that describe the target user group. To understand what kind of personas LLMs generate, we investigate the diversity and bias in 450 LLM-generated personas with the help of internal evaluators (n=4) and subject-matter experts (SMEs) (n=5). The research findings reveal biases in LLM-generated personas, particularly in age, occupation, and pain points, as well as a strong bias towards personas from the United States. Human evaluations demonstrate that LLM persona descriptions were informative, believable, positive, relatable, and not stereotyped. The SMEs rated the personas slightly more stereotypical, less positive, and less relatable than the internal evaluators. The findings suggest that LLMs can generate consistent personas perceived as believable, relatable, and informative while containing relatively low amounts of stereotyping.
Joni Salminen, Chang Liu 0007, Wenjing Pian, Jianxing Chi, Essi Häyhänen, Jim Jansen
CHI6
2024 Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas
abstract
Seungjong Sun, Eungu Lee, Seo Yeon Baek, Seunghyun Hwang, Wonbyung Lee, Dongyan Nan, Bernard J Jansen, Jang Hyun Kim. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Seungjong Sun, Eungu Lee, Seo Yeon Baek, Seunghyun Hwang, Wonbyung Lee, Dongyan Nan, Jim Jansen, Jang-Hyun Kim 0001
EMNLP7
2024 Using Cipherbot: An Exploratory Analysis of Student Interaction with an LLM-Based Educational Chatbot
abstract
Cipherbot, an educational chatbot using large language models to answer student questions concerning learning materials uploaded by the educator, was pilot tested in a classroom setting. Forty-four students used Cipherbot for seven weeks, sending 8077 messages. The average number of messages sent per student was 184 (SD = 80), with an average length of 98 characters (SD = 80). The engagement followed a non-normal distribution, with few power users, implying that most students are still hesitant to adopt tools like Cipherbot. Cipherbot was able to answer 82.5% of the student questions, demonstrating a scalable ability to address students' learning queries, with some room for improvement.
Joni Salminen, Soon-Gyo Jung, Johanne Medina, Kholoud Khalil Aldous, Jinan Y. Azem, Waleed Akhtar, Jim Jansen
L@S7
2024 Evaluating LLM-Generated Topics from Survey Responses: Identifying Challenges in Recruiting Participants through Crowdsourcing
abstract
The evolution of generative artificial intelligence (AI) technologies, particularly large language models (LLMs), has lead to consequences for the field of Human-Computer Interaction (HCI) in areas such as personalization, predictive analytics, automation, and data analysis. This research aims to evaluate LLM-generated topics derived from survey responses in comparison with topics suggested by humans, particularly participants recruited through a crowdsourcing experiment. We present an evaluation results to compare LLM-generated topics with human-generated topics in terms of Quality, Usefulness, Accuracy, Interestingness, and Completeness. This involves three stages: (1) Design and Generate Topics with an LLM (OpenAI’s GPT-4); (2) Crowdsourcing Human-Generated Topics; and (3) Evaluation of Human-Generated Topics and LLM-Generated Topics. However, a feasibility study with 33 crowdworkers indicated challenges in using participants for LLM evaluation, particularly in inviting humans participants to suggest topics based on open-ended survey answers. We highlight several challenges in recruiting crowdsourcing participants for generating topics from survey responses. We recommend using well-trained human experts rather than crowdsourcing to generate human baselines for LLM evaluation.
Reham Al Tamime, Joni Salminen, Soon-Gyo Jung, Jim Jansen
VL/HCC4
2024 Leveraging Personas for Social Impact: A Review of Their Applications to Social Good in Design
abstract
Personas inform design by representing diverse user needs. Since their initial application in commercial technology contexts, personas have been adopted in several research domains for public good, such as health, accessibility, politics and civic society, education, sustainability, cybersecurity, and criminology. In this review paper, we analyzed 58 research studies that created personas in these domains, referred to as Personas for Social Good (PFSG). In most studies, PFSG was primarily exploratory and focused on initial methodology development. More than half (59%) neglected to discuss concerns with stereotyping or evaluate how personas contributed to improving social concerns in their respective domains. To facilitate a shift towards more socially conscious persona applications, we identified and critically examined the most comprehensive PFSG domain applications in our sample. Based on their strengths, we present an ecological framework to guide researchers in holistically aligning persona creation efforts with addressing critical social challenges.
Kathleen W. Guan, Joni Salminen, Soon-Gyo Jung, Jim Jansen
Int. J. Hum. Comput. Interact.4
2024 Beyond Avatar Coolness: Exploring the Effects of Avatar Attributes on Continuance Intention to Play Massively Multiplayer Online Role-Playing Games
abstract
This research investigates players’ continuance intentions to play massively multiplayer online role-playing games (MMORPGs) by constructing a model based on the concepts of avatar coolness (i.e., avatar attractiveness, avatar originality, and avatar subculture appeal), social identity theory, and flow theory. Analyzing survey-based data from 375 Korean MMORPG players, we found that avatar attractiveness, avatar originality, and avatar subculture appeal were positively related to avatar coolness. In addition, avatar originality positively affects avatar subculture appeal. Moreover, avatar coolness positively affects the continuance intention to play MMORPGs via avatar identification and flow state. This study is the first to develop avatar coolness and explore its role in affecting the intention to play MMORPGs. To offer “cool” avatars, the implications are those game designers should continually update avatars for freshness based on current trends, provide a variety of skins for personalization for user preferences, and offer avatars that are visually appealing to the gaming population. These require continual assessments of the MMORPG player population.
Dongyan Nan, Seungjong Sun, Jim Jansen, Jang-Hyun Kim 0001
Int. J. Hum. Comput. Interact.3
2023 Measuring Engagement Through Remote Interactions of Customers: Introducing METRIC
abstract
In this article, we present METRIC. Measuring Engagement Through Remote Interactions of Customers (METRIC) (https://metric.qcri.org/) is a tool for collecting, measuring, analyzing, and reporting the engagement of online systems through actual interactions of customers or users, either remote or in the lab. METRIC enables system stakeholders to enhance their understanding of their audience, customer, or users' actual behavior on pages, images, videos, interfaces, and online systems, including the gaze and interaction with sub-elements on a page within a system or comparisons via A/B testing. Along with eye-tracking devices, METRIC uses a webcam-based eye-tracking JavaScript library for the ability to monitor the users' real visual attention during interaction with the online system. METRIC provides sophisticated reporting features throughout the collecting, measuring, and analyzing process. METRIC can also be deployed in user experiments toward the design of better cooperation technologies, primarily due to its online nature.
Jinan Y. Azem, Joni Salminen, Soon-Gyo Jung, Jim Jansen
ISNCC4
2023 What really matters?: characterising and predicting user engagement of news postings using multiple platforms, sentiments and topics
abstract
This research characterises user engagement of approximately 3,000,000 news postings of 53 news outlets and 50,000,000 associated user comments during 8 months on 5 social media platforms (i.e. Facebook, Instagram, Twitter, YouTube, and Reddit). We investigate the effect of sentiments and topics on user engagement across four levels of user engagement expressions (i.e. views, likes, comments, cross-platform posting). We find that sentiments and topics differ by both news outlets and social media platforms, and both sentiments and topics by the four levels of user engagement expression. Finally, we predict a volume of four user engagement levels for given news content, with an 83% maximum average F1-score for the external posting of news articles from one platform to another using language and metadata features. Implications are that news outlets can benefit by developing a platform, sentiment and topic, and strategies to best achieve user engagement objectives.
Kholoud Khalil Aldous, Jisun An, Jim Jansen
Behav. Inf. Technol.3
2023 Fair compensation of crowdsourcing work: the problem of flat rates
abstract
Compensating crowdworkers for their research participation often entails paying a flat rate to all participants, regardless of the amount of time they spend on the task or skill level. If the actual time required varies considerably between workers, flat rates may yield unfair compensation. To study this matter, we analyzed three survey studies with varying complexity. Based on the United Kingdom minimum wage and actual task completion times, we found that more than 3 in 4 (76.5%) of the crowdworkers studied were paid more than the intended hourly wage, and around one in four (23.5%) was paid less than the intended hourly wage when using a flat rate compensation model based on estimated completion time. The results indicate that the popular flat rate model falls short as a form of equitable remuneration, when perceiving fairness in the form of compensating one’s time. Flat rate compensation would not be problematic if the workers’ completion times were similar, but this is not the case in reality, as skills and motivation can vary. To overcome this problem, the study proposes three alternative compensation models: Compensation by Normal Distribution, Multi-Objective Fairness, and Post-Hoc Bonuses.
Joni Salminen, Ahmed Mohamed Sayed Kamel, Soon-Gyo Jung, Mekhail Mustak, Jim Jansen
Behav. Inf. Technol.5
2023 Will they take this offer? A machine learning price elasticity model for predicting upselling acceptance of premium airline seating
abstract
Employing customer information from one of the world's largest airline companies, we develop a price elasticity model (PREM) using machine learning to identify customers likely to purchase an upgrade offer from economy to premium class and predict a customer's acceptable price range. A simulation of 64.3 million flight bookings and 14.1 million email offers over three years mirroring actual data indicates that PREM implementation results in approximately 1.12 million (7.94%) fewer non-relevant customer email messages, a predicted increase of 72,200 (37.2%) offers accepted, and an estimated $72.2 million (37.2%) of increased revenue. Our results illustrate the potential of automated pricing information and targeting marketing messages for upselling acceptance. We also identified three customer segments: (1) Never Upgrades are those who never take the upgrade offer, (2) Upgrade Lovers are those who generally upgrade, and (3) Upgrade Lover Lookalikes have no historical record but fit the profile of those that tend to upgrade. We discuss the implications for airline companies and related travel and tourism industries.
Saravanan Thirumuruganathan, Noora Al Emadi, Soon-Gyo Jung, Joni Salminen, Dianne Ramirez Robillos, Jim Jansen
Inf. Manag.6
2023 The realness of fakes: Primary evidence of the effect of deepfake personas on user perceptions in a design task
abstract
Deepfakes, realistic portrayals of people that do not exist, have garnered interest in research and industry. Yet, the contributions of deepfake technology to human-computer interaction remain unclear. One possible value of deepfake technology is to create more immersive user personas. To test this premise, we use a commercial-grade service to generate three deepfake personas (DFs). We also create counterparts of the same persona in two traditional modalities: classic and narrative personas. We then investigate how persona modality affects the perceptions and task performance of the persona user. Our findings show that the DFs were perceived as less empathetic, credible, complete, clear, and immersive than other modalities. Participants also indicated less willingness to use the DFs and less sense of control, but there were no differences in task performance. We also found a strong correlation between the uncanny valley effect and other user perceptions, implying that the tested deepfake technology might lack maturity for personas, negatively affecting user experience. Designers might also be accustomed to using traditional persona profiles. Further research is needed to investigate the potential and downsides of DFs.
Ilkka Kaate, Joni Salminen, João M. Santos 0001, Soon-Gyo Jung, Rami Olkkonen, Jim Jansen
Int. J. Hum. Comput. Stud.6
2022 Use Cases for Design Personas: A Systematic Review and New Frontiers
abstract
Personas represent the needs of users in diverse populations and impact design by endearing empathy and improving communication. While personas have been lauded for their benefits, we could locate no prior review of persona use cases in design, prompting the question: how are personas actually used to achieve these benefits? To address this question, we review 95 articles containing persona application across multiple domains, and identify software development, healthcare, and higher education as the top domains that employ personas. We then present a three-stage design hierarchy of persona usage to describe how personas are used in design tasks. Finally, we assess the increasing trend of persona initiatives aimed towards social good rather than solely commercial interests. Our findings establish a roadmap of best practices for how practitioners can innovatively employ personas to increase the value of designs and highlight avenues of using personas for socially impactful purposes.
Joni Salminen, Kathleen W. Guan, Soon-Gyo Jung, Jim Jansen
CHI4
2022 Developing Persona Analytics Towards Persona Science
abstract
Much of the reported work on personas suffers from the lack of empirical evidence. To address this issue, we introduce Persona Analytics (PA), a system that tracks how users interact with data-driven personas. PA captures users’ mouse and gaze behavior to measure users’ interaction with algorithmically generated personas and use of system features for an interactive persona system. Measuring these activities grants an understanding of the behaviors of a persona user, required for quantitative measurement of persona use to obtain scientifically valid evidence. Conducting a study with 144 participants, we demonstrate how PA can be deployed for remote user studies during exceptional times when physical user studies are difficult, if not impossible.
Joni Salminen, Soon-Gyo Jung, Jim Jansen
IUI3
2022 Using artificially generated pictures in customer-facing systems: an evaluation study with data-driven personas
abstract
We conduct two studies to evaluate the suitability of artificially generated facial pictures for use in a customer-facing system using data-driven personas. STUDY 1 investigates the quality of a sample of 1,000 artificially generated facial pictures. Obtaining 6,812 crowd judgments, we find that 90% of the images are rated medium quality or better. STUDY 2 examines the application of artificially generated facial pictures in data-driven personas using an experimental setting where the high-quality pictures are implemented in persona profiles. Based on 496 participants using 4 persona treatments (2 × 2 research design), findings of Bayesian analysis show that using the artificial pictures in persona profiles did not decrease the scores for Authenticity, Clarity, Empathy, and Willingness to Use of the data-driven personas.
Joni Salminen, Soon-Gyo Jung, Ahmed Mohamed Sayed Kamel, João M. Santos 0001, Jim Jansen
Behav. Inf. Technol.5
2022 Time-varying effects of search engine advertising on sales-An empirical investigation in E-commerce
Kang Zhao 0001, Daniel Dajun Zeng, Jim Jansen
Decis. Support Syst.4
2022 How does varying the number of personas affect user perceptions and behavior? Challenging the 'small personas' hypothesis!
abstract
Studies in human-computer interaction recommend creating fewer than ten personas, based on stakeholders’ limitations to cognitively process and use personas. However, no existing studies offer empirical support for having fewer rather than more personas. Investigating this matter, thirty-seven participants interacted with five and fifteen personas using an interactive persona system, choosing one persona to design for. Our study results from eye-tracking and survey data suggest that when using interactive persona systems, the number of personas can be increased from the conventionally suggested ‘less than ten’, without significant negative effects on user perceptions or task performance, and with the positive effects of increasing engagement with the personas, having a more diverse representation of the end-user population, as well as users accessing personas from more varied demographic groups for a design task. Using the interactive persona system, users adjusted their information processing style by spending less time on each persona when presented with fifteen personas, while still absorbing a similar amount of information than with five personas, implying that more efficient information processing strategies are applied with more personas. The results highlight the importance of designing interactive persona systems to support users’ browsing of more personas.
Joni Salminen, Soon-Gyo Jung, Lene Nielsen, Sercan Sengün, Jim Jansen
Int. J. Hum. Comput. Stud.5
2022 Engineers, Aware! Commercial Tools Disagree on Social Media Sentiment: Analyzing the Sentiment Bias of Four Major Tools
abstract
Large commercial sentiment analysis tools are often deployed in software engineering due to their ease of use. However, it is not known how accurate these tools are, and whether the sentiment ratings given by one tool agree with those given by another tool. We use two datasets - (1) NEWS consisting of 5,880 news stories and 60K comments from four social media platforms: Twitter, Instagram, YouTube, and Facebook; and (2) IMDB consisting of 7,500 positive and 7,500 negative movie reviews - to investigate the agreement and bias of four widely used sentiment analysis (SA) tools: Microsoft Azure (MS), IBM Watson, Google Cloud, and Amazon Web Services (AWS). We find that the four tools assign the same sentiment on less than half (48.1%) of the analyzed content. We also find that AWS exhibits neutrality bias in both datasets, Google exhibits bi-polarity bias in the NEWS dataset but neutrality bias in the IMDB dataset, and IBM and MS exhibit no clear bias in the NEWS dataset but have bi-polarity bias in the IMDB dataset. Overall, IBM has the highest accuracy relative to the known ground truth in the IMDB dataset. Findings indicate that psycholinguistic features - especially affect, tone, and use of adjectives - explain why the tools disagree. Engineers are urged caution when implementing SA tools for applications, as the tool selection affects the obtained sentiment labels.
Soon-Gyo Jung, Joni Salminen, Jim Jansen
Proc. ACM Hum. Comput. Interact.3
2022 Can Unhappy Pictures Enhance the Effect of Personas? A User Experiment
abstract
There has been little research into whether a persona's picture should portray a happy or unhappy individual. We report a user experiment with 235 participants, testing the effects of happy and unhappy image styles on user perceptions, engagement, and personality traits attributed to personas using a mixed-methods analysis. Results indicate that the participant's perceptions of the persona's realism and pain point severity increase with the use of unhappy pictures. In contrast, personas with happy pictures are perceived as more extroverted, agreeable, open, conscientious, and emotionally stable. The participants’ proposed design ideas for the personas scored more lexical empathy scores for happy personas. There were also significant perception changes along with the gender and ethnic lines regarding both empathy and perceptions of pain points. Implications are the facial expression in the persona profile can affect the perceptions of those employing the personas. Therefore, persona designers should align facial expressions with the task for which the personas will be employed. Generally, unhappy images emphasize realism and pain point severity, and happy images invoke positive perceptions.
Joni Salminen, Sercan Sengün, João M. Santos 0001, Soon-Gyo Jung, Jim Jansen
ACM Trans. Comput. Hum. Interact.5
2021 Picturing It!: The Effect of Image Styles on User Perceptions of Personas
abstract
Though photographs of real people are typically used to portray personas, there is little research into the potential advantages or disadvantages of using such images, relative to other image styles. We conducted an experiment with 149 participants, testing the effects of six different image styles on user perceptions and personality traits that are attributed to personas by the participants. Results show that perceptions of clarity, completeness, consistency, credibility, and empathy for a persona increase with picture realism. Personas with more realistic pictures are also perceived as more agreeable, open, and emotionally stable, with higher confidence in these assessments. We also find evidence of the uncanny valley effect, with realistic cartoon personas experiencing a decrease in the user perception scores.
Joni Salminen, Soon-Gyo Jung, João M. Santos 0001, Ahmed Mohamed Sayed Kamel, Jim Jansen
CHI5
2021 The Problem of Majority Voting in Crowdsourcing with Binary Classes
Joni Salminen, Ahmed Mohamed Sayed Kamel, Soon-Gyo Jung, Jim Jansen
ECSCW4
2021 Taking Back Control of Social Media Feeds with Take Back Control
abstract
Controlling the quality of social media feeds poses an issue for many users. Platforms such as Twitter give users some options to influence their feeds. Still, the selection of content predominantly relies on implicit rather than explicit user actions, as manual options for "cleaning the feed" are often cumbersome and difficult to use for most users. Here, we present Take Back Control, a web browser extension that gives users control to hide undesirable content from their social media feeds. The extension combines JavaScript (for hiding the content) and machine learning (for deciding what content to hide). Our current demonstration includes three filter types: Toxic, Political, and Negative content, with a possibility to add more filters, all of this with the overarching aim of helping end users control the information visible in their social media feeds.
Joni Salminen, Juan Corporan, Soon-Gyo Jung, Jim Jansen
INISTA4
2021 Think-Aloud Surveys - A Method for Eliciting Enhanced Insights During User Studies
Lene Nielsen, Joni Salminen, Soon-Gyo Jung, Jim Jansen
INTERACT (5)4
2021 Helping Professionals Select Persona Interview Questions Using Natural Language Processing
Joni Salminen, Kamal Chhirang, Soon-Gyo Jung, Jim Jansen
INTERACT (3)4
2021 Persona analytics: Analyzing the stability of online segments and content interests over time using non-negative matrix factorization
Jim Jansen, Soon-Gyo Jung, Shammur Absar Chowdhury, Joni Salminen
Expert Syst. Appl.1
2021 A Survey of 15 Years of Data-Driven Persona Development
abstract
Data-driven persona development unifies methodologies for creating robust personas from the behaviors and demographics of user segments. Data-driven personas have gained popularity in human-computer interaction due to digital trends such as personified big data, online analytics, and the evolution of data science algorithms. Even with its increasing popularity, there is a lack of a systematic understanding of the research on the topic. To address this gap, we review 77 data-driven persona research articles from 2005–2020. The results indicate three periods: (1) Quantification (2005–2008), which consists of the first experiments with data-driven methods, (2) Diversification (2009–2014), which involves more pluralistic use of data and algorithms, and (3) Digitalization (2015–present), marked by the abundance of online user data and the rapid development of data science algorithms and software. Despite consistent work on data-driven personas, there remain many research gaps concerning (a) shared resources, (b) evaluation methods, (c) standardization, (d) consideration for inclusivity, and (e) risk of losing in-depth user insights. We encourage organizations to realistically assess their data-driven persona development readiness to gain value from data-driven personas.
Joni Salminen, Kathleen W. Guan, Soon-Gyo Jung, Jim Jansen
Int. J. Hum. Comput. Interact.4
2021 How Does Personification Impact Ad Performance and Empathy? An Experiment with Online Advertising
abstract
This research explores the value of personas for supporting professional advertisers to design adverts for social media. We test if a personified user group (PUG), when provided to online ad designers, results in better ad performance than when using a non-personified user group (NUG) that had no face picture or name. Our experiment has 30 participants that created Facebook ads using both PUG and NUG. We found that using PUG did increase advertising click performance of ads created by people who are more experienced with ads and personas. Moreover, an analysis of the ad texts showed that the use of PUG increased the empathy of the created ads, supporting the foundational empathy benefit cited in HCI literature. However, the use of PUG did not significantly increase purchase intent. The results imply that using PUG for online ad design evokes more empathy and improves click-through performance. More empathetic ads can have a positive impact on social media users, given that they appear to increase relevance.
Joni Salminen, Ilkka Kaate, Ahmed Mohamed Sayed Kamel, Soon-Gyo Jung, Jim Jansen
Int. J. Hum. Comput. Interact.5
2021 The ability of personas: An empirical evaluation of altering incorrect preconceptions about users
Joni Salminen, Soon-Gyo Jung, Shammur Absar Chowdhury, Dianne Ramirez Robillos, Jim Jansen
Int. J. Hum. Comput. Stud.5
2020 A Literature Review of Quantitative Persona Creation
abstract
Quantitative persona creation (QPC) has tremendous potential, as HCI researchers and practitioners can leverage user data from online analytics and digital media platforms to better understand their users and customers. However, there is a lack of a systematic overview of the QPC methods and progress made, with no standard methodology or known best practices. To address this gap, we review 49 QPC research articles from 2005 to 2019. Results indicate three stages of QPC research: Emergence, Diversification, and Sophistication. Sharing resources, such as datasets, code, and algorithms, is crucial to achieving the next stage (Maturity). For practitioners, we provide guiding questions for assessing QPC readiness in organizations.
Joni Salminen, Kathleen W. Guan, Soon-Gyo Jung, Shammur Absar Chowdhury, Jim Jansen
CHI5
2020 Personas and Analytics: A Comparative User Study of Efficiency and Effectiveness for a User Identification Task
abstract
Personas are a well-known technique in human computer interaction. However, there is a lack of rigorous empirical research evaluating personas relative to other methods. In this 34-participant experiment, we compare a persona system and an analytics system, both using identical user data, for efficiency and effectiveness for a user identification task. Results show that personas afford faster task completion than the analytics system, as well as outperforming analytics with significantly higher user identification accuracy. Qualitative analysis of think-aloud transcripts shows that personas have other benefits regarding learnability and consistency. However, the analytics system affords insights and capabilities that personas cannot due to inherent design differences. Findings support the use of personas to learn about users, empirically confirming some of the stated benefits in the literature, while also highlighting the limitations of personas that may necessitate the use of accompanying methods.
Joni Salminen, Soon-Gyo Jung, Shammur Absar Chowdhury, Sercan Sengün, Jim Jansen
CHI5
2020 The effect of numerical and textual information on visual engagement and perceptions of AI-driven persona interfaces
abstract
In an experiment, we present 38 marketing and data analysts professionals with two online AI-driven persona interfaces, one using numbers and the other using text. We employ eye tracking, think-aloud, and a post-engagement survey for data collection to measure perception and visual engagement with the personas along 7 constructs. Results show that the use of numbers has a mixed effect on the perceptions and visual engagement of the persona profile, with job role as a determining factor on whether numbers/text affect end users for 2 of the constructs. The use of numbers has a significant positive effect on user perceptions of usefulness by analysts but a significantly negative effect on user perceptions of completeness for both marketers and analysts. The use of numbers decreases the perceived completeness of the personas for both marketer and analysts. This research has both theoretical and practical consequences for AI-driven persona development and their interface design, suggesting that the inclusion of numbers can have a desirable effect for certain roles but with possible negative effects on user perceptions.
Joni Salminen, Ying-Hsang Liu, Sercan Sengün, João M. Santos 0001, Soon-Gyo Jung, Jim Jansen
IUI6
2020 A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection
abstract
Access to social media often enables users to engage in conversation with limited accountability. This allows a user to share their opinions and ideology, especially regarding public content, occasionally adopting offensive language. This may encourage hate crimes or cause mental harm to targeted individuals or groups. Hence, it is important to detect offensive comments in social media platforms. Typically, most studies focus on offensive commenting in one platform only, even though the problem of offensive language is observed across multiple platforms. Therefore, in this paper, we introduce and make publicly available a new dialectal Arabic news comment dataset, collected from multiple social media platforms, including Twitter, Facebook, and YouTube. We follow two-step crowd-annotator selection criteria for low-representative language annotation task in a crowdsourcing platform. Furthermore, we analyze the distinctive lexical content along with the use of emojis in offensive comments. We train and evaluate the classifiers using the annotated multi-platform dataset along with other publicly available data. Our results highlight the importance of multiple platform dataset for (a) cross-platform, (b) cross-domain, and (c) cross-dialect generalization of classifier performance.
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-Gyo Jung, Jim Jansen, Joni Salminen
LREC5
2020 Are These Comments Triggering? Predicting Triggers of Toxicity in Online Discussions
abstract
Understanding the causes or triggers of toxicity adds a new dimension to the prevention of toxic behavior in online discussions. In this research, we define toxicity triggers in online discussions as a non-toxic comment that lead to toxic replies. Then, we build a neural network-based prediction model for toxicity trigger. The prediction model incorporates text-based features and derived features from previous studies that pertain to shifts in sentiment, topic flow, and discussion context. Our findings show that triggers of toxicity contain identifiable features and that incorporating shift features with the discussion context can be detected with a ROC-AUC score of 0.87. We discuss implications for online communities and also possible further analysis of online toxicity and its root causes.
Hind A. Al-Merekhi, Haewoon Kwak, Joni Salminen, Jim Jansen
WWW4
2020 Does a Smile Matter if the Person Is Not Real?: The Effect of a Smile and Stock Photos on Persona Perceptions
abstract
We analyze the effect of using smiling/non-smiling and stock photo/non-stock photo pictures in persona profiles on four key persona perceptions, including credibility, likability, similarity, and willingness to use. For this, we collect data from an experiment with 2,400 participants using a 16-item survey instrument and multiple persona profile treatments of which half have a smiling photo/stock photo and half do not. The results from structural equation modeling, supplemented by a qualitative analysis, show that a smile enhances the perceived similarity with the persona, similar personas are more liked, and that likability increases the willingness to use a persona. In contrast, the use of stock photos decreases the perceived similarity with the persona as well as persona credibility, both of which are significant predictors to a willingness to use a persona. These professionally crafted stock-photos seem to diminish the sense of identification with the persona. The above effects are consistent across the tested ages, genders, and races of the persona picture, although the effect sizes tend to be small. The results suggest that persona creators should use smiling pictures of real people to evoke positive perceptions toward the personas. In addition to presenting quantitative evidence on the predictors of willingness to use a persona, our research has implications for the design of persona profiles, showing that the picture choice influences individuals’ persona perceptions even when the other persona information is identical.
Joni Salminen, Soon-Gyo Jung, João M. Santos 0001, Jim Jansen
Int. J. Hum. Comput. Interact.4
2020 Persona Transparency: Analyzing the Impact of Explanations on Perceptions of Data-Driven Personas
abstract
Computational techniques are becoming more common in persona development. However, users of personas may question the information in persona profiles because they are unsure of how it was created. This problem is especially vexing for data-driven personas because their creation is an opaque algorithmic process. In this research, we analyze the effect of increased transparency – i.e., explanations of how the information in data-driven personas was produced – on user perceptions. We find that higher transparency through these explanations increases the perceived completeness and clarity of the personas. Contrary to our hypothesis, the perceived credibility of the personas decreases with the increased transparency, possibly due to the technical complexity of the persona profiles disrupting the facade of the personas being real people. This finding suggests that explaining the algorithmic process of data-driven persona creation involves a “transparency trade-off”. We also find that the gender of the persona affects the perceptions, with transparency increasing perceived completeness and empathy of the female persona, but not for the male persona. Therefore, transparency may specifically assist in the acceptance of female personas. We provide practical implication for persona creators regarding transparency in persona profiles.
Joni Salminen, João M. Santos 0001, Soon-Gyo Jung, Motahhare Eslami, Jim Jansen
Int. J. Hum. Comput. Interact.5
2020 Persona Perception Scale: Development and Exploratory Validation of an Instrument for Evaluating Individuals' Perceptions of Personas
Joni Salminen, João M. Santos 0001, Haewoon Kwak, Jisun An, Soon-Gyo Jung, Jim Jansen
Int. J. Hum. Comput. Stud.6
2019 Online Hate Ratings Vary by Extremes: A Statistical Analysis
abstract
Analyzing 5,665 crowd ratings on 1,133 social media comments, we find that individuals tend to agree on the extremes of a hate rating scale more than in the middle when evaluating the hatefulness of online comments. The agreement is higher for less hateful comments and lowest on moderately hateful comments. The results have implications for researchers developing machine learning models for online hate processing, as the extreme classes are likely to require fewer annotations for reaching statistical stability. Our findings suggest that the models developed in this domain should consider the distributions of hate ratings rather than average hate scores.
Joni Salminen, Hind A. Al-Merekhi, Ahmed Mohamed Sayed Kamel, Soon-Gyo Jung, Jim Jansen
CHIIR5
2019 Design Issues in Automatically Generated Persona Profiles: A Qualitative Analysis from 38 Think-Aloud Transcripts
abstract
Increased access to data and computational techniques enable innovations in the space of automated customer analytics, for example, automatic persona generation. Automatic persona generation is the process of creating data-driven representations from user or customer statistics. Even though automatic persona generation is technically possible and provides advantages compared to manual persona creation regarding the speed and freshness of the personas, it is not clear (a) what information to include in the persona profiles and (b) how to display that information. To query into these aspects relating information design of personas, we conducted a user study with 38 participants. In the findings, we report several challenges relating to the design of automatically generated persona profiles, including usability issues, perceptual issues, and issues relating to information content. Our research has implications for the information design of data-driven personas.
Joni Salminen, Sercan Sengün, Soon-Gyo Jung, Jim Jansen
CHIIR4
2019 View, Like, Comment, Post: Analyzing User Engagement by Topic at 4 Levels across 5 Social Media Platforms for 53 News Organizations
Kholoud Khalil Aldous, Jisun An, Jim Jansen
ICWSM3
2019 Confusion and information triggered by photos in persona profiles
Joni Salminen, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Lene Nielsen, Jim Jansen
Int. J. Hum. Comput. Stud.6
2018 "Is More Better?": Impact of Multiple Photos on Perception of Persona Profiles
abstract
In this research, we investigate if and how more photos than a single headshot can heighten the level of information provided by persona profiles. We conduct eye-tracking experiments and qualitative interviews with variations in the photos: a single headshot, a headshot and images of the persona in different contexts, and a headshot with pictures of different people representing key persona attributes. The results show that more contextual photos significantly improve the information end users derive from a persona profile; however, showing images of different people creates confusion and lowers the informativeness. Moreover, we discover that choice of pictures results in various interpretations of the persona that are biased by the end users' experiences and preconceptions. The results imply that persona creators should consider the design power of photos when creating persona profiles.
Joni Salminen, Lene Nielsen, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Jim Jansen
CHI6
2018 Automatic Persona Generation (APG): A Rationale and Demonstration
abstract
We present Automatic Persona Generation (APG), a methodology and system for quantitative persona generation using large amounts of online social media data. The system is operational, beta deployed with several client organizations in multiple industry verticals and ranging from small-to-medium sized enterprises to large multi-national corporations. Using a robust web framework and stable back-end database, APG is currently processing tens of millions of user interactions with thousands of online digital products on multiple social media platforms, such as Facebook and YouTube. APG identifies both distinct and impactful user segments and then creates persona descriptions by automatically adding pertinent features, such as names, photos, and personal attributes. We present the overall methodological approach, architecture development, and main system features. APG has a potential value for organizations distributing content via online platforms and is unique in its approach to persona generation. APG can be found online at https://persona.qcri.org.
Soon-Gyo Jung, Joni Salminen, Haewoon Kwak, Jisun An, Jim Jansen
CHIIR5
2018 Fixation and Confusion: Investigating Eye-tracking Participants' Exposure to Information in Personas
abstract
To more effectively convey relevant information to end users of persona profiles, we conducted a user study consisting of 29 participants engaging with three persona layout treatments. We were interested in confusion engendered by the treatments on the participants, and conducted a within-subjects study in the actual work environment, using eye-tracking and talk-aloud data collection. We coded the verbal data into classes of informativeness and confusion and correlated it with fixations and durations on the Areas of Interests recorded by the eye-tracking device. We used various analysis techniques, including Mann-Whitney, regression, and Levenshtein distance, to investigate how confused users differed from non-confused users, what information of the personas caused confusion, and what were the predictors of confusion of end users of personas. We consolidate our various findings into a confusion ratio measure, which highlights in a succinct manner the most confusing elements of the personas. Findings show that inconsistencies among the informational elements of the persona generate the most confusion, especially with the elements of images and social media quotes. The research has implications for the design of personas and related information products, such as user profiling and customer segmentation.
Joni Salminen, Jim Jansen, Jisun An, Soon-Gyo Jung, Lene Nielsen, Haewoon Kwak
CHIIR2
2018 Assessing the Accuracy of Four Popular Face Recognition Tools for Inferring Gender, Age, and Race
Soon-Gyo Jung, Jisun An, Haewoon Kwak, Joni Salminen, Jim Jansen
ICWSM5
2018 Automatically Conceptualizing Social Media Analytics Data via Personas
Soon-Gyo Jung, Joni Salminen, Jisun An, Haewoon Kwak, Jim Jansen
ICWSM5
2018 Anatomy of Online Hate: Developing a Taxonomy and Machine Learning Models for Identifying and Classifying Hate in Online News Media
Joni Salminen, Hind A. Al-Merekhi, Milica Milenkovic, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Jim Jansen
ICWSM7
2018 What We Read, What We Search: Media Attention and Public Attention Among 193 Countries
abstract
We investigate the alignment of international attention of news media organizations within 193 countries with the expressed international interests of the public within those same countries from March 7, 2016 to April 14, 2017. We collect fourteen months of longitudinal data of online news from Unfiltered News and web search volume data from Google Trends and build a multiplex network of media attention and public attention in order to study its structural and dynamic properties. Structurally, the media attention and the public attention are both similar and different depending on the resolution of the analysis. For example, we find that 63.2% of the country-specific media and the public pay attention to different countries, but local attention flow patterns, which are measured by network motifs, are very similar. We also show that there are strong regional similarities with both media and public attention that is only disrupted by significantly major worldwide incidents (e.g., Brexit). Using Granger causality, we show that there are a substantial number of countries where media attention and public attention are dissimilar by topical interest. Our findings show that the media and public attention toward specific countries are often at odds, indicating that the public within these countries may be ignoring their country-specific news outlets and seeking other online sources to address their media needs and desires.
Haewoon Kwak, Jisun An, Joni Salminen, Soon-Gyo Jung, Jim Jansen
WWW5
2018 Questioner or question: Predicting the response rate in social question and answering on Sina Weibo
Zhe Liu 0002, Jim Jansen
Inf. Process. Manag.2
2018 Imaginary People Representing Real Numbers: Generating Personas from Online Social Media Data
abstract
We develop a methodology to automate creating imaginary people, referred to as personas, by processing complex behavioral and demographic data of social media audiences. From a popular social media account containing more than 30 million interactions by viewers from 198 countries engaging with more than 4,200 online videos produced by a global media corporation, we demonstrate that our methodology has several novel accomplishments, including: (a) identifying distinct user behavioral segments based on the user content consumption patterns; (b) identifying impactful demographics groupings; and (c) creating rich persona descriptions by automatically adding pertinent attributes, such as names, photos, and personal characteristics. We validate our approach by implementing the methodology into an actual working system; we then evaluate it via quantitative methods by examining the accuracy of predicting content preference of personas, the stability of the personas over time, and the generalizability of the method via applying to two other datasets. Research findings show the approach can develop rich personas representing the behavior and demographics of real audiences using privacy-preserving aggregated online social media data from major online platforms. Results have implications for media companies and other organizations distributing content via online platforms.
Jisun An, Haewoon Kwak, Soon-Gyo Jung, Joni Salminen, M. Admad, Jim Jansen
ACM Trans. Web6
2017 Leveraging Social Analytics Data for Identifying Customer Segments for Online News Media
abstract
In this work, we describe a methodology for leveraging large amounts of customer interaction data with online content from major social media platforms in order to isolate meaningful customer segments. The methodology is robust in that it can rapidly identify diverse customer segments using solely online behaviors and then associate these behavioral customer segments with the related distinct demographic segments, presenting a holistic picture of the customer base of an organization. We validate our methodology via the implementation of a working system that rapidly and in near real-time processes tens of millions of online customer interactions with content posted on major social media platforms in order to identify both the distinct behavioral segments and corresponding impactful demographic segments. We illustrate the functionality of the methodology with real data from a major online content provider with millions of online interactions from more than thirty countries. We further show one possible use for such information via the automatic generation of personas for an organization, which can be used for the formulation of marketing strategy, implementation of advertising plans, or development of products. The research results offer insights into competitive marketing and product preferences for the consumers of online digital content. We conclude with a discussion of areas for future work.
Jim Jansen, Soon-Gyo Jung, Joni Salminen, Jisun An, Haewoon Kwak
AICCSA1
2017 Personas for Content Creators via Decomposed Aggregate Audience Statistics
abstract
We propose a novel method for generating personas based on online user data for the increasingly common situation of content creators distributing products via online platforms. We use non-negative matrix factorization to identify user segments and develop personas by adding personality such as names and photos. Our approach can develop accurate personas representing real groups of people using online user data, versus relying on manually gathered data.
Jisun An, Haewoon Kwak, Jim Jansen
ASONAM3
2017 Identifying and predicting the desire to help in social question and answering
Zhe Liu 0002, Jim Jansen
Inf. Process. Manag.2
2017 ASK: A taxonomy of accuracy, social, and knowledge information seeking posts in social question and answering
abstract
Many people turn to their social networks to find information through the practice of question and answering. We believe it is necessary to use different answering strategies based on the type of questions to accommodate the different information needs. In this research, we propose the ASK taxonomy that categorizes questions posted on social networking sites into three types according to the nature of the questioner's inquiry of accuracy, social, or knowledge. To automatically decide which answering strategy to use, we develop a predictive model based on ASK question types using question features from the perspectives of lexical, topical, contextual, and syntactic as well as answer features. By applying the classifier on an annotated data set, we present a comprehensive analysis to compare questions in terms of their word usage, topical interests, temporal and spatial restrictions, syntactic structure, and response characteristics. Our research results show that the three types of questions exhibited different characteristics in the way they are asked. Our automatic classification algorithm achieves an 83% correct labeling result, showing the value of the ASK taxonomy for the design of social question and answering systems.
Zhe Liu 0002, Jim Jansen
J. Assoc. Inf. Sci. Technol.2
2017 Information Sharing by Viewers Via Second Screens for In-Real-Life Events
abstract
The use of second screen devices with social media facilitates conversational interaction concerning broadcast media events, creating what we refer to as the social soundtrack. In this research, we evaluate the change of the Super Bowl XLIX social soundtrack across three social media platforms on the topical categories of commercials, music, and game at three game phases ( Pre , During , and Post ). We perform statistical analysis on more than 3M, 800K, and 50K posts from Twitter, Instagram, and Tumblr, respectively. Findings show that the volume of posts in the During phase is fewer compared to Pre and Post phases; however, the hourly mean in the During phase is considerably higher than it is in the other two phases. We identify the predominant phase and category of interaction across all three social media sites. We also determine the significance of change in absolute scale across the Super Bowl categories (commercials, music, game) and in both absolute and relative scales across Super Bowl phases ( Pre , During , Post ) for the three social network platforms (Twitter, Tumblr, Instagram). Results show that significant phase-category relationships exist for all three social networks. The results identify the During phase as the predominant one for all three categories on all social media sites with respect to the absolute volume of conversations in a continuous scale. From the relative volume perspective, the During phase is highest for the music category for most social networks. For the commercials and game categories, however, the Post phase is higher than the During phase for Twitter and Instagram, respectively. Regarding category identification, the game category is the highest for Twitter and Instagram but not for Tumblr, which has dominant peaks for music and/or commercials in all three phases. It is apparent that different social media platforms offer various phase and category affordances. These results are important in identifying the influence that second screen technology has on information sharing across different social media platforms and indicates that the viewer role is transitioning from passive to more active.
Partha Mukherjee, Jim Jansen
ACM Trans. Web2
2016 Validating social media data for automatic persona generation
abstract
Using personas during interactive design has considerable potential for product and content development. Unfortunately, personas have typically been a fairly static technique. In this research, we validate an approach for creating personas in real time, based on analysis of actual social media data in an effort to automate the generation of personas. We validate that social media data can be implemented as an approach for automating generating personas in real time using actual YouTube social media data from a global media corporation that produces online digital content. Using the organization's YouTube channel, we collect demographic data, customer interactions, and topical interests, leveraging more than 188,000 subscriber profiles and more than 30 million user interactions. Then, we conduct statistical analysis on the social media data to determine whether the data could lead to the generation of valid personas based on statistically difference market segments. Findings show that customers can be segmented using product topics by gender and age based using social media data. However, our findings also show that the data is biased by the content created. The results offer insights into competitive marketing and product preferences for the consumers of the online digital content. Implications are that personas can be generated in real-time using social media data, instead of a time-consuming manual development process.
Jisun An, Haewoon Kwak, Jim Jansen
AICCSA3
2016 Some commericial concerns of Amazon's community forums: The case of the Kindle
abstract
Ecommerce consumers have many platforms for discussing products, one of which is the online product forum. These forums exist across the web on a variety of ecommerce shopping sites, including Amazon.com. However, the value of an online product forum for consumers, ecommerce, or the shopping site business is unclear. In this research, we examine 1,027 threads and 12,467 posts from the Amazon Kindle Forum, which is Amazon's most active product forum. We use text analytics techniques to categorize the social media posts according to phases of buying decision process, simplified into pre-purchase and post-purchase activity. We use the metric of post helpfulness rating to measure the post value to the customer. Our analysis shows that customers primarily use the Kindle Forum for post-purchase activity. In addition, customers appear to find embedded links particularly helpful in the forum. Activity in the forum is related to the release dates of different Kindle models. Finally, results show that the forum can be used to identify customer complaints and product issues, so online discussion forums are a valuable tool for customer relationship management. The implications of the findings are that ecommerce companies could benefit from providing greater post-purchase customer support in their product forums.
Allie Whitman, Jim Jansen
AICCSA2
2016 Detecting Rumors from Microblogs with Recurrent Neural Networks
Jing Ma 0004, Wei Gao 0001, Prasenjit Mitra 0001, Sejeong Kwon, Jim Jansen, Kam-Fai Wong, Meeyoung Cha
IJCAI5
2016 A web analytics approach for appraising electronic resources in academic libraries
abstract
University libraries provide access to thousands of journals and spend millions of dollars annually on electronic resources. With several commercial entities providing these electronic resources, the result can be silo systems and processes to evaluate cost and usage of these resources, making it difficult to provide meaningful analytics. In this research, we examine a subset of journals from a large research library using a web analytics approach with the goal of developing a framework for the analysis of library subscriptions. This foundational approach is implemented by comparing the impact to the cost, titles, and usage for the subset of journals and by assessing the funding area. Overall, the results highlight the benefit of a web analytics evaluation framework for university libraries and the impact of classifying titles based on the funding area. Furthermore, they show the statistical difference in both use and cost among the various funding areas when ranked by cost, eliminating the outliers of heavily used and highly expensive journals. Future work includes refining this model for a larger scale analysis tying metrics to library organizational objectives and for the creation of an online application to automate this analysis.
Daniel M. Coughlin, Mark C. Campbell, Jim Jansen
J. Assoc. Inf. Sci. Technol.3
2016 Modeling journal bibliometrics to predict downloads and inform purchase decisions at university research libraries
abstract
University libraries provide access to thousands of online journals and other content, spending millions of dollars annually on these electronic resources. Providing access to these online resources is costly, and it is difficult both to analyze the value of this content to the institution and to discern those journals that comparatively provide more value. In this research, we examine 1,510 journals from a large research university library, representing more than 40% of the university's annual subscription cost for electronic resources at the time of the study. We utilize a web analytics approach for the creation of a linear regression model to predict usage among these journals. We categorize metrics into two classes: global (journal focused) and local (institution dependent). Using 275 journals for our training set, our analysis shows that a combination of global and local metrics creates the strongest model for predicting full‐text downloads. Our linear regression model has an accuracy of more than 80% in predicting downloads for the 1,235 journals in our test set. The implications of the findings are that university libraries that use local metrics have better insight into the value of a journal and therefore more efficient cost content management.
Daniel M. Coughlin, Jim Jansen
J. Assoc. Inf. Sci. Technol.2
2016 Understanding and Predicting Question Subjectivity in Social Question and Answering
abstract
The explosive popularity of social networking sites has provided an additional venue for online information seeking. By posting questions in their status updates, more and more people are turning to social networks to fulfill their information needs. Given that understanding individuals' information needs could improve the performance of question answering, in this paper, we model the task of intent detection as a binary classification problem, and thus for each question, two classes are defined: subjective and objective. We use a comprehensive set of lexical, syntactical, and contextual features to build the classifier and the experimental results show satisfactory classification performance. By applying the classifier on a larger dataset, we then present in-depth analyses to compare subjective and objective questions, in terms of the way they are being asked and answered. We find that the two types of questions exhibited very different characteristics, and further validate the expected benefits of differentiating questions according to their subjectivity orientations.
Zhe Liu 0002, Jim Jansen
IEEE Trans. Comput. Soc. Syst.2
2015 External to internal search: Associating searching on search engines with searching on sites
Adan Ortiz-Cordova, Jim Jansen
Inf. Process. Manag.3
2013 Factors influencing the response rate in social question and answering behavior
abstract
With the increasing growth and popularity of social networking sites, social question and answering has become a venue for individuals to seek and share information. This study evaluates eleven extrinsic factors that may influence the response rate in social question and answering. These factors include the number of followers, the frequency of posting, the number of at-mentioned recipients, whether or not a question contains an at-mentioned verified account, unverified account, hashtag, emoticon, expression of gratitude, repeated punctuation or interjections, as well as the topic and the posting time of a question. We collected and analyzed over 10,000 questions from Sina Weibo. Eight out of all eleven features were found to significantly predict the number of responses received. We believe that our study is of significant value in providing insights for the design and development of future social question and answering tools, as well as enhancing the collaboration among social network users in supporting social information seeking activities.
Zhe Liu 0002, Jim Jansen
CSCW2
2013 Evaluating the performance of demographic targeting using gender in sponsored search
Jim Jansen, Kathleen A. Moore, Stephen Carman
Inf. Process. Manag.1
2013 The effect of ad rank on the performance of keyword advertising campaigns
abstract
The goal of this research is to evaluate the effect of ad rank on the performance of keyword advertising campaigns. We examined a large‐scale data file comprised of nearly 7,000,000 records spanning 33 consecutive months of a majorUSretailer's search engine marketing campaign. The theoretical foundation is serial position effect to explain searcher behavior when interacting with ranked ad listings. We control for temporal effects and use one‐way analysis of variance (ANOVA) withTamhane'sT2 tests to examine the effect of ad rank on critical keyword advertising metrics, including clicks, cost‐per‐click, sales revenue, orders, items sold, and advertising return on investment. Our findings show significant ad rank effect on most of those metrics, although less effect on conversion rates. A primacy effect was found on both clicks and sales, indicating a general compelling performance of top‐ranked ads listed on the first results page. Conversion rates, on the other hand, follow a relatively stable distribution except for the top 2 ads, which had significantly higher conversion rates. However, examining conversion potential (the effect of both clicks and conversion rate), we show that ad rank has a significant effect on the performance of keyword advertising campaigns. Conversion potential is a more accurate measure of the impact of an ad's position. In fact, the first ad position generates about 80% of the total profits, after controlling for advertising costs. In addition to providing theoretical grounding, the research results reported in this paper are beneficial to companies using search engine marketing as they strive to design more effective advertising campaigns.
Jim Jansen, Zhe Liu 0002, Zach Simon
J. Assoc. Inf. Sci. Technol.1
2012 An Integrated Conceptual Model to Incorporate Information Tasks in Workflow Models
Sandeep Purao, Wolfgang Maass 0002, Veda C. Storey, Jim Jansen, Madhu C. Reddy
ER4
2012 Classifying web search queries to identify high revenue generating customers
abstract
Traffic from search engines is important for most online businesses, with the majority of visitors to many websites being referred by search engines. Therefore, an understanding of this search engine traffic is critical to the success of these websites. Understanding search engine traffic means understanding the underlying intent of the query terms and the corresponding user behaviors of searchers submitting keywords. In this research, using 712,643 query keywords from a popular Spanish music website relying on contextual advertising as its business model, we use a k‐means clustering algorithm to categorize the referral keywords with similar characteristics of onsite customer behavior, including attributes such as clickthrough rate and revenue. We identified 6 clusters of consumer keywords. Clusters range from a large number of users who are low impact to a small number of high impact users. We demonstrate how online businesses can leverage this segmentation clustering approach to provide a more tailored consumer experience. Implications are that businesses can effectively segment customers to develop better business models to increase advertising conversion rates.
Adan Ortiz-Cordova, Jim Jansen
J. Assoc. Inf. Sci. Technol.2
2011 Real time search on the web: Queries, topics, and economic value
Jim Jansen, Zhe Liu 0002, Courtney Weaver, Gerry Campbell, Matthew Gregg
Inf. Process. Manag.1
2010 Gender demographic targeting in sponsored search
abstract
In this research, we evaluate the effect of gender in analyzing the performance of sponsored search advertising. We examine a log file with data comprised of nearly 7,000,000 records spanning 33 consecutive months of a search engine marketing campaign from a major US retailer. We classify key phrases selected for the campaign with a probability of being targeted for a specific gender and then compare the consumer actions using the critical sponsored search metrics of impressions, clicks, cost-per-click, sales revenue, orders, and items sold. Findings from our analysis show that the gender-orientation of the key phrase is a significant determinant in predicting behaviors and performance, with statistically different consumer behaviors for all attributes as the probability of a male or female keyword phrase changes. However, gender neutral phrases perform the best overall, calling into question the benefits of demographic targeting. Insight from this research could result in sponsored results being more effectively targeted to searchers and potential consumers.
Jim Jansen, Lauren Solomon
CHI1
2010 The seventeen theoretical constructs of information searching and information retrieval
abstract
Abstract In this article, we identify, compare, and contrast theoretical constructs for the fields of information searching and information retrieval to emphasize the uniqueness of and synergy between the fields. Theoretical constructs are the foundational elements that underpin a field's core theories, models, assumptions, methodologies, and evaluation metrics. We provide a framework to compare and contrast the theoretical constructs in the fields of information searching and information retrieval usingintellectual perspectiveandtheoretical orientation. The intellectual perspectives areinformation searching,information retrieval, andcross‐cutting; and the theoretical orientations areinformation,people, andtechnology. Using this framework, we identify 17 significant constructs in these fields contrasting the differences and comparing the similarities. We discuss the impact of the interplay among these constructs for moving research forward within both fields. Although there is tension between the fields due to contradictory constructs, an examination shows a trend toward convergence. We discuss the implications for future research within the information searching and information retrieval fields.
Jim Jansen, Soo Young Rieh
J. Assoc. Inf. Sci. Technol.1
2009 Using the taxonomy of cognitive learning to model online searching
Jim Jansen, Danielle L. Booth, Brian Keith Smith
Inf. Process. Manag.1
2009 Time series analysis of a Web search engine transaction log
Jim Jansen, Amanda Spink
Inf. Process. Manag.2
2009 Web search
Amanda Spink, Jim Jansen
Inf. Sci.2
2009 Patterns of query reformulation during Web searching
abstract
Abstract Query reformulation is a key user behavior during Web search. Our research goal is to develop predictive models of query reformulation during Web searching. This article reports results from a study in which we automatically classified the query‐reformulation patterns for 964,780 Web searching sessions, composed of 1,523,072 queries, to predict the next query reformulation. We employed an n‐gram modeling approach to describe the probability of users transitioning from one query‐reformulation state to another to predict their next state. We developed first‐, second‐, third‐, and fourth‐order models and evaluated each model for accuracy of prediction, coverage of the dataset, and complexity of the possible pattern set. The results show that Reformulation and Assistance account for approximately 45% of all query reformulations; furthermore, the results demonstrate that the first‐ and second‐order models provide the best predictability, between 28 and 40% overall and higher than 70% for some patterns. Implications are that the n‐gram approach can be used for improving searching systems and searching assistance.
Jim Jansen, Danielle L. Booth, Amanda Spink
J. Assoc. Inf. Sci. Technol.1
2009 Brand and its effect on user perception of search engine performance
abstract
Abstract In this research we investigate the effect of search engine brand on the evaluation of searching performance. Our research is motivated by the large amount of search traffic directed to a handful of Web search engines, even though many have similar interfaces and performance. We conducted a laboratory experiment with 32 participants using a 42 factorial design confounded in four blocks to measure the effect of four search engine brands (Google, MSN, Yahoo!, and a locally developed search engine) while controlling for the quality and presentation of search engine results. We found brand indeed played a role in the searching process. Brand effect varied in different domains. Users seemed to place a high degree of trust in major search engine brands; however, they were more engaged in the searching process when using lesser‐known search engines. It appears that branding affects overall Web search at four stages: (a) search engine selection, (b) search engine results page evaluation, (c) individual link evaluation, and (d) evaluation of the landing page. We discuss the implications for search engine marketing and the design of empirical studies measuring search engine performance.
Jim Jansen, Mimi Zhang, Carsten D. Schultz
J. Assoc. Inf. Sci. Technol.1
2009 Twitter power: Tweets as electronic word of mouth
abstract
Abstract In this paper we report research results investigating microblogging as a form of electronic word‐of‐mouth for sharing consumer opinions concerning brands. We analyzed more than 150,000 microblog postings containing branding comments, sentiments, and opinions. We investigated the overall structure of these microblog postings, the types of expressions, and the movement in positive or negative sentiment. We compared automated methods of classifying sentiment in these microblogs with manual coding. Using a case study approach, we analyzed the range, frequency, timing, and content of tweets in a corporate account. Our research findings show that 19% of microblogs contain mention of a brand. Of the branding microblogs, nearly 20% contained some expression of brand sentiments. Of these, more than 50% were positive and 33% were critical of the company or product. Our comparison of automated and manual coding showed no significant differences between the two approaches. In analyzing microblogs for structure and composition, the linguistic structure of tweets approximate the linguistic patterns of natural language expressions. We find that microblogging is an online tool for customer word of mouth communications and discuss the implications for corporations using microblogging as part of their overall marketing strategy.
Jim Jansen, Mimi Zhang, Kate Sobel, Abdur Chowdhury
J. Assoc. Inf. Sci. Technol.1
2009 A study and comparison of multimedia Web searching: 1997-2006
abstract
Abstract Searching for multimedia is an important activity for users of Web search engines. Studying user's interactions with Web search engine multimedia buttons, including image, audio, and video, is important for the development of multimedia Web search systems. This article provides results from a Weblog analysis study of multimedia Web searching by Dogpile users in 2006. The study analyzes the (a) duration, size, and structure of Web search queries and sessions; (b) user demographics; (c) most popular multimedia Web searching terms; and (d) use of advanced Web search techniques including Boolean and natural language. The current study findings are compared with results from previous multimedia Web searching studies. The key findings are: (a) Since 1997, image search consistently is the dominant media type searched followed by audio and video; (b) multimedia search duration is still short (>50% of searching episodes are <1 min), using few search terms; (c) many multimedia searches are for information about people, especially in audio search; and (d) multimedia search has begun to shift from entertainment to other categories such as medical, sports, and technology (based on the most repeated terms). Implications for design of Web multimedia search engines are discussed.
Dian Tjondronegoro, Amanda Spink, Jim Jansen
J. Assoc. Inf. Sci. Technol.3
2009 Identification of factors predicting clickthrough in Web searching using neural network analysis
abstract
Abstract In this research, we aim to identify factors that significantly affect the clickthrough of Web searchers. Our underlying goal is determine more efficient methods to optimize the clickthrough rate. We devise a clickthrough metric for measuring customer satisfaction of search engine results using the number of links visited, number of queries a user submits, and rank of clicked links. We use a neural network to detect the significant influence of searching characteristics on future user clickthrough. Our results show that high occurrences of query reformulation, lengthy searching duration, longer query length, and the higher ranking of prior clicked links correlate positively with future clickthrough. We provide recommendations for leveraging these findings for improving the performance of search engine retrieval and result ranking, along with implications for search engine marketing.
Jim Jansen, Amanda Spink
J. Assoc. Inf. Sci. Technol.2
2008 Myatt, G.J. (2007). Making sense of data: a practical guide to exploratory data analysis and data mining (pp. 280). Wiley
Jim Jansen
Inf. Process. Manag.1
2008 Determining the informational, navigational, and transactional intent of Web queries
Jim Jansen, Danielle L. Booth, Amanda Spink
Inf. Process. Manag.1
2008 A model for understanding collaborative information behavior in context: A study of two healthcare teams
Madhu C. Reddy, Jim Jansen
Inf. Process. Manag.2
2007 Investigating the relevance of sponsored results for web ecommerce queries
abstract
Are sponsored links, the primary business model for Web search engines, providing Web consumers with relevant results? This research addresses this issue by investigating the relevance of sponsored and non-sponsored links for ecommerce queries from the major search engines. The results show that average relevance ratings for sponsored and non-sponsored links are virtually the same, although the relevance ratings for sponsored links are statistically higher. We used 108 ecommerce queries and 8,256 retrieved links for these queries from three major Web search engines, Google, MSN, and Yahoo!. We present the implications for Web search engines and sponsored search as a long-term business model as well as a mechanism for finding relevant information for searchers.
Jim Jansen
SIGIR1
2007 Viewing online searching within a learning paradigm
abstract
In this research, we investigate whether one can model online searching as a learning paradigm. We examined the searching characteristics of 41 participants engaged in 246 searching tasks. We classified the searching tasks according to Anderson and Krathwohl's Taxonomy, an updated version of Bloom's taxonomy. Anderson and Krathwohl is a six level categorization of cognitive learning. Research results show that Applying takes the most searching effort as measured by queries per session and specific topics searched per sessions. The categories of Remembering and Understanding, which are lower-order learning levels, exhibit searching characteristics similar to the higher order categories of Evaluating and Creating. It seems that searchers rely primarily on their internal knowledge and use searching primarily as fact checking and verification when engaged in Evaluating and Creating. Implications are that the commonly held notions of Web searchers having simple information goals may not be correct. We discuss the implications for Web searching, including designing interfaces to support exploration and alternate views.
Jim Jansen, Brian Keith Smith, Danielle L. Booth
SIGIR1
2007 Determining the user intent of web search engine queries
abstract
Determining the user intent of Web searches is a difficult problem due to the sparse data available concerning the searcher. In this paper, we examine a method to determine the user intent underlying Web search engine queries. We qualitatively analyze samples of queries from seven transaction logs from three different Web search engines containing more than five million queries. From this analysis, we identified characteristics of user queries based on three broad classifications of user intent. The classifications of informational, navigational, and transactional represent the type of content destination the searcher desired as expressed by their query. We implemented our classification algorithm and automatically classified a separate Web search engine transaction log of over a million queries submitted by several hundred thousand users. Our findings show that more than 80% of Web queries are informational in nature, with about 10% each being navigational and transactional. In order to validate the accuracy of our algorithm, we manually coded 400 queries and compared the classification to the results from our algorithm. This comparison showed that our automatic classification has an accuracy of 74%. Of the remaining 25% of the queries, the user intent is generally vague or multi-faceted, pointing to the need to for probabilistic classification. We illustrate how knowledge of searcher intent might be used to enhance future Web search engines.
Jim Jansen, Danielle L. Booth, Amanda Spink
WWW1
2007 Understanding web search via a learning paradigm
abstract
Investigating whether one can view Web searching as a learning process, we examined the searching characteristics of 41 participants engaged in 246 searching tasks. We classified the searching tasks according an updated version of Bloom.s taxonomy, a six level categorization of cognitive learning. Results show that Applying takes the most searching effort as measured by queries per session and specific topics searched per sessions. The lower level categories of Remembering and Understanding exhibit searching characteristics similar to the higher order learning of Evaluating and Creating. It appears that searchers rely primarily on their internal knowledge for Evaluating and Creating, using searching primarily as fact checking and verification. Implications are that the commonly held notion that Web searchers have simple information needs may not be correct. We discuss the implications for Web searching, including designing interfaces to support exploration.
Jim Jansen, Brian Keith Smith, Danielle L. Booth
WWW1
2007 Brand awareness and the evaluation of search results
abstract
We investigate the effect of search engine brand (i.e., the identifying name or logo that distinguishes a product from its competitors) on evaluation of system performance. This research is motivated by the large amount of search traffic directed to a handful of Web search engines, even though most are of equal technical quality with similar interfaces. We conducted a laboratory study with 32 participants to measure the effect of four search engine brands while controlling for the quality of search engine results. There was a 25% difference between the most highly rated search engine and the lowest using average relevance ratings, even though search engine results were identical in both content and presentation. Qualitative analysis suggests branding affects user views of popularity, trust and specialization. We discuss implications for search engine marketing and the design of search engine quality studies.
Jim Jansen, Mimi Zhang
WWW1
2007 Factors relating to the decision to click on a sponsored link
Jim Jansen, Anna Brown, Marc Resnick
Decis. Support Syst.1
2007 The Craft of Research, 2nd edition (Chicago Guides to Writing, Editing, and Publishing) (Paperback) by Wayne C. Booth, Joseph M. Williams, Gregory G. Colomb. Paperback: 336 pages. University of Chicago Press; 2nd edition. ISBN: 0226065685
Jim Jansen
Inf. Process. Manag.1
2007 Effective Expert Witnessing, fourth ed., Matson, Jack V., Daou, Suha F., Soper, Jeffrey G. CRC, 160 p., ISBN: 0849313015
Jim Jansen
Inf. Process. Manag.1
2007 Chris Anderson, The Long Tail: Why the Future of Business is Selling Less or More, Hyperion, New York (2006) ISBN 1-4013-0237-8 $24.95
Jim Jansen
Inf. Process. Manag.1
2007 Defining a session on Web search engines
abstract
Abstract Detecting query reformulations within a session by a Web searcher is an important area of research for designing more helpful searching systems and targeting content to particular users. Methods explored by other researchers include both qualitative (i.e., the use of human judges to manually analyze query patterns on usually small samples) and nondeterministic algorithms, typically using large amounts of training data to predict query modification during sessions. In this article, we explore three alternative methods for detection of session boundaries. All three methods are computationally straightforward and therefore easily implemented for detection of session changes. We examine 2,465,145 interactions from 534,507 users of Dogpile.com on May 6, 2005. We compare session analysis using (a) Internet Protocol address and cookie; (b) Internet Protocol address, cookie, and a temporal limit on intrasession interactions; and (c) Internet Protocol address, cookie, and query reformulation patterns. Overall, our analysis shows that defining sessions by query reformulation along with Internet Protocol address and cookie provides the best measure, resulting in an 82% increase in the count of sessions. Regardless of the method used, the mean session length was fewer than three queries, and the mean session duration was less than 30 min. Searchers most often modified their query by changing query terms (nearly 23% of all query modifications) rather than adding or deleting terms. Implications are that for measuring searching traffic, unique sessions may be a better indicator than the common metric of unique visitors. This research also sheds light on the more complex aspects of Web searching involving query modifications and may lead to advances in searching tools.
Jim Jansen, Amanda Spink, Chris Blakely, Sherry Koshman
J. Assoc. Inf. Sci. Technol.1
2007 Web searcher interaction with the Dogpile.com metasearch engine
abstract
Abstract Metasearch engines are an intuitive method for improving the performance of Web search by increasing coverage, returning large numbers of results with a focus on relevance, and presenting alternative views of information needs. However, the use of metasearch engines in an operational environment is not well understood. In this study, we investigate the usage of Dogpile.com, a major Web metasearch engine, with the aim of discovering how Web searchers interact with metasearch engines. We report results examining 2,465,145 interactions from 534,507 users of Dogpile.com on May 6, 2005 and compare these results with findings from other Web searching studies. We collect data on geographical location of searchers, use of system feedback, content selection, sessions, queries, and term usage. Findings show that Dogpile.com searchers are mainly from the USA (84% of searchers), use about 3 terms per query (mean 5 2.85), implement system feedback moderately (8.4% of users), and generally (56% of users) spend less than one minute interacting with the Web search engine. Overall, metasearchers seem to have higher degrees of interaction than searchers on non‐metasearch engines, but their sessions are for a shorter period of time. These aspects of metasearching may be what define the differences from other forms of Web searching. We discuss the implications of our findings in relation to metasearch for Web searchers, search engines, and content providers.
Jim Jansen, Amanda Spink, Sherry Koshman
J. Assoc. Inf. Sci. Technol.1
2007 Editorial
Kirstie Hawkey, Jim Jansen, Melanie Kellar, Don Turnbull
J. Web Eng.3
2007 The comparative effectiveness of sponsored and nonsponsored links for Web e-commerce queries
abstract
The predominant business model for Web search engines is sponsored search, which generates billions in yearly revenue. But are sponsored links providing online consumers with relevant choices for products and services? We address this and related issues by investigating the relevance of sponsored and nonsponsored links for e-commerce queries on the major search engines. The results show that average relevance ratings for sponsored and nonsponsored links are practically the same, although the relevance ratings for sponsored links are statistically higher. We used 108 ecommerce queries and 8,256 retrieved links for these queries from three major Web search engines: Yahoo!, Google, and MSN. In addition to relevance measures, we qualitatively analyzed the e-commerce queries, deriving five categorizations of underlying information needs. Product-specific queries are the most prevalent (48%). Title (62%) and summary (33%) are the primary basis for evaluating sponsored links with URL a distant third (2%). To gauge the effectiveness of sponsored search campaigns, we analyzed the sponsored links from various viewpoints. It appears that links from organizations with large sponsored search campaigns are more relevant than the average sponsored link. We discuss the implications for Web search engines and sponsored search as a long-term business model and as a mechanism for finding relevant information for searchers.
Jim Jansen
ACM Trans. Web1
2006 Soumen Chakrabarti, Mining the Web: Discovering Knowledge from Hypertext Data, 2002, Morgan-Kaufmann Publishers, 352 pp., ISBN: 1-55860-754-4
Jim Jansen
Inf. Process. Manag.1
2006 Karen E. Fisher, Sanda Erdelez, Lynn (E.F.) McKechnie, Theories of Information Science Behavior, ASIST Monograph Series, Information Today
Jim Jansen
Inf. Process. Manag.1
2006 Book review
Jim Jansen
Inf. Process. Manag.1
2006 The effectiveness of Web search engines for retrieving relevant ecommerce links
Jim Jansen, Paulo R. Molina
Inf. Process. Manag.1
2006 How are we searching the World Wide Web? A comparison of nine search engine transaction logs
Jim Jansen, Amanda Spink
Inf. Process. Manag.1
2006 A study of results overlap and uniqueness among major Web search engines
Amanda Spink, Jim Jansen, Chris Blakely, Sherry Koshman
Inf. Process. Manag.2
2006 Multitasking during Web search sessions
Amanda Spink, Minsoo Park, Jim Jansen, Jan O. Pedersen 0001
Inf. Process. Manag.3
2006 An examination of searcher's perceptions of nonsponsored and sponsored links during ecommerce Web searching
abstract
Abstract In this article, we report results of an investigation into the effect of sponsored links on ecommerce information seeking on the Web. In this research, 56 participants each engaged in six ecommerce Web searching tasks. We extracted these tasks from the transaction log of a Web search engine, so they represent actual ecommerce searching information needs. Using 60 organic and 30 sponsored Web links, the quality of the Web search engine results was controlled by switching nonsponsored and sponsored links on half of the tasks for each participant. This allowed for investigating the bias toward sponsored links while controlling for quality of content. The study also investigated the relationship between searching self‐efficacy, searching experience, types of ecommerce information needs, and the order of links on the viewing of sponsored links. Data included 2,453 interactions with links from result pages and 961 utterances evaluating these links. The results of the study indicate that there is a strong preference for nonsponsored links, with searchers viewing these results first more than 82% of the time. Searching self‐efficacy and experience does not increase the likelihood of viewing sponsored links, and the order of the result listing does not appear to affect searcher evaluation of sponsored links. The implications for sponsored links as a long‐term business model are discussed.
Jim Jansen, Marc Resnick
J. Assoc. Inf. Sci. Technol.1
2006 Web searching on the Vivisimo search engine
abstract
Abstract The application of clustering to Web search engine technology is a novel approach that offers structure to the information deluge often faced by Web searchers. Clustering methods have been well studied in research labs; however, real user searching with clustering systems in operational Web environments is not well understood. This article reports on results from a transaction log analysis of Vivisimo.com, which is a Web meta‐search engine that dynamically clusters users' search results. A transaction log analysis was conducted on 2‐week's worth of data collected from March 28 to April 4 and April 25 to May 2, 2004, representing 100% of site traffic during these periods and 2,029,734 queries overall. The results show that the highest percentage of queries contained two terms. The highest percentage of search sessions contained one query and was less than 1 minute in duration. Almost half of user interactions with clusters consisted of displaying a cluster's result set, and a small percentage of interactions showed cluster tree expansion. Findings show that 11.1% of search sessions were multitasking searches, and there are a broad variety of search topics in multitasking search sessions. Other searching interactions and statistics on repeat users of the search engine are reported. These results provide insights into search characteristics with a cluster‐based Web search engine and extend research into Web searching trends.
Sherry Koshman, Amanda Spink, Jim Jansen
J. Assoc. Inf. Sci. Technol.3
2006 Automated gathering of Web information: An in-depth examination of agents interacting with search engines
abstract
The Web has become a worldwide repository of information which individuals, companies, and organizations utilize to solve or address various information problems. Many of these Web users utilize automated agents to gather this information for them. Some assume that this approach represents a more sophisticated method of searching. However, there is little research investigating how Web agents search for online information. In this research, we first provide a classification for information agent using stages of information gathering, gathering approaches, and agent architecture. We then examine an implementation of one of the resulting classifications in detail, investigating how agents search for information on Web search engines, including the session, query, term, duration and frequency of interactions. For this temporal study, we analyzed three data sets of queries and page views from agents interacting with the Excite and AltaVista search engines from 1997 to 2002, examining approximately 900,000 queries submitted by over 3,000 agents. Findings include: (1) agent sessions are extremely interactive, with sometimes hundreds of interactions per second (2) agent queries are comparable to human searchers, with little use of query operators, (3) Web agents are searching for a relatively limited variety of information, wherein only 18% of the terms used are unique, and (4) the duration of agent-Web search engine interaction typically spans several hours. We discuss the implications for Web information agents and search engines.
Jim Jansen, Tracy Mullen, Amanda Spink, Jan O. Pedersen 0001
ACM Trans. Internet Techn.1
2005 Automated evaluation of search engine performance via implicit user feedback
abstract
Measuring the information retrieval effectiveness of Web search engines can be expensive if human relevance judgments are required to evaluate search results. Using implicit user feedback for search engine evaluation provides a cost and time effective manner of addressing this problem. Web search engines can use human evaluation of search results without the expense of human evaluators. An additional advantage of this approach is the availability of real time data regarding system performance. Wecapture user relevance judgments actions such as print, save and bookmark, sending these actions and the corresponding document identifiers to a central server via a client application. We use this implicit feedback to calculate performance metrics, such as precision. We can calculate an overall system performance metric based on a collection of weighted metrics.
Jim Jansen
SIGIR2
2005 Seeking and implementing automated assistance during the search process
Jim Jansen
Inf. Process. Manag.1
2005 An analysis of Web searching by European AlltheWeb.com users
Jim Jansen, Amanda Spink
Inf. Process. Manag.1
2005 Evaluating the effectiveness of and patterns of interactions with automated searching assistance
abstract
Abstract We report quantitative and qualitative results of an empirical evaluation to determine whether automated assistance improves searching performance and when searchers desire system intervention in the search process. Forty participants interacted with two fully functional information retrieval systems in a counterbalanced, within‐participant study. The systems were identical in all respects except that one offered automated assistance and the other did not. The study used a client‐side automated assistance application, an approximately 500,000‐document Text REtrieval Conference content collection, and six topics. Results indicate that automated assistance can improve searching performance. However, the improvement is less dramatic than one might expect, with an approximately 20% performance increase, as measured by the number of user‐selected relevant documents. Concerning patterns of interaction, we identified 1,879 occurrences of searcher– system interactions and classified them into 9 major categories and 27 subcategories or states. Results indicate that there are predictable patterns of times when searchers desire and implement searching assistance. The most common three‐state pattern isExecute Query–View Results: With Scrolling–View Assistance.Searchers appear receptive to automated assistance; there is a 71% implementation rate. There does not seem to be a correlation between the use of assistance and previous searching performance. We discuss the implications for the design of information retrieval systems and future research directions.
Jim Jansen, Michael D. McNeese
J. Assoc. Inf. Sci. Technol.1
2005 A temporal comparison of AltaVista Web searching
abstract
Abstract Major Web search engines, such as AltaVista, are essential tools in the quest to locate online information. This article reports research that used transaction log analysis to examine the characteristics and changes in AltaVista Web searching that occurred from 1998 to 2002. The research questions we examined are (1) What are the changes in AltaVista Web searching from 1998 to 2002? (2) What are the current characteristics of AltaVista searching, including the duration and frequency of search sessions? (3) What changes in the information needs of AltaVista users occurred between 1998 and 2002? The results of our research show (1) a move toward more interactivity with increases in session and query length, (2) with 70% of session durations at 5 minutes or less, the frequency of interaction is increasing, but it is happening very quickly, and (3) a broadening range of Web searchers' information needs, with the most frequent terms accounting for less than 1% of total term usage. We discuss the implications of these findings for the development of Web search engines.
Jim Jansen, Amanda Spink, Jan O. Pedersen 0001
J. Assoc. Inf. Sci. Technol.1
2004 The Effect of Specialized Multimedia Collections on Web Searching
Jim Jansen, Amanda Spink, Jan O. Pedersen 0001
J. Web Eng.1
2003 Designing automated help using searcher system dialogues
abstract
This research utilizes a cognitive model of interactive information retrieval seeking to improve the performance of Web information retrieval (IR) systems. Building on the stratified model, we define interactions at the surface stratum that shed light on the cognitive, affective, and situational strata during the information retrieval process. We propose that one can utilize these interactions to improve the design information retrieval systems. This paper presents the development technique used to modify an existing IR system that monitors these interactions and, using associated assumptions about situational relevance, recommends search tactics to the user. The result is an increase in the performance of an IR system as measured by precision. The system design and results of an evaluation are presented. Research thus far indicates that user-system interactions at the surface stratum can be used to improve system performance. Using these interactions, one can develop the stratified model to a level of granularity useful for the design of Web IR systems.
Jim Jansen
SMC1
2003 Web searching agents, what are they doing out there?
abstract
The Web has become a worldwide repository of information, which individuals, companies, and organizations utilize to solve or address various information problems. Many of these Web users utilize automated agents to gather this information for them. It is assumed that this approach represents a more sophisticated method of searching. However, there is little research investigating how Web agents search for online information. In this research, we examine how agents search for information on Web search engines, including the session, query, term, duration and frequency of interactions. For this study, we analyzed queries that 2,717 agents submitted to the Alta Vista search engine on 8 September 2002. Findings include: (1) agents interacting with Web search engines use queries comparable to human searchers, (2) Web agents are searching for a relatively limited variety of information, with only 18% of the terms used being unique, and (3) agent-Web search engine interaction typically spans several hours with multiple instances of interaction per second.
Jim Jansen, Amanda Spink, Jan O. Pedersen 0001
SMC1
2003 Coverage, relevance, and ranking: The impact of query operators on Web search engine results
abstract
Research has reported that about 10% of Web searchers utilize advanced query operators, with the other 90% using extremely simple queries. It is often assumed that the use of query operators, such as Boolean operators and phrase searching, improves the effectiveness of Web searching. We test this assumption by examining the effects of query operators on the performance of three major Web search engines. We selected one hundred queries from the transaction log of a Web search service. Each of these original queries contained query operators such as AND, OR, MUST APPEAR (+), or PHRASE (" "). We then removed the operators from these one hundred advanced queries. We submitted both the original and modified queries to three major Web search engines; a total of 600 queries were submitted and 5,748 documents evaluated. We compared the results from the original queries with the operators to the results from the modified queries without the operators. We examined the results for changes in coverage, relative precision, and ranking of relevant documents. The use of most query operators had no significant effect on coverage, relative precision, or ranking, although the effect varied depending on the search engine. We discuss implications for the effectiveness of searching techniques as currently taught, for future information retrieval system design, and for future research.
Caroline M. Eastman, Jim Jansen
ACM Trans. Inf. Syst.2
2001 A review of Web searching studies and a framework for future research
abstract
Research on Web searching is at an incipient stage. This aspect provides a unique opportunity to review the current state of research in the field, identify common trends, develop a methodological framework, and define terminology for future Web searching studies. In this article, the results from published studies of Web searching are reviewed to present the current state of research. The analysis of the limited Web searching studies available indicates that research methods and terminology are already diverging. A framework is proposed for future studies that will facilitate comparison of results. The advantages of such a framework are presented, and the implications for the design of Web information retrieval systems studies are discussed. Additionally, the searching characteristics of Web users are compared and contrasted with users of traditional information retrieval and online public access systems to discover if there is a need for more studies that focus predominantly or exclusively on Web searching. The comparison indicates that Web searching differs from searching in other environments.
Jim Jansen, Udo W. Pooch
J. Assoc. Inf. Sci. Technol.1
2001 Vox populi: The public searching of the web
abstract
this paper we compare and contrast results from our two previous studies of Excite queries data sets, each containing over 1 million queries submitted by over 200,000 Excite users collected in September 1997 and December 1999. We examine how public Web searching changing during that two-year time period
Dietmar Wolfram, Amanda Spink, Jim Jansen, Tefko Saracevic
J. Assoc. Inf. Sci. Technol.3
2000 Real life, real users, and real needs: a study and analysis of user queries on the web
Jim Jansen, Amanda Spink, Tefko Saracevic
Inf. Process. Manag.1
2000 Searching for multimedia: analysis of audio, video and image Web queries
Jim Jansen, Abby Goodrum, Amanda Spink
World Wide Web1
1999 A Software Agent for Performance Improvement of Existing Information Retrieval Systems
abstract
No abstract available.
Jim Jansen
IUI1
1999 Digital video in education
abstract
Digital Video is an exciting new medium with the potential to revolutionize the way organizations train their employees. However, there are questions that must be answered. How practical is video? What is the demand? What is the best use of video? In this paper, we compare the performance and quality of common digital formats, analyze 851,770 queries from an Excite database, and present the results of a study that explores the value of digital video in an educational environment.
Todd Smith, Anthony Ruocco, Jim Jansen
SIGCSE3
1998 A Proxy Server Experiment : an Indication of the Changing Nature of the Web
abstract
With the growing reliance on connectivity to the World-Wide Web (Web), many organizations have been experiencing trouble servicing their users with adequate access and response time. Increase bandwidth on more connections to the Web can relieve the access problem, but this approach may not decrease the access time. Additionally, increase bandwidth comes at greatly increased cost. Therefore, many organizations have turned to the use of proxy servers. A proxy server is a Web server that caches Internet resources for re-use by a set of client machines. The performance increases of proxy servers has been widely reported; however, we could not locate any test of proxy server performance. Given the exponential growth of the Web in just the last year, we wondered if this would have an effect on the performance of proxy servers. Therefore, we conducted a 14-day proxy server experiment. The results of our experiment showed that the proxy servers actually decreased performance, i.e. access time. We review this experiment, analyze why the proxy server failed to decrease the access time, and draw conclusions on the changing nature of the Web and its impact on proxy servers.
Richard Howard, Jim Jansen
ICCCN2