Soya Park

dblp:173/9168 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-2149-4420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 "Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-Improving
abstract
Specifying model behavior is challenging—especially when the desired behavior is unpopular relative to the model’s training data. Reversing the influence of massive training corpora is both time-consuming and costly, and such interventions are typically inaccessible to end users. While Large Language Models (LLMs) make it easier to write instructions using natural language, specifying unpopular behaviors remains a difficult task.
Soya Park, J. D. Zamfirescu-Pereira, Chinmay Kulkarni 0001
IUI1
2024 Who2chat: A Social Networking System for Academic Researchers in Virtual Social Hours Enabling Coordinating, Overcoming Barriers and Social Signaling
abstract
Virtual academic networking is socio-technically challenging, however, fruitful for researchers' success. We introduce a system called Who2chat to tackle the challenge and facilitate connections of researchers in virtual social hours. Who2chat allows academic researchers to create a research profile and express their research interests, find researchers with similar interests, overcome social barriers, and coordinate and start video chats, all within a single interface. We engaged in an iterative design process by deploying Who2chat at academic conferences. In our preliminary deployment (N=80), we found that researchers often have difficulty finding other researchers who share similar interests, and they are shy about reaching out to other researchers. Inspired by this, we implemented social-signaling features to Who2chat and ran our first deployment (N=220). Our results highlight that the interface allowed users to find relevant researchers and helped them feel confident in joining conversations. However, this led to large group conversations where discussion topics were more superficial. In response, we developed and deployed our second interface (N=81). Key improvements were managing the size of conversations, dynamically determining and allowing individuals to join a conversation based on their relevance to the ongoing discussion, and maintaining the ratio of senior and junior members, to further enhance the quality of discussions. As a result, participants were able to meet more people and engage in more meaningful conversations. Our work demonstrates an interface design for social networking in academic settings and how to lower social barriers in virtual networking.
Soya Park, Jaeyoon Song 0001, David R. Karger, Thomas W. Malone
Proc. ACM Hum. Comput. Interact.1
2024 "How fancy you are to make us use your fancy tool": Coordinating Individuals' Tool Preference over Group Boundaries
abstract
When a group makes a decision, it necessitates the understanding and amalgamation of information from different group members. This process becomes particularly intricate in cross-boundary teams, which consist of individuals from diverse organizational backgrounds, each bringing in unique informational tools and representation modalities. People share information generated from their personal tools, and the variance in representation of such information makes it challenging to form cohesive group decisions. We conducted workshop studies with 11 knowledge workers to understand current practices of tool adaptation and negotiation in such teams. The results indicate a reluctance to adopt new tools due to perceived violations of social acceptance, often leading to negative judgments of those suggesting new tools. Consequently, participants in cross-boundary teams gravitated towards their preferred tools, complicating the aggregation of inputs and impeding cohesive decision-making. To address these challenges, we developed a platform facilitating sensemaking and decision-making without necessitating compromises on tool preferences. In our mixed-method within-subject experiments, this approach enabled faster, more informed decision-making with reduced mental load and increased engagement through enhanced social interaction and acknowledgment of diverse contributions.
Qianqia (queenie) Zhang, Soya Park, Michael J. Muller, David R. Karger
Proc. ACM Hum. Comput. Interact.2
2024 "I Really Need Your Help with This Work...": A System for Navigating the Tricky Terrain of Managing Up by Leveraging One's Motivation to Get Things Done
abstract
When people need help from their supervisors or peers, they often have to manage up to get things done. However, unlike managing subordinates (managing down), managing people of equal or higher status (managing up) are not obligated to help. These requests often involve collaborative tasks between requesters and performers. Through interviews, we found that these collaborative tasks require coordination work that is not materialized in existing management tools. We also found that requesters are willing to take on this coordination work to see their requests fulfilled. To address this issue, we propose a system called TaskLight , which allows requesters to handle coordination work themselves. For example, requesters can collect useful context and information for their performers. We conducted two deployment studies and found that TaskLight leads to better outcomes because requesters are able to assist performers more effectively. Our findings demonstrate a new way to reduce the social burdens of managing up and improve collaboration.
Soya Park, Stuti Vishwabhan, Michael J. Muller, David R. Karger
ACM Trans. Comput. Hum. Interact.1
2023 Retrospector: Rapid collaborative reflection to improve collaborative practices
abstract
Online platforms for freelancing allow teams performing complex work to be assembled in a matter of minutes and dispersed nearly as quickly. With such short time frames, ad hoc and virtual teams have few opportunities to learn strategies and effective team practices to work with their colleagues. Without such practices, teams are prone to work sub-optimally and lack direction. One key challenge in virtual teams discovering effective team practices is that because the practices ought to involve situated knowledge, it takes time to coalesce, as team members learn about each other over time. This work introduces Retrospector, that ad hoc teams can use to reflect collaboratively and reinforce effective team practices. Our interface accelerates the discovery of practices in situ and then guides them in reinforcing and applying these practices to future tasks. We conducted a between-subjects experiment (N=75) to assess our design with crowdworkers from the Amazon Mechanical Turk platform. This randomized controlled experiment showed that teams using our system for approximately six minutes of collaborative reflection were able to discover effective practices more successfully and had significantly improved team performance and viability. These results indicate that deliberate support for improving team practices can improve outcomes even through very short interaction. We conclude with design implications and opportunities for future work.
Soya Park, Chinmay Kulkarni 0001
Proc. ACM Hum. Comput. Interact.1
2022 Exploring Team-Sourced Hyperlinks to Address Navigation Challenges for Low-Vision Readers of Scientific Papers
abstract
Reading academic papers is a fundamental part of higher education and research, but navigating these information-dense texts can be challenging. In particular, low-vision readers using magnification encounter additional barriers to quickly skimming and visually locating information. In this work, we explored the design of interfaces to enable readers to: 1) navigate papers more easily, and 2) input the required navigation hooks that AI cannot currently automate. To explore this design space, we ran two exploratory studies. The first focused on current practices of low-vision paper readers, the challenges they encounter, and the interfaces they desire. During this study, low-vision participants were interviewed, and tried out four new paper navigation prototypes. Results from this study grounded the design of our end-to-end system prototype Ocean, which provides an accessible front-end for low-vision readers, and enables all readers to contribute to the backend by leaving traces of their reading paths for others to leverage. Our second study used this exploratory interface in a field study with groups of low-vision and sighted readers to probe the user experience of reading and creating traces. Our findings suggest that it may be possible for readers of all abilities to organically leave traces in papers, and that these traces can be used to facilitate navigation tasks, in particular for low-vision readers. Based on our findings, we present design considerations for creating future paper-reading tools that improve access, and organically source the required data from readers.
Soya Park, Jonathan Bragg, Kevin Larson, Danielle Bragg
Proc. ACM Hum. Comput. Interact.1
2022 Documentation Matters: Human-Centered AI System to Assist Data Science Code Documentation in Computational Notebooks
abstract
Computational notebooks allow data scientists to express their ideas through a combination of code and documentation. However, data scientists often pay attention only to the code, and neglect creating or updating their documentation during quick iterations. Inspired by human documentation practices learned from 80 highly-voted Kaggle notebooks, we design and implement Themisto, an automated documentation generation system to explore how human-centered AI systems can support human data scientists in the machine learning code documentation scenario. Themisto facilitates the creation of documentation via three approaches: a deep-learning-based approach to generate documentation for source code, a query-based approach to retrieve online API documentation for source code, and a user prompt approach to nudge users to write documentation. We evaluated Themisto in a within-subjects experiment with 24 data science practitioners, and found that automated documentation generation techniques reduced the time for writing documentation, reminded participants to document code they would have ignored, and improved participants’ satisfaction with their computational notebook.
April Yi Wang, Dakuo Wang, Jaimie Drozdal, Michael J. Muller, Soya Park, Justin D. Weisz, Xuye Liu, Lingfei Wu 0001, Casey Dugan
ACM Trans. Comput. Hum. Interact.5
2021 Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
abstract
Data scientists face a steep learning curve in understanding a new domain for which they want to build machine learning (ML) models. While input from domain experts could offer valuable help, such input is often limited, expensive, and generally not in a form readily consumable by a model development pipeline. In this paper, we propose Ziva, a framework to guide domain experts in sharing essential domain knowledge to data scientists for building NLP models. With Ziva, experts are able to distill and share their domain knowledge using domain concept extractors and five types of label justification over a representative data sample. The design of Ziva is informed by preliminary interviews with data scientists, in order to understand current practices of domain knowledge acquisition process for ML development projects. To assess our design, we run a mix-method case-study to evaluate how Ziva can facilitate interaction between domain experts and data scientists. Our results highlight that (1) domain experts are able to use Ziva to provide rich domain knowledge, while maintaining low mental load and stress levels; and (2) data scientists find Ziva’s output helpful for learning essential information about the domain, offering scalability of information, and lowering the burden on domain experts to share knowledge. We conclude this work by experimenting with building NLP models using the Ziva output for our case study.
Soya Park, April Yi Wang, Ban Kawas, Qingzi Vera Liao, David Piorkowski, Marina Danilevsky
IUI1
2021 How AI Developers Overcome Communication Challenges in a Multidisciplinary Team: A Case Study
abstract
The development of AI applications is a multidisciplinary effort, involving multiple roles collaborating with the AI developers, an umbrella term we use to include data scientists and other AI-adjacent roles on the same team. During these collaborations, there is a knowledge mismatch between AI developers, who are skilled in data science, and external stakeholders who are typically not. This difference leads to communication gaps, and the onus falls on AI developers to explain data science concepts to their collaborators. In this paper, we report on a study including analyses of both interviews with AI developers and artifacts they produced for communication. Using the analytic lens of shared mental models, we report on the types of communication gaps that AI developers face, how AI developers communicate across disciplinary and organizational boundaries, and how they simultaneously manage issues regarding trust and expectations.
David Piorkowski, Soya Park, April Yi Wang, Dakuo Wang, Michael J. Muller, Felix Portnoy
Proc. ACM Hum. Comput. Interact.2
2019 Opportunities for Automating Email Processing: A Need-Finding Study
abstract
Email management consumes significant effort from senders and recipients. Some of this work might be automatable. We performed a mixed-methods need-finding study to learn: (i) what sort of automatic email handling users want, and (ii) what kinds of information and computation are needed to support that automation. Our investigation included a design workshop to identify categories of needs, a survey to better understand those categories, and a classification of existing email automation software to determine which needs have been addressed. Our results highlight the need for: a richer data model for rules, more ways to manage attention, leveraging internal and external email context, complex processing such as response aggregation, and affordances for senders. To further investigate our findings, we developed a platform for authoring small scripts over a user's inbox. Of the automations found in our studies, half are impossible in popular email clients, motivating new design directions.
Soya Park, Amy X. Zhang, Luke S. Murray, David R. Karger
CHI1
2015 Practical message-passing framework for large-scale combinatorial optimization
abstract
Graphical Model (GM) has provided a popular framework for big data analytics because it often lends itself to distributed and parallel processing by utilizing graph-based ‘local’ structures. It models correlated random variables where in particular, the max-product Belief Propagation (BP) is the most popular heuristic to compute the most-likely assignment in GMs. In the past years, it has been proven that BP can solve a few classes of combinatorial optimization problems under certain conditions. Motivated by this, we explore the prospect of using BP to solve generic combinatorial optimization problems. The challenge is that, in practice, BP may converge very slowly and even if it does converge, the BP decision often violates the constraints of the original problem. This paper proposes a generic framework that enables us to apply BP-based algorithms to compute an approximate feasible solution for an arbitrary combinatorial optimization task. The main novel ingredients include (a) careful initialization of BP messages, (b) hybrid damping on BP updates, and (c) post-processing using BP beliefs. Utilizing the framework, we develop parallel algorithms for several large-scale combinatorial optimization problems including maximum weight matching, vertex cover and independent set. We demonstrate that our framework delivers high approximation ratio, speeds up the process by parallelization, and allows large-scale processing involving billions of variables.
Inho Cho, Soya Park, Dongsu Han, Jinwoo Shin
IEEE BigData2