Nischal Shrestha

dblp:225/7825 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2023 Detangler: Helping Data Scientists Explore, Understand, and Debug Data Wrangling Pipelines
abstract
Data scientists spend significant time on data wrangling-a process involving data cleaning, shaping, and pre-processing. Data wrangling requires meticulous exploration and backtracking to assess data quality by applying and validating numerous data transformation chains, making it a tedious and error-prone process. In this paper, we present Detangler, an interactive tool within the RStudio IDE that helps data scientists identify and debug data quality issues and wrangling code. The design of Detangler is informed via formative interviews, and it presents data scientists with (i) insights into potential data quality issues, and (ii) always-on visual summaries of the effects of individual data transformations, enabling interactive exploration of data and wrangling code. Through a laboratory study with 18 data scientists, triangulated with telemetry data, we find that Detangler improves exploration and debugging of data smells and data wrangling code. We discuss design implications for future tools for data science programming.
Nischal Shrestha, Bhavya Chopra, Austin Z. Henley, Chris Parnin
VL/HCC1
2022 ReBOC: Recommending Bespoke Open Source Software Projects to Contributors
abstract
Open Source Software for Social Good (OSS4SG) projects are projects that address a societal need and target people who need help. These projects often address high-impact humanitarian causes such as curating local health resources during a global pandemic, informing the public on the structural integrity of buildings, and encouraging civic engagement in times of strife. These projects carry a high intrinsic reward for contributing but are hard to find—prior research has shown that one of the the top challenges for contributors is not knowing where to find good projects to work on. Currently, contributors must manually search and assess whether projects align with their growing technical skills and intended impact interests.In this paper, we describe a recommendation system that automatically recommends OSS4SG projects for contributors based on their activity and project-related information. To score and rank projects, we calculated scores based on four signals: technical skills, interests, social ties, and recency of project activity. We performed an offline validation of the recommendation system using standard evaluation metrics such as the hit rate ratio. Results show that the signals are effective in producing a ranked list of OSS4SG projects for contributors, with room for improvement. Finally, we conducted a formative study with contributors to better understand their process of project discovery, validate our findings, and identify additional signals for future work to improve recommendations.
Denae Ford, Nischal Shrestha, Thomas Zimmermann 0001
VL/HCC2
2021 Unravel: A Fluent Code Explorer for Data Wrangling
abstract
Data scientists have adopted a popular design pattern in programming called the fluent interface for composing data wrangling code. The fluent interface works by combining multiple transformations on a data table—or dataframes—with a single chain of expressions, which produces an output. Although fluent code promotes legibility, the intermediate dataframes are lost, forcing data scientists to unravel the chain through tedious code edits and re-execution. Existing tools for data scientists do not allow easy exploration or support understanding of fluent code. To address this gap, we designed a tool called Unravel that enables structural edits via drag-and-drop and toggle switch interactions to help data scientists explore and understand fluent code. Data scientists can apply simple structural edits via drag-and-drop and toggle switch interactions to reorder and (un)comment lines. To help data scientists understand fluent code, Unravel provides function summaries and always-on visualizations highlighting important changes to a dataframe. We discuss the design motivations behind Unravel and how it helps understand and explore fluent code. In a first-use study with 14 data scientists, we found that Unravel facilitated diverse activities such as validating assumptions about the code or data, exploring alternatives, and revealing function behavior.
Nischal Shrestha, Titus Barik, Chris Parnin
UIST1
2021 Remote, but Connected: How #TidyTuesday Provides an Online Community of Practice for Data Scientists
abstract
Data science practitioners face the challenge of continually honing their skills such as data wrangling and visualization. As data scientists seek online spaces to network, learn and share resources with one another, each individual has to employ their own ad-hoc strategy to practice their data science skills. Given these disjointed efforts, it is crucial to ask: how can we build an inclusive, welcoming online community of practice that unites data scientists in their collective efforts to become experts? Daily hashtags on Twitter are used on specific days and have shown promise in forming a community of practice (CoP) in social networking sites like Twitter, but how do they benefit the community and its members? To understand how daily hashtags benefit data scientists and form an online CoP, we conducted a qualitative study on #TidyTuesday---a daily hashtag project for data scientists using R---using the framework of CoP as a lens for analysis. We conducted semi-structured interviews with 26 participants and uncovered motivations behind their participation in #TidyTuesday, how the project benefited them, and how it cultivated an online CoP. Our findings contribute to the CSCW research on community of practices by providing design trade-offs of using daily hashtags on Twitter, and guidelines on growing and sustaining an online community of practice for data scientists.
Nischal Shrestha, Titus Barik, Chris Parnin
Proc. ACM Hum. Comput. Interact.1
2020 Here we go again: why is it difficult for developers to learn another programming language?
abstract
Once a programmer knows one language, they can leverage concepts and knowledge already learned, and easily pick up another programming language. But is that always the case? To understand if programmers have difficulty learning additional programming languages, we conductedan empirical study of Stack Overflow questions across 18 different programming languages. We hypothesized that previous knowledge could potentially interfere with learning a new programming language. From our inspection of 450 Stack Overflow questions, we found 276 instances of interference that occurred due to faulty assumptions originating from knowledge about a different language. To understand why these difficulties occurred, we conducted semi-structured interviews with 16 professional programmers. The interviews revealed that programmers make failed attempts to relate a new programming language with what they already know. Our findings inform design implications for technical authors, toolsmiths, and language designers, such as designing documentation and automated tools that reduce interference, anticipating uncommon language transitions during language design, and welcoming programmers not just into a language, but its entire ecosystem.
Nischal Shrestha, Colton Botta, Titus Barik, Chris Parnin
ICSE1
2019 Exploring tools and strategies used during regular expression composition tasks
abstract
Regular expressions are frequently found in programming projects. Studies have found that developers can accurately determine whether a string matches a regular expression. However, we still do not know the challenges associated with composing regular expressions. We conduct an exploratory case study to reveal the tools and strategies developers use during regular expression composition. In this study, 29 students are tasked with composing regular expressions that pass unit tests illustrating the intended behavior. The tasks are in Java and the Eclipse IDE was set up with JUnit tests. Participants had one hour to work and could use any Eclipse tools, web search, or web-based tools they desired. Screen-capture software recorded all interactions with browsers and the IDE. We analyzed the videos quantitatively by transcribing logs and extracting personas. Our results show that participants were 30% successful (28 of 94 attempts) at achieving a 100% pass rate on the unit tests. When participants used tools frequently, as in the case of the novice tester and the knowledgeable tester personas, or when they guess at a solution prior to searching, they are more likely to pass all the unit tests. We also found that compile errors often arise when participants searched for a result and copy/pasted the regular expression from another language into their Java files. These results point to future research into making regular expression composition easier for programmers, such as integrating visualization into the IDE to reduce context switching or providing language migration support when reusing regular expressions written in another language to reduce compile errors.
Gina R. Bai, Brian Clee, Nischal Shrestha, Carl Chapman, Cimone Wright-Hamor, Kathryn T. Stolee
ICPC3
2019 Instrument Designs for Validating Cross-Language Behavioral Differences
abstract
Programmers are expected to use multiple programming languages frequently. Studies have found that programmers try to reuse existing knowledge from their previous languages. However, this strategy can result in misconceptions from previous languages. Current learning resources that support the strategy are limited because there is no systematic way to produce or validate the material. We designed three instruments that can help identify and validate meaningful behavior differences between two languages to pinpoint potential misconceptions. To validate the instruments, we examined how Python programmers predict behavior in a less familiar language like R, and whether they expect various R semantics. We found that the instruments are effective in validating differences between Python and R which were linked to misconceptions. We discuss design trade-offs between the three instruments and provide guidelines for researchers and educators in systematically validating programming misconceptions when switching to a new language.
Nischal Shrestha, Chris Parnin
VL/HCC1
2018 Towards Supporting Knowledge Transfer of Programming Languages
abstract
Today, there are hundreds of programming languages that are widely used. Programmers at all levels are expected to become proficient in multiple languages. Experienced programmers who have knowledge of at least one language are able to learn a second language much quicker than novices. However, the transfer process can still be difficult when there exists numerous differences from their previous language. Documentation, online courses and tutorials tend to present information geared towards novices. This type of presentation might suffice for beginners, but it doesn't support learning for experienced programmers [1] who would benefit from leveraging their knowledge of previous programming languages. In my work, I explore teaching programming languages through the lens of learning transfer, which occurs when learning in one context either enhances (positive transfer) or undermines (negative transfer) a related performance in another context. To investigate this approach, I created and evaluated a research tool called Transfer Tutor that teaches programmers R in terms of Python and Pandas, a data analysis library. The following design choices were made to explore learning transfer, applied to the topic of data frame manipulation: 1) highlighting similarities between syntax elements to support learning transfer 2) explicit tutoring on potential misconceptions 3) stepping through and highlighting elements of the snippets incrementally.
Nischal Shrestha
VL/HCC1
2018 It's Like Python But: Towards Supporting Transfer of Programming Language Knowledge
abstract
Expertise in programming traditionally assumes a binary novice-expert divide. Learning resources typically target programmers who are learning programming for the first time, or expert programmers for that language. An underrepresented, yet important group of programmers are those that are experienced in one programming language, but desire to author code in a different language. For this scenario, we postulate that an effective form of feedback is presented as a transfer from concepts in the first language to the second. Current programming environments do not support this form of feedback. In this study, we apply the theory of learning transfer to teach a language that programmers are less familiar with-such as R-in terms of a programming language they already know-such as Python. We investigate learning transfer using a new tool called Transfer Tutor that presents explanations for R code in terms of the equivalent Python code. Our study found that participants leveraged learning transfer as a cognitive strategy, even when unprompted. Participants found Transfer Tutor to be useful across a number of affordances like stepping through and highlighting facts that may have been missed or misunderstood. However, participants were reluctant to accept facts without code execution or sometimes had difficulty reading explanations that are verbose or complex. These results provide guidance for future designs and research directions that can support learning transfer when learning new programming languages.
Nischal Shrestha, Titus Barik, Chris Parnin
VL/HCC1