VLDB 2026 Research / reviewers in the wild / expert
William Thies
dblp:24/1462
· DBLP profile ↗
52ranked-venue papers
9as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 19 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-authorSystems, architecture and hardware · 11 · 3 first-authorSoftware engineering, systems software and programming languages · 11 · 3 first-authorArtificial intelligence and machine learning · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | "Is it Even Giving the Correct Reading or Not?": How Trust and Relationships Mediate Blood Pressure Management in IndiaabstractWhile chronic disease afflicts a large Indian population, the technologies used to manage chronic diseases have largely been informed by studies conducted in other sociocultural contexts. To address this gap, we conducted qualitative interviews with 21 patients clinically diagnosed with abnormal blood pressure (BP) living in low-resourced communities of Haryana, Uttarakhand and Uttar Pradesh in India. We found that patients’ trust in the BP ecosystem and social ties plays a significant role in shaping their perceptions of technology and chronic care. Trust in one actor of the ecosystem fosters trust in another, e.g., trust in BP reading depended on the type of device and the person measuring the BP. We also observed nuanced sharing and intermediation of BP devices. Based on our findings, we recommend designs to boost patients’ trust, familiarity, and access to technologies used in BP management and improve their experience of care in low-resource settings in India. Nimisha Karnatak, Brooke Loughrin, Tiffany Kuo, Odeline Mateu-Silvernail, Indrani Medhi-Thies, William Thies |
ACM Trans. Comput. Hum. Interact. | 6 |
| 2022 | Exploring Collection of Sign Language Videos through CrowdsourcingabstractInadequate sign language data currently impedes advancement of sign language ML and AI. Training on existing datasets results in limited models due to small size, and lack of diverse signers in real-world settings. Complex labeling problems in particular often limit scale. In this work, we explore the potential for crowdsourcing to help overcome these barriers. To do this, we ran a user study with exploratory crowdsourcing tasks designed to support scalability: 1) to record videos of specific content -- thereby enabling automatic, scalable labeling -- and 2) to perform quality control checks for execution consistency -- further reducing post-processing requirements. We also provided workers with a searchable view of the crowdsourced dataset, to boost engagement and transparency and align with Deaf community values. Our user study included 29 participants using our exploratory tasks to record 1906 videos and perform 2331 quality control checks. Our results suggest that a crowd of signers may be able to generate high-quality recordings and perform reliable quality control, and that the signing community values visibility into the resulting dataset. Danielle Bragg, Abraham Glasser, Fyodor O. Minakov, Naomi Caselli, William Thies |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2021 | ASL Sea Battle: Gamifying Sign Language Data CollectionabstractThe development of accurate machine learning models for sign languages like American Sign Language (ASL) has the potential to break down communication barriers for deaf signers. However, to date, no such models have been robust enough for real-world use. The primary barrier to enabling real-world applications is the lack of appropriate training data. Existing training sets suffer from several shortcomings: small size, limited signer diversity, lack of real-world settings, and missing or inaccurate labels. In this work, we present ASL Sea Battle, a sign language game designed to collect datasets that overcome these barriers, while also providing fun and education to users. We conduct a user study to explore the data quality that the game collects, and the user experience of playing the game. Our results suggest that ASL Sea Battle can reliably collect and label real-world sign language videos, and provides fun and education at the expense of data throughput. Danielle Bragg, Naomi Caselli, John W. Gallagher, Miriam Goldberg, Courtney J. Oka, William Thies |
CHI | 6 |
| 2020 | Chat in the Hat: A Portable Interpreter for Sign Language UsersabstractMany Deaf and Hard-of-Hearing (DHH) individuals rely on sign language interpreting to communicate with hearing peers. If on-site interpreting is not available, DHH individuals may use remote interpreting over a smartphone video-call. However, this solution requires the DHH individual to give up either 1) the use of one signing hand by holding the smartphone or 2) their ability to multitask and move around by propping the smartphone up in a fixed location. We explore this problem within the context of the workplace, and present a prototype hands-free device using augmented reality glasses with a hat-mounted fisheye camera and mic/speaker. To explore the validity of our design, we conducted 1) a video interpretability experiment, and 2) a user study with 18 participants (9 DHH, 9 hearing) in a workplace environment. Our results suggest that a hands-free device can support accurate interpretation while enhancing personal interactions. Larwan Berke, William Thies, Danielle Bragg |
ASSETS | 2 |
| 2020 | Exploring Collection of Sign Language Datasets: Privacy, Participation, and Model PerformanceabstractAs machine learning algorithms continue to improve, collecting training data becomes increasingly valuable. At the same time, increased focus on data collection may introduce compounding privacy concerns. Accessibility projects in particular may put vulnerable populations at risk, as disability status is sensitive, and collecting data from small populations limits anonymity. To help address privacy concerns while maintaining algorithmic performance on machine learning tasks, we propose privacy-enhancing distortions of training datasets. We explore this idea through the lens of sign language video collection, which is crucial for advancing sign language recognition and translation. We present a web study exploring signers’ concerns in contributing to video corpora and their attitudes about using filters, and a computer vision experiment exploring sign language recognition performance with filtered data. Our results suggest that privacy concerns may exist in contributing to sign language corpora, that filters (especially expressive avatars and blurred faces) may impact willingness to participate, and that training on more filtered data may boost recognition accuracy in some cases. Danielle Bragg, Oscar Koller, Naomi Caselli, William Thies |
ASSETS | 4 |
| 2020 | Using Mobile Airtime Credits to Incentivize Learning, Sharing and Survey Response: Experiences from the FieldabstractIn the Global South, mobile airtime payment has emerged as a popular way to incentivize different research studies, including ones on survey completion or disseminating information to people. Building on this literature, we report deployment experiences from three different studies in India that used airtime incentives. The first was used to promote awareness about HIV/AIDS, the second for promoting awareness and surveying preparedness for an upcoming election, and the third to measure learning and encourage people to vote in a conflict-hit region for a different election. Unlike past work, we found that a delivery mechanism that focuses on asking questions first, rather than presenting a tutorial and then asking questions, worked well in practice. In addition, we found multiple challenges in adoption of the technology and tried different ways to incentivize peer sharing of our system. Between the three deployments, we also addressed other technical and human-centered challenges such as delayed airtime payments and people using the system on behalf of someone else. We hope that our experiences and insights can be helpful to others seeking to deploy applications that utilize mobile airtime payments for learning, sharing, and survey response. Devansh Mehta, Ramaravind Kommiya Mothilal, Alok Sharma, William Thies, Amit Sharma 0007 |
COMPASS | 4 |
| 2019 | Exploring Crowdsourced Work in Low-Resource SettingsabstractWhile researchers have studied the benefits and hazards of crowdsourcing for diverse classes of workers, most work has focused on those having high familiarity with both computers and English. We explore whether paid crowdsourcing can be inclusive of individuals in rural India, who are relatively new to digital devices and literate mainly in local languages. We built an Android application to measure the accuracy with which participants can digitize handwritten Marathi/Hindi words. The tasks were based on the real-world need for digitizing handwritten Devanagari script documents. Results from a two-week, mixed-methods study show that participants achieved 96.7% accuracy in digitizing handwritten words on low-end smartphones. A crowdsourcing platform that employs these users performs comparably to a professional transcription firm. Participants showed overwhelming enthusiasm for completing tasks, so much so that we recommend imposing limits to prevent overuse of the application. We discuss the implications of these results for crowdsourcing in low-resource areas. Manu Chopra, Indrani Medhi-Thies, Joyojeet Pal, Colin Scott, William Thies, Vivek Seshadri |
CHI | 5 |
| 2019 | 99DOTS: a low-cost approach to monitoring and improving medication adherenceabstractEnsuring that patients adhere to prescribed medication remains an important challenge in global health. While technology has been utilized to monitor and improve adherence, solutions to date have been too costly for large-scale deployment in developing regions. This paper describes 99DOTS, a low-cost approach for tracking adherence using a combination of paper packaging and low-end mobile phones. Every day, patients reveal an unpredictable phone number behind the pills and send a free call to that number to indicate that drugs were dispensed and taken. Within five years of its inception, 99DOTS has become a standard of care for tuberculosis in India and has enrolled over 200,000 patients. We provide a holistic account of the project's evolution, including its iterative design, scaled implementation, and lessons learned along the way. We hope this account will serve as a useful case study for anyone seeking to establish and scale new low-cost technologies for a global audience. Nakull Gupta, Brandon Liu, Vineet Nair, Reena Kuttan, Priyanka Ivatury, Amy Chen, Kshama Lakshman, Rashmi Rodrigues, George D'Souza, Deepti Chittamuru, Raghuram Rao, Kiran Rade, Bhavin Vadera, Daksha Shah, Vinod Choudhary, Vineet Chadha, Amar Shah 0003, Sameer Kumta, Puneet Dewan, Bruce Thomas, William Thies |
ICTD | 23 |
| 2019 | Learn2Earn: Using Mobile Airtime Incentives to Bolster Public Awareness CampaignsabstractIn rural parts of the developing world, spreading awareness about critical issues in health, governance, and other topics is challenging and costly. Traditional media such as print, radio and TV each have limitations and offer little guarantee that new information is absorbed or retained by the target population. This paper describes Learn2Earn, a system that leverages mobile payments to bolster public awareness campaigns in rural India. Users call an Interactive Voice Response (IVR) system, listen to a brief audio tutorial, and take a multiple-choice quiz to check their understanding. People who pass the quiz receive a mobile top-up (about $0.14) and have the opportunity to earn additional credits by referring others to the system. We describe a pilot deployment of Learn2Earn in rural India that spread via word-of-mouth to over 15,000 people within seven weeks. Usage was concentrated among young men, many of them students. In a mixed-methods study, we draw upon call logs, electronic surveys, qualitative interviews, and other sources of data to suggest that Learn2Earn could be an effective way to build awareness about important topics. Sai Swaminathan, Indrani Medhi-Thies, Devansh Mehta, Edward Cutrell, Amit Sharma 0007, William Thies |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2016 | ICT-Enabled Grievance Redressal in Central India: A Comparative AnalysisabstractHelping citizens to resolve grievances is an important part of many e-governance initiatives. In this paper, we examine two contemporary initiatives that use ICTs to help citizens resolve grievances in central India. One system is a state-run call center (the CM Helpline), while the other is an independent citizen journalism service (CGNet Swara). Despite similarities in their high-level goals, approach, and geographies served, the systems have key differences in their use of technology, their level of transparency, and their relationship to government. Using qualitative interviews, field immersions, and other data, we analyze how these differences impact the experiences of citizens, officials, and the intermediaries between them. We synthesize our observations into a set of recommendations for the design of future ICT-enabled grievance redressal systems. Megh Marathe, Jacki O'Neill, Paromita Pain, William Thies |
ICTD | 4 |
| 2015 | Sangeet Swara: A Community-Moderated Voice Forum in Rural IndiaabstractInteractive voice forums have emerged as a promising platform for people in developing regions to record and share audio messages using low-end mobile phones. However, one of the barriers to the scalability of voice forums is the process of screening and categorizing content, often done by a dedicated team of moderators. We present Sangeet Swara, a voice forum for songs and cultural content that relies on the community of callers to curate high-quality posts that are prioritized for playback to others. An 11-week deployment of Sangeet Swara found broad and impassioned usage, especially among visually impaired users. We also conducted a follow-up experiment, called Talent Hunt, that sought to reduce reliance on toll-free telephone lines. Together, our deployments span about 53,000 calls from 13,000 callers, who submitted 6,000 posts and 150,000 judgments of other content. Using a mixed-methods analysis of call logs, audio content, comparison with outside judges, and 204 automated phone surveys, we evaluate the user experience, the strengths and weaknesses of community moderation, financial sustainability, and the implications for future systems. Aditya Vashistha, Edward Cutrell, Gaetano Borriello, William Thies |
CHI | 4 |
| 2015 | Increasing the Reach of Snowball Sampling: The Impact of Fixed versus Lottery IncentivesabstractThough many researchers have studied how to incentivize people to respond to surveys, little is known about how these incentives impact respondents' willingness to recruit others to participate as well. In this paper, we show that the incentives offered for individual survey responses can have a dramatic impact on the overall reach of a survey through a network of peers. In a field experiment in India, we made a survey accessible via mobile phones and offered respondents either a fixed incentive (guaranteed payment of about $0.17) or a lottery incentive (1% chance of winning $17). When asked to choose, a significant fraction of respondents preferred the lottery incentive. However, when encouraged to spread the survey, the fixed incentive spread over 100 times further, reaching about 800 people in a day. We interpret this surprising result and discuss the implications for HCI. Aditya Vashistha, Edward Cutrell, William Thies |
CSCW | 3 |
| 2015 | Revisiting CGNet Swara and its impact in rural IndiaabstractCGNet Swara is a voice-based platform for citizen journalism, launched in rural India in 2010. Since then, CGNet Swara has logged over 575,000 phone calls, over 6,900 published stories, and 287 reports of specific problems that were solved via the system. In this paper, we characterize the ongoing impact of CGNet Swara using a mixed-methods approach that includes 70 interviews with contributors, listeners, moderators, journalists, officials, and other actors. Our analysis also draws on the content of published posts, two focus groups, and a 9-day field immersion. Our results highlight personal narratives of the transformative benefits CGNet Swara has brought to rural communities. While the resolution of grievances is the most visible impact, we also uncover a diverse portfolio of other impacts connected to contributing and listening to the platform, as well as opportunities to further enhance impact. Our work contributes to the dialogue surrounding the impact of ICTD projects, especially those that span multiple years. Megh Marathe, Jacki O'Neill, Paromita Pain, William Thies |
ICTD | 4 |
| 2015 | The Whodunit Challenge: Mobilizing the Crowd in India
Aditya Vashistha, Rajan Vaish, Edward Cutrell, William Thies |
INTERACT (2) | 4 |
| 2015 | Measuring and Maximizing the Effectiveness of Honor Codes in Online CoursesabstractWe measure the effectiveness of a traditional honor code at deterring cheating in an online examination, and we compare it to that of a stern warning. Through experimental evaluation in a 409-student online course, we find that a pre-task warning leads to a significant decrease in the rate of cheating while an honor code has a smaller (non-significant) effect. Unlike much prior work, we measure the rate of cheating directly and we do not rely on potentially inaccurate post-examination surveys. Our findings demonstrate that replacing traditional honor codes with warnings could be a simple and effective way to deter cheating in online courses. Henry Corrigan-Gibbs, Nakull Gupta, Curtis G. Northcutt, Edward Cutrell, William Thies |
L@S | 5 |
| 2015 | Blended Learning in Indian Colleges with Massively Empowered ClassroomabstractStudents in the developing world are frequently cited as being among the most important beneficiaries of online education initiatives such as massive open online courses (MOOCs). While some predict that online classrooms will replace physical classrooms, our experience suggests that blending online and in-person instruction is more likely to succeed in developing regions. However, very little research has actually been done on the effects of online education or blended learning in these environments. In this paper we describe a blended learning initiative that combines videos from a large online course with peer-led sessions for undergraduate technical education in India. We performed a randomized controlled trial (RCT) that indicates our intervention was associated with a small but significant improvement in performance on a summative exam. We discuss the results of the RCT and an ethnographic study of the intervention to make recommendations for future, scalable blended learning initiatives for places such as India. Edward Cutrell, Jacki O'Neill, Srinath Bala, B. Nitish, Nakull Gupta, Viraj Kumar, William Thies |
L@S | 8 |
| 2015 | Source Effects in Online EducationabstractWhile most MOOCs rely on world-famous experts to teach the masses, in many circumstances students may learn more from people who share their context such as local teachers or peers. Here, we describe an experiment to explore how the "source" of video content, the teacher, affects online learning, specifically in the context of higher education in Indian colleges. The proposed experiment will compare three content sources -- a local lecturer (teacher from an Indian engineering college), a local peer (both male and female students similar to the targeted audience), and an internationally recognized expert (a Stanford lecturer). Students will watch videos by the various source authors, after which we will measure differences in their preference, engagement, and learning. In addition, we discuss our experiences with helping students prepare video lectures and describe the support and processes we used to curate interesting and clear peer-generated content. Nakull Gupta, Jacki O'Neill, Edward Cutrell, William Thies |
L@S | 5 |
| 2015 | Deterring Cheating in Online EnvironmentsabstractMany Internet services depend on the integrity of their users, even when these users have strong incentives to behave dishonestly. Drawing on experiments in two different online contexts, this study measures the prevalence of cheating and evaluates two different methods for deterring it. Our first experiment investigates cheating behavior in a pair of online exams spanning 632 students in India. Our second experiment examines dishonest behavior on Mechanical Turk through an online task with 2,378 total participants. Using direct measurements that are not dependent on self-reports, we detect significant rates of cheating in both environments. We confirm that honor codes--despite frequent use in massive open online courses (MOOCs)--lead to only a small and insignificant reduction in online cheating behaviors. To overcome these challenges, we propose a new intervention: a stern warning that spells out the potential consequences of cheating. We show that the warning leads to a significant (about twofold) reduction in cheating, consistent across experiments. We also characterize the demographic correlates of cheating on Mechanical Turk. Our findings advance the understanding of cheating in online environments, and suggest that replacing traditional honor codes with warnings could be a simple and effective way to deter cheating in online courses and online labor marketplaces. Henry Corrigan-Gibbs, Nakull Gupta, Curtis G. Northcutt, Edward Cutrell, William Thies |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2014 | VidWiki: enabling the crowd to improve the legibility of online educational videosabstractVideos are becoming an increasingly popular medium for communicating information, especially for online education. Recent efforts by organizations like Coursera, edX, Udacity and Khan Academy have produced thousands of educational videos with hundreds of millions of views in their attempt to make high quality teaching available to the masses. As a medium, videos are time-consuming to produce and cannot be easily modified after release. As a result, errors or problems with legibility are common. While text-based information platforms like Wikipedia have benefitted enormously from crowdsourced contributions for the creation and improvement of content, the various limitations of video hinder the collaborative editing and improvement of educational videos. To address this issue, we present VidWiki, an online platform that enables students to iteratively improve the presentation quality and content of educational videos. Through the platform, users can improve the legibility of handwriting, correct errors, or translate text in videos by overlaying typeset content such as text, shapes, equations, or images. We conducted a small user study in which 13 novice users annotated and revised Khan Academy videos. Our results suggest that with only a small investment of time on the part of viewers, it may be possible to make meaningful improvements in online educational videos. Mydhili Bayyapunedi, Dilip Ravindran, Edward Cutrell, William Thies |
CSCW | 5 |
| 2014 | Online learning versus blended learning: an exploratory studyabstractDue to the recent emergence of massive open online courses (MOOCs), students and teachers are gaining unprecedented access to high-quality educational content. However, many questions remain on how best to utilize that content in a classroom environment. In this small-scale, exploratory study, we compared two ways of using a recorded video lecture. In the online learning condition, students viewed the video on a personal computer, and also viewed a follow-up tutorial (a quiz review) on the computer. In the blended learning condition, students viewed the video as a group in a classroom, and received the follow-up tutorial from a live lecturer. We randomly assigned 102 students to these conditions, and assessed learning outcomes via a series of quizzes. While we saw significant learning gains after each session conducted, we did not observe any significant differences between the online and blended learning groups. We discuss these findings as well as areas for future work. Balasubramanyan Ashok, Srinath Bala, Edward Cutrell, Naren Datha, Rahul Kumar 0002, Viraj Kumar, P. Madhusudan, Siddharth Prakash, Sriram K. Rajamani, Satish Sangameswaran, William Thies |
L@S | 13 |
| 2013 | TypeRighting: combining the benefits of handwriting and typeface in online educational videosabstractRecent years have seen enormous growth of online educational videos, spanning K-12 tutorials to university lectures. As this content has grown, so too has grown the number of presentation styles. Some educators have strong allegiance to handwritten recordings (using pen and tablet), while others use only typed (PowerPoint) presentations. In this paper, we present the first systematic comparison of these two presentation styles and how they are perceived by viewers. Surveys on edX and Mechanical Turk suggest that users enjoy handwriting because it is personal and engaging, yet they also enjoy typeface because it is clear and legible. Based on these observations, we propose a new presentation style, TypeRighting, that combines the benefits of handwriting and typeface. Each phrase is written by hand, but fades into typeface soon after it appears. Our surveys suggest that about 80% of respondents prefer TypeRighting over handwriting. The same fraction of respondents prefer TypeRighting over typeface, for videos in which the handwriting is sufficiently legible. Mydhili Bayyapunedi, Edward Cutrell, Anant Agarwal, William Thies |
CHI | 5 |
| 2013 | Using automated voice calls to improve adherence to iron supplements during pregnancy: a pilot studyabstractFor years, researchers have explored the use of mobile phone reminders to improve adherence to medication. However, few studies have measured the direct medical benefit of those reminders, especially for low-literate populations in the developing world. This paper describes the use of automated voice calls to promote adherence to iron supplements among pregnant women in urban India. Unlike prior studies, we assess impact via a direct measurement of hemoglobin (Hb) levels in the blood. We enrolled 130 pregnant women from a low-income area of Mumbai, India and randomly assigned them to control and treatment groups. Both groups received a counseling session and a free supply of medication. The treatment group also received short audio messages, three times per week for a period of three months, encouraging them to take iron supplements. Results suggest that automated calls positively impacted Hb levels. However, because we could only recover 79 women for follow-up, and the effect size was small, our results lack statistical power (average change in Hb = 0.43 g/dL, 95% CI = -0.13 0.98 g/dL, p=0.13). We conclude that automated calls deserve further consideration for reducing maternal anemia, and we share our lessons learned for the benefit of future interventions. Niranjan Pai, Pradnya Supe, Shailesh Kore, Y. S. Nandanwar, Aparna Hegde, Edward Cutrell, William Thies |
ICTD (1) | 7 |
| 2012 | "Yours is better!": participant response bias in HCIabstractAlthough HCI researchers and practitioners frequently work with groups of people that differ significantly from themselves, little attention has been paid to the effects these differences have on the evaluation of HCI systems. Via 450 interviews in Bangalore, India, we measure participant response bias due to interviewer demand characteristics and the role of social and demographic factors in influencing that bias. We find that respondents are about 2.5x more likely to prefer a technological artifact they believe to be developed by the interviewer, even when the alternative is identical. When the interviewer is a foreign researcher requiring a translator, the bias towards the interviewer's artifact increases to 5x. In fact, the interviewer's artifact is preferred even when it is degraded to be obviously inferior to the alternative. We conclude that participant response bias should receive more attention within the CHI community, especially when designing for underprivileged populations. Nicola Dell, Vidya Vaidyanathan, Indrani Medhi-Thies, Edward Cutrell, William Thies |
CHI | 5 |
| 2012 | mClerk: enabling mobile crowdsourcing in developing regionsabstractGlobal crowdsourcing platforms could offer new employment opportunities to low-income workers in developing countries. However, the impact to date has been limited because poor communities usually lack access to computers and the Internet. Aakar Gupta, William Thies, Edward Cutrell, Ravin Balakrishnan |
CHI | 2 |
| 2012 | Emergent practices around CGNet Swara, voice forum for citizen journalism in rural IndiaabstractRural communities in India are often underserved by the mainstream media. While there is a public discourse surrounding the issues they face, this dialogue typically takes place on television, in newspaper editorials, and on the Internet. Unfortunately, participation in such forums is limited to the most privileged members of society, excluding those individuals who have the largest stake in the conversation. Preeti Mudliar, Jonathan Donner, William Thies |
ICTD | 3 |
| 2012 | Biometric Monitoring as a Persuasive Technology: Ensuring Patients Visit Health Centers in India's Slums
Nupur Bhatnagar, Abhishek Sinha, Navkar Samdaria, Aakar Gupta, Shelly Batra, Manish Bhardwaj, William Thies |
PERSUASIVE | 7 |
| 2012 | An Empirical Study of License Violations in Open Source ProjectsabstractThe use of Open Source Software (OSS) components in building applications has presented the challenge of integrating them in a way such that the licenses of the individual components do not conflict with each other and if applicable, the overall license of the application. These conflicts lead to violations, with many having far reaching legal consequences. While proprietary software firms are often plagued with the risks of not satisfying the clauses of OSS licenses, we hypothesize that a large degree of code reuse within the OSS community poses similar threats too. Through an analysis of 1423 projects, consisting of approximately 69 million non-blank lines of code from Google Code project hosting, we validate instances of code reuse between projects by comparing their licenses. Our results discover four violations, evaluated by searching for files that share similar content. Additionally, we present statistics on code reuse within the set of projects. Arunesh Mathur, Harshal Choudhary, Priyank Vashist, William Thies, P. Santhi Thilagam |
SEW | 4 |
| 2012 | Low-cost audience polling using computer visionabstractElectronic response systems known as "clickers" have demonstrated educational benefits in well-resourced classrooms, but remain out-of-reach for most schools due to their prohibitive cost. We propose a new, low-cost technique that utilizes computer vision for real-time polling of a classroom. Our approach allows teachers to ask a multiple-choice question. Students respond by holding up a qCard: a sheet of paper that contains a printed code, similar to a QR code, encoding their student IDs. Students indicate their answers (A, B, C or D) by holding the card in one of four orientations. Using a laptop and an off-the-shelf webcam, our software automatically recognizes and aggregates the students' responses and displays them to the teacher. We built this system and performed initial trials in secondary schools in Bangalore, India. In a 25-student classroom, our system offers 99.8% recognition accuracy, captures 97% of responses within 10 seconds, and costs 15 times less than existing electronic solutions. Edward Cutrell, William Thies |
UIST | 3 |
| 2011 | Utilizing DVD players as low-cost offline internet browsersabstractIn the developing world, computers and Internet access remain rare. However, there are other devices that can be used to deliver information, including TVs and DVD players. In this paper, we work to bridge this gap by delivering offline Internet content on DVD, for interactive playback on ordinary DVD players. Using the remote control, users can accomplish all of the major functions available in a Web browser, including navigation, hyperlinks, and search. Gaurav Paruthi, William Thies |
CHI | 2 |
| 2011 | ALTER: exploiting breakable dependences for parallelizationabstractFor decades, compilers have relied on dependence analysis to determine the legality of their transformations. While this conservative approach has enabled many robust optimizations, when it comes to parallelization there are many opportunities that can only be exploited by changing or re-ordering the dependences in the program. Abhishek Udupa, Kaushik Rajan, William Thies |
PLDI | 3 |
| 2011 | Designing mobile interfaces for novice and low-literacy usersabstractWhile mobile phones have found broad application in bringing health, financial, and other services to the developing world, usability remains a major hurdle for novice and low-literacy populations. In this article, we take two steps to evaluate and improve the usability of mobile interfaces for such users. First, we offer an ethnographic study of the usability barriers facing 90 low-literacy subjects in India, Kenya, the Philippines, and South Africa. Then, via two studies involving over 70 subjects in India, we quantitatively compare the usability of different points in the mobile design space. In addition to text interfaces such as electronic forms, SMS, and USSD, we consider three text-free interfaces: a spoken dialog system, a graphical interface, and a live operator. Our results confirm that textual interfaces are unusable by first-time low-literacy users, and error prone for literate but novice users. In the context of healthcare, we find that a live operator is up to ten times more accurate than text-based interfaces, and can also be cost effective in countries such as India. In the context of mobile banking, we find that task completion is highest with a graphical interface, but those who understand the spoken dialog system can use it more quickly due to their comfort and familiarity with speech. We synthesize our findings into a set of design recommendations. Indrani Medhi-Thies, Somani Patnaik, Emma Brunskill, S. N. Nagasena Gautama, William Thies, Kentaro Toyama |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2010 | An empirical characterization of stream programs and its implications for language and compiler designabstractStream programs represent an important class of high-performance computations. Defined by their regular processing of sequences of data, stream programs appear most commonly in the context of audio, video, and digital signal processing, though also in networking, encryption, and other areas. In order to develop effective compilation techniques for the streaming domain, it is important to understand the common characteristics of these programs. Prior characterizations of stream programs have examined legacy implementations in C, C++, or FORTRAN, making it difficult to extract the high-level properties of the algorithms. William Thies, Saman P. Amarasinghe |
PACT | 1 |
| 2010 | Interactive DVDs as a platform for educationabstractWhile many technologies remain out-of-reach for households in the developing world, one exception to this rule is that of entertainment technologies. Even in poor communities, there is a strong drive to own devices such as TVs and, increasingly, DVD players. Though they are typically used for video content, ordinary DVD players also support rich interactivity and programmability, including the capability to browse over 100,000 menus using the remote control. Our vision is to leverage these capabilities to support interactive applications -- such as encyclopedias, language tutoring, and medical decision systems -- without any dependence on a computer. Kiran Gaikwad, Gaurav Paruthi, William Thies |
ICTD | 3 |
| 2009 | Computer-aided design for microfluidic chips based on multilayer soft lithographyabstractMicrofluidic chips are emerging as a powerful platform for automating biology experiments. As it becomes possible to integrate tens of thousands of components on a single chip, researchers will require design automation tools to push the scale and complexity of their designs to match the capabilities of the substrate. However, to date such tools have focused only on droplet-based devices, leaving out the popular class of chips that are based on multilayer soft lithography. In this paper, we develop design automation techniques for microfluidic chips based on multilayer soft lithography. We focus our attention on the control layer, which is driven by pressure actuators to invoke the desired flows on chip. We present a language in which designers can specify the Instruction Set Architecture (ISA) of a microfluidic device. Given an ISA, we automatically infer the locations of valves needed to implement the ISA. We also present novel algorithms for minimizing the number of control lines needed to drive the valves, as well as for routing valves to control ports while admitting sharing between the control lines. To the microfluidic community, we offer a free computer-aided design tool, Micado, which implements a subset of our algorithms as a practical plug-in to AutoCAD. Micado is being used successfully by microfluidic designers. We demonstrate its performance on three realistic chips. Nada Amin, William Thies, Saman P. Amarasinghe |
ICCD | 2 |
| 2009 | Evaluating the accuracy of data collection on mobile phones: A study of forms, SMS, and voiceabstractWhile mobile phones have found broad application in reporting health, financial, and environmental data, there has been little study of the possible errors incurred during mobile data collection. This paper provides the first (to our knowledge) quantitative evaluation of data entry accuracy on mobile phones in a resource-poor setting. Via a study of 13 users in Gujarat, India, we evaluated three user interfaces: 1) electronic forms, containing numeric fields and multiple-choice menus, 2) SMS, where users enter delimited text messages according to printed cue cards, and 3) voice, where users call an operator and dictate the data in real-time. Our results indicate error rates (per datum entered) of 4.2% for electronic forms, 4.8% for SMS, and 0.45% for voice. These results caused us to migrate our own initiative (a tuberculosis treatment program in rural India) from electronic forms to voice, in order to avoid errors on critical health data. While our study has some limitations, including varied backgrounds and training of participants, it suggests that some care is needed in deploying electronic interfaces in resource-poor settings. Further, it raises the possibility of using voice as a low-tech, high-accuracy, and cost-effective interface for mobile data collection. Somani Patnaik, Emma Brunskill, William Thies |
ICTD | 3 |
| 2009 | Manipulating lossless video in the compressed domainabstractA compressed-domain transformation is one that operates directly on the compressed format, rather than requiring conversion to an uncompressed format prior to processing. Performing operations in the compressed domain offers large speedups, as it reduces the volume of data processed and avoids the overhead of re-compression. William Thies, Steven Hall, Saman P. Amarasinghe |
ACM Multimedia | 1 |
| 2008 | Abstraction layers for scalable microfluidic biocomputing
William Thies, John Paul Urbanski, Todd Thorsen, Saman P. Amarasinghe |
Nat. Comput. | 1 |
| 2007 | A Practical Approach to Exploiting Coarse-Grained Pipeline Parallelism in C ProgramsabstractThe emergence of multicore processors has heightened the need for effective parallel programming practices. In addition to writing new parallel programs, the next generation of programmers will be faced with the overwhelming task of migrating decades' worth of legacy C code into a parallel representation. Addressing this problem requires a toolset of parallel programming primitives that can broadly apply to both new and existing programs. While tools such as threads and OpenMP allow programmers to express task and data parallelism, support for pipeline parallelism is distinctly lacking. In this paper, we offer a new and pragmatic approach to leveraging coarse-grained pipeline parallelism in C programs. We target the domain of streaming applications, such as audio, video, and digital signal processing, which exhibit regular flows of data. To exploit pipeline parallelism, we equip the programmer with a simple set of annotations (indicating pipeline boundaries) and a dynamic analysis that tracks all communication across those boundaries. Our analysis outputs a stream graph of the application as well as a set of macros for parallelizing the program and communicating the data needed. We apply our methodology to six case studies, including MPEG-2 decoding, MP3 decoding, GMTI radar processing, and three SPEC benchmarks. Our analysis extracts a useful block diagram for each application, and the parallelized versions offer a 2.78x mean speedup on a 4-core machine. William Thies, Vikram Chandrasekhar, Saman P. Amarasinghe |
MICRO | 1 |
| 2007 | Learning biophysically-motivated parameters for alpha helix predictionabstractBACKGROUND: Our goal is to develop a state-of-the-art protein secondary structure predictor, with an intuitive and biophysically-motivated energy model. We treat structure prediction as an optimization problem, using parameterizable cost functions representing biological "pseudo-energies". Machine learning methods are applied to estimate the values of the parameters to correctly predict known protein structures. RESULTS: Focusing on the prediction of alpha helices in proteins, we show that a model with 302 parameters can achieve a Qalpha value of 77.6% and an SOValpha value of 73.4%. Such performance numbers are among the best for techniques that do not rely on external databases (such as multiple sequence alignments). Further, it is easier to extract biological significance from a model with so few parameters. CONCLUSION: The method presented shows promise for the prediction of protein secondary structure. Biophysically-motivated elementary free-energies can be learned using SVM techniques to construct an energy cost function whose predictive performance rivals state-of-the-art. This method is general and can be extended beyond the all-alpha case described here. Blaise Gassend, Charles W. O'Donnell, William Thies, Marten van Dijk, Srini Devadas |
BMC Bioinform. | 3 |
| 2007 | A step towards unifying schedule and storage optimizationabstractWe present a unified mathematical framework for analyzing the tradeoffs between parallelism and storage allocation within a parallelizing compiler. Using this framework, we show how to find a good storage mapping for a given schedule, a good schedule for a given storage mapping, and a good storage mapping that is valid for all legal (one-dimensional affine) schedules. We consider storage mappings that collapse one dimension of a multidimensional array, and programs that are in a single assignment form and accept a one-dimensional affine schedule. Our method combines affine scheduling techniques with occupancy vector analysis and incorporates general affine dependences across statements and loop nests. We formulate the constraints imposed by the data dependences and storage mappings as a set of linear inequalities, and apply numerical programming techniques to solve for the shortest occupancy vector. We consider our method to be a first step towards automating a procedure that finds the optimal tradeoff between parallelism and storage space. William Thies, Frédéric Vivien, Saman P. Amarasinghe |
ACM Trans. Program. Lang. Syst. | 1 |
| 2006 | Exploiting coarse-grained task, data, and pipeline parallelism in stream programsabstractAs multicore architectures enter the mainstream, there is a pressing demand for high-level programming models that can effectively map to them. Stream programming offers an attractive way to expose coarse-grained parallelism, as streaming applications (image, video, DSP, etc.) are naturally represented by independent filters that communicate over explicit data channels.In this paper, we demonstrate an end-to-end stream compiler that attains robust multicore performance in the face of varying application characteristics. As benchmarks exhibit different amounts of task, data, and pipeline parallelism, we exploit all types of parallelism in a unified manner in order to achieve this generality. Our compiler, which maps from the StreamIt language to the 16-core Raw architecture, attains a 11.2x mean speedup over a single-core baseline, and a 1.84x speedup over our previous work. Michael I. Gordon, William Thies, Saman P. Amarasinghe |
ASPLOS | 2 |
| 2006 | Abstraction Layers for Scalable Microfluidic Biocomputers
William Thies, John Paul Urbanski, Todd Thorsen, Saman P. Amarasinghe |
DNA | 1 |
| 2005 | Optimizing stream programs using linear state space analysisabstractDigital Signal Processing (DSP) is becoming increasingly widespread in portable devices. Due to harsh constraints on power, latency, and throughput in embedded environments, developers often appeal to signal processing experts to hand-optimize algorithmic aspects of the application. However, such DSP optimizations are tedious, error-prone, and expensive, as they require sophisticated domain-specific knowledge.We present a general model for automatically representing and optimizing a large class of signal processing applications. The model is based on linear state space systems. A program is viewed as a set of filters, each of which has an input stream, an output stream, and a set of internal states. At each time step, the filter produces some outputs that are a linear combination of the inputs and the state values; the state values are also updated in a linear fashion. Examples of linear state space filters include IIR filters and linear difference equations.Using the state space representation, we describe a novel set of program transformations, including combination of adjacent filters, elimination of redundant states and reduction of the number of system parameters. We have implemented the optimizations in the StreamIt compiler and demonstrate improved generality over previous techniques. Sitij Agrawal, William Thies, Saman P. Amarasinghe |
CASES | 2 |
| 2005 | Static Deadlock Detection for Java Libraries
Amy L. Williams, William Thies, Michael D. Ernst |
ECOOP | 2 |
| 2005 | Cache aware optimization of stream programsabstractEffective use of the memory hierarchy is critical for achieving high performance on embedded systems. We focus on the class of streaming applications, which is increasingly prevalent in the embedded domain. We exploit the widespread parallelism and regular communication patterns in stream programs to formulate a set of cache aware optimizations that automatically improve instruction and data locality. Our work is in the context of the Synchronous Dataflow model, in which a program is described as a graph of independent actors that communicate over channels. The communication rates between actors are known at compile time, allowing the compiler to statically model the caching behavior.We present three cache aware optimizations: 1) execution scaling, which judiciously repeats actor executions to improve instruction locality, 2) cache aware fusion, which combines adjacent actors while respecting instruction cache constraints, and 3) scalar replacement, which converts certain data buffers into a sequence of scalar variables that can be register allocated. The optimizations are founded upon a simple and intuitive model that quantifies the temporal locality for a sequence of actor executions. Our implementation of cache aware optimizations in the StreamIt compiler yields a 249% average speedup (over unoptimized code) for our streaming benchmark suite on a StrongARM 1110 processor. The optimizations also yield a 154% speedup on a Pentium 3 and a 152% speedup on an Itanium 2. Janis Sermulins, William Thies, Rodric M. Rabbah, Saman P. Amarasinghe |
LCTES | 2 |
| 2005 | Teleport messaging for distributed stream programsabstractIn this paper, we develop a new language construct to address one of the pitfalls of parallel programming: precise handling of events across parallel components. The construct, termed teleport messaging, uses data dependences between components to provide a common notion of time in a parallel system. Our work is done in the context of the Synchronous Dataflow (SDF) model, in which computation is expressed as a graph of independent components (or actors) that communicate in regular patterns over data channels. We leverage the static properties of SDF to compute a stream dependence function, SDEP, that compactly describes the ordering constraints between actor executions.Teleport messaging utilizes SDEP to provide powerful and precise event handling. For example, an actor A can specify that an event should be processed by a downstream actor B as soon as B sees the "effects" of the current execution of A. We argue that teleport messaging improves readability and robustness over existing practices. We have implemented messaging as part of the StreamIt compiler, with a backend for a cluster of workstations. As teleport messaging exposes optimization opportunities to the compiler, it also results in a 49% performance improvement for a software radio benchmark. William Thies, Michal Karczmarek, Janis Sermulins, Rodric M. Rabbah, Saman P. Amarasinghe |
PPoPP | 1 |
| 2003 | Phased scheduling of stream programsabstractAs embedded DSP applications become more complex, it is increasingly important to provide high-level stream abstractions that can be compiled without sacrificing efficiency. In this paper, we describe scheduler support for StreamIt, a high-level language for signal processing applications. A StreamIt program consists of a set of autonomous filters that communicate with each other via FIFO queues. As in Synchronous Dataflow (SDF), the input and output rates of each filter are known at compile time. However, unlike SDF, the stream graph is represented using hierarchical structures, each of which has a single input and a single output.We describe a scheduling algorithm that leverages the structure of StreamIt to provide a flexible tradeoff between code size and buffer size. The algorithm describes the execution of each hierarchical unit as a set of phases. A complete cycle through the phases represents a single steady-state execution. By varying the granularity of a phase, our algorithm provides a continuum between single appearance schedules and minimum latency schedules. We demonstrate that a minimal latency schedule is effective in decreasing buffer requirements for some applications, while the phased representation mitigates the associated increase in code size. Michal Karczmarek, William Thies, Saman P. Amarasinghe |
LCTES | 2 |
| 2003 | Linear analysis and optimization of stream programsabstractAs more complex DSP algorithms are realized in practice, there is an increasing need for high-level stream abstractions that can be compiled without sacrificing efficiency. Toward this end, we present a set of aggressive optimizations that target linear sections of a stream program. Our input language is StreamIt, which represents programs as a hierarchical graph of autonomous filters. A filter is linear if each of its outputs can be represented as an affine combination of its inputs. Linearity is common in DSP components; examples include FIR filters, expanders, compressors, FFTs and DCTs.We demonstrate that several algorithmic transformations, traditionally hand-tuned by DSP experts, can be completely automated by the compiler. First, we present a linear extraction analysis that automatically detects linear filters from the C-like code in their work function. Then, we give a procedure for combining adjacent linear filters into a single filter, as well as for translating a linear filter to operate in the frequency domain. We also present an optimization selection algorithm, which finds the sequence of combination and frequency transformations that will give the maximal benefit.We have completed a fully-automatic implementation of the above techniques as part of the StreamIt compiler, and we demonstrate a 450% performance improvement over our benchmark suite. Andrew A. Lamb, William Thies, Saman P. Amarasinghe |
PLDI | 2 |
| 2002 | A stream compiler for communication-exposed architecturesabstractWith the increasing miniaturization of transistors, wire delays are becoming a dominant factor in microprocessor performance. To address this issue, a number of emerging architectures contain replicated processing units with software-exposed communication between one unit and another (e.g., Raw, SmartMemories, TRIPS). However, for their use to be widespread, it will be necessary to develop compiler technology that enables a portable, high-level language to execute efficiently across a range of wire-exposed architectures.In this paper, we describe our compiler for StreamIt: a high-level, architecture-independent language for streaming applications. We focus on our backend for the Raw processor. Though StreamIt exposes the parallelism and communication patterns of stream programs, some analysis is needed to adapt a stream program to a software-exposed processor. We describe a partitioning algorithm that employs fission and fusion transformations to adjust the granularity of a stream graph, a layout algorithm that maps a stream graph to a given network topology, and a scheduling strategy that generates a fine-grained static communication pattern for each computational element.We have implemented a fully functional compiler that parallelizes StreamIt applications for Raw, including several load-balancing transformations. Using the cycle-accurate Raw simulator, we demonstrate that the StreamIt compiler can automatically map a high-level stream abstraction to Raw without losing performance. We consider this work to be a first step towards a portable programming model for communication-exposed architectures. Michael I. Gordon, William Thies, Michal Karczmarek, Jasper Lin, Ali S. Meli, Andrew A. Lamb, Chris Leger, Jeremy Wong, Henry Hoffmann, David Maze, Saman P. Amarasinghe |
ASPLOS | 2 |
| 2002 | StreamIt: A Language for Streaming Applications
William Thies, Michal Karczmarek, Saman P. Amarasinghe |
CC | 1 |
| 2002 | Providing Web search capability for low-connectivity communitiesabstractThere are many technical problems that exist in communities other than our own. These problems both deserve our attention and require focused research. We present one example: the TEK Search Engine. TEK is an email-based search engine designed to deliver low-bandwidth information to low-connectivity communities. Libby Levison, William Thies, Saman P. Amarasinghe |
ISTAS | 2 |
| 2001 | A Unified Framework for Schedule and Storage OptimizationabstractWe present a unified mathematical framework for analyzing the tradeoffs between parallelism and storage allocation within a parallelizing compiler. Using this framework, we show how to find a good storage mapping for a given schedule, a good schedule for a given storage mapping, and a good storage mapping that is valid for all legal schedules. We consider storage mappings that collapse one dimension of a multi-dimensional array, and programs that are in a single assignment form with a one-dimensional schedule. Our technique combines affine scheduling techniques with occupancy vector analysis and incorporates general affine dependences across statements and loop nests. We formulate the constraints imposed by the data dependences and storage mappings as a set of linear inequalities, and apply numerical programming techniques to efficiently solve for the shortest occupancy vector. We consider our method to be a first step towards automating a procedure that finds the optimal tradeoff between parallelism and storage space. William Thies, Frédéric Vivien, Jeffrey Sheldon, Saman P. Amarasinghe |
PLDI | 1 |