Andrew Tran

dblp:97/7560 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-0094-1113ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hacking Student Leadership: Peer Mentorship and Leadership Skill Development Among Hackathon Organizers
abstract
While hackathons are often celebrated for their impact on participants, less attention has been given to the unique leadership development opportunities for the student organizers who create and run these events. Unlike traditional classroom settings, where leadership and collaboration skills are typically delayed until upper-level courses, hackathon organizers must tackle these challenges earlier on. In managing a large-scale, formative event, student organizers take on roles that require decision-making, teamwork, and project management. This poster explores the experiences of student organizers at Anonymous Hackathon, emphasizing how this informal learning opportunity complements gaps in the traditional computer science (CS) curriculum by fostering essential leadership and collaboration skills earlier in students' academic careers.
Kush Patel, Andrew Tran, Christine Kapp, Daniel Bicalho, Yatri Patel, Chiku Okechukwu, Egi Rama, Stephen MacNeil
SIGCSE (2)2
2024 MolLM: a unified language model for integrating biomedical text with 2D and 3D molecular representations
abstract
MOTIVATION: The current paradigm of deep learning models for the joint representation of molecules and text primarily relies on 1D or 2D molecular formats, neglecting significant 3D structural information that offers valuable physical insight. This narrow focus inhibits the models' versatility and adaptability across a wide range of modalities. Conversely, the limited research focusing on explicit 3D representation tends to overlook textual data within the biomedical domain. RESULTS: We present a unified pre-trained language model, MolLM, that concurrently captures 2D and 3D molecular information alongside biomedical text. MolLM consists of a text Transformer encoder and a molecular Transformer encoder, designed to encode both 2D and 3D molecular structures. To support MolLM's self-supervised pre-training, we constructed 160K molecule-text pairings. Employing contrastive learning as a supervisory signal for learning, MolLM demonstrates robust molecular representation capabilities across four downstream tasks, including cross-modal molecule and text matching, property prediction, captioning, and text-prompted molecular editing. Through ablation, we demonstrate that the inclusion of explicit 3D representations improves performance in these downstream tasks. AVAILABILITY AND IMPLEMENTATION: Our code, data, pre-trained model weights, and examples of using our model are all available at https://github.com/gersteinlab/MolLM. In particular, we provide Jupyter Notebooks offering step-by-step guidance on how to use MolLM to extract embeddings for both molecules and text.
Xiangru Tang, Andrew Tran, Jeffrey Tan, Mark Gerstein
Bioinform.2
2023 Generating Multiple Choice Questions for Computing Courses Using Large Language Models
abstract
Generating high-quality multiple-choice questions (MCQs) is a time-consuming activity that has led practitioners and researchers to develop community question banks and reuse the same questions from semester to semester. This results in generic MCQs which are not relevant to every course. Template-based methods for generating MCQs require less effort but are similarly limited. At the same time, advances in natural language processing have resulted in large language models (LLMs) that are capable of doing tasks previously reserved for people, such as generating code, code explanations, and programming assignments. In this paper, we investigate whether these generative capabilities of LLMs can be used to craft high-quality M CQs more efficiently, thereby enabling instructors to focus on personalizing MCQs to each course and the associated learning goals. We used two LLMs, GPT-3 and GPT-4, to generate isomorphic MCQs based on MCQs from the Canterbury Question Bank and an Introductory to Low-level C Programming Course. We evaluated the resulting MCQs to assess their ability to generate correct answers based on the question stem, a task that was previously not possible. Finally, we investigate whether there is a correlation between model performance and the discrimination score of the associated MCQ to understand whether low discrimination questions required the model to do more inference and therefore perform poorly. GPT-4 correctly generated the answer for 78.5% of MCQs based only on the question stem. This suggests that instructors could use these models to quickly draft quizzes, such as during a live class, to identify misconceptions in real-time. We also replicate previous findings that GPT-3 performs poorly on answering, or in our case generating, correct answers to MCQs. We also present cases we observed where LLMs struggled to produce correct answers. Finally, we discuss implications for computing education.
Andrew Tran, Kenneth Angelikas, Egi Rama, Chiku Okechukwu, David H. Smith IV, Stephen MacNeil
FIE1
2023 Using Large Language Models to Automatically Identify Programming Concepts in Code Snippets
abstract
Curating course material that aligns with students’ learning goals is a challenging and time-consuming task that instructors undergo when preparing their curricula. For instance, it is a challenge to find multiple-choice questions or example codes that demonstrate recursion in an unlabeled question bank or repository. Recently, Large Language Models (LLMs) have demonstrated the capability to generate high-quality learning materials at scale. In this poster, we use LLMs to identify programming concepts found within code snippets, allowing instructors to quickly curate their course materials. We compare programming concepts generated by LLMs with concepts generated by experts to see the extent to which they agree. The agreement was calculated using Cohen’s Kappa.
Andrew Tran, Egi Rama, Kenneth Angelikas, Stephen MacNeil
ICER (2)1
2023 Comparing Code Explanations Created by Students and Large Language Models
abstract
Reasoning about code and explaining its purpose are fundamental skills for computer scientists. There has been extensive research in the field of computing education on the relationship between a student's ability to explain code and other skills such as writing and tracing code. In particular, the ability to describe at a high-level of abstraction how code will behave over all possible inputs correlates strongly with code writing skills. However, developing the expertise to comprehend and explain code accurately and succinctly is a challenge for many students. Existing pedagogical approaches that scaffold the ability to explain code, such as producing exemplar code explanations on demand, do not currently scale well to large classrooms. The recent emergence of powerful large language models (LLMs) may offer a solution. In this paper, we explore the potential of LLMs in generating explanations that can serve as examples to scaffold students' ability to understand and explain code. To evaluate LLM-created explanations, we compare them with explanations created by students in a large course (n ≈ 1000) with respect to accuracy, understandability and length. We find that LLM-created explanations, which can be produced automatically on demand, are rated as being significantly easier to understand and more accurate summaries of code than student-created explanations. We discuss the significance of this finding, and suggest how such models can be incorporated into introductory programming education.
Juho Leinonen 0001, Paul Denny 0001, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, Arto Hellas
ITiCSE (1)7
2023 The Implications of Large Language Models for CS Teachers and Students
abstract
The introduction of Large Language Models (LLMs) has generated a significant amount of excitement both in industry and among researchers. Recently, tools that leverage LLMs have made their way into the classroom where they help students generate code and help instructors generate learning materials. There are likely many more uses of these tools -- both beneficial to learning and possibly detrimental to learning. To help ensure that these tools are used to enhance learning, educators need to not only be familiar with these tools, but with their use and potential misuse. The goal of this BoF is to raise awareness about LLMs and to build a learning community around their use in computing education. Aligned with this goal of building an inclusive learning community, our BoF is led by globally distributed discussion leaders, including undergraduate researchers, to facilitate multiple coordinated discussions that can lead to a broader conversation about the role of LLMs in CS education.
Stephen MacNeil, Joanne Kim, Juho Leinonen 0001, Paul Denny 0001, Seth Bernstein, Brett A. Becker, Michel Wermelinger, Arto Hellas, Andrew Tran, Sami Sarsa, James Prather, Viraj Kumar
SIGCSE (2)9
2023 Automatically Generating CS Learning Materials with Large Language Models
abstract
Recent breakthroughs in Large Language Models (LLMs), such as GPT-3 and Codex, now enable software developers to generate code based on a natural language prompt. Within computer science education, researchers are exploring the potential for LLMs to generate code explanations and programming assignments using carefully crafted prompts. These advances may enable students to interact with code in new ways while helping instructors scale their learning materials. However, LLMs also introduce new implications for academic integrity, curriculum design, and software engineering careers. This workshop will demonstrate the capabilities of LLMs to help attendees evaluate whether and how LLMs might be integrated into their pedagogy and research. We will also engage attendees in brainstorming to consider how LLMs will impact our field.
Stephen MacNeil, Andrew Tran, Juho Leinonen 0001, Paul Denny 0001, Joanne Kim, Arto Hellas, Seth Bernstein, Sami Sarsa
SIGCSE (2)2
2023 Experiences from Using Code Explanations Generated by Large Language Models in a Web Software Development E-Book
abstract
Advances in natural language processing have resulted in large language models (LLMs) that can generate code and code explanations. In this paper, we report on our experiences generating multiple code explanation types using LLMs and integrating them into an interactive e-book on web software development. Three different types of explanations -- a line-by-line explanation, a list of important concepts, and a high-level summary of the code -- were created. Students could view explanations by clicking a button next to code snippets, which showed the explanation and asked about its utility. Our results show that all explanation types were viewed by students and that the majority of students perceived the code explanations as helpful to them. However, student engagement varied by code snippet complexity, explanation type, and code snippet length. Drawing on our experiences, we discuss future directions for integrating explanations generated by LLMs into CS classrooms.
Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny 0001, Seth Bernstein, Juho Leinonen 0001
SIGCSE (1)2
2022 Generating Diverse Code Explanations using the GPT-3 Large Language Model
abstract
Good explanations are essential to efficiently learning introductory programming concepts [10]. To provide high-quality explanations at scale, numerous systems automate the process by tracing the execution of code [8, 12], defining terms [9], giving hints [16], and providing error-specific feedback [10, 16]. However, these approaches often require manual effort to configure and only explain a single aspect of a given code segment. Large language models (LLMs) are also changing how students interact with code [7]. For example, Github's Copilot can generate code for programmers [4], leading researchers to raise concerns about cheating [7]. Instead, our work focuses on LLMs' potential to support learning by explaining numerous aspects of a given code snippet. This poster features a systematic analysis of the diverse natural language explanations that GPT-3 can generate automatically for a given code snippet. We present a subset of three use cases from our evolving design space of AI Explanations of Code.
Stephen MacNeil, Andrew Tran, Dan Mogil, Seth Bernstein, Erin Ross, Ziheng Huang 0002
ICER (2)2
2020 Extend: A Framework for Increasing Energy Access by Interconnecting Solar Home Systems
abstract
The means of electrifying households and the resulting electricity networks are rapidly evolving. Traditionally, an extension of existing centralized grids was the only prominent technique, but now electrification is seeing massive expansion via decentralized solar home systems (SHSs). These systems consist of a low-wattage photovoltaic (PV) panel (typically 5-100W), a battery, a collection of energy-efficient DC appliances, and a charge controller. Spurred by significant advances and reduced costs in solar, batteries, energyefficient appliances, and mobile money-driven business models, SHSs have proliferated rapidly, with tens of millions of systems now deployed, primarily in regions with otherwise low rates of electricity access.
Santiago Correa, Noman Bashir, Andrew Tran, David Irwin 0001, Jay Taneja
COMPASS3
2009 Scalable onion routing with torsk
abstract
We introduce Torsk, a structured peer-to-peer low-latency anonymity protocol. Torsk is designed as an interoperable replacement for the relay selection and directory service of the popular Tor anonymity network, that decreases the bandwidth cost of relay selection and maintenance from quadratic to quasilinear while introducing no new attacks on the anonymity provided by Tor, and no additional delay to connections made via Tor. The resulting bandwidth savings make a modest-sized Torsk network significantly cheaper to operate, and allows low-bandwidth clients to join the network. Unlike previous proposals for P2P anonymity schemes, Torsk does not require all users to relay traffic for others. Torsk utilizes a combination of two P2P lookup mechanisms with complementary strengths in order to avoid attacks on the confidentiality and integrity of lookups. We show by analysis that previously known attacks on P2P anonymity schemes do not apply to Torsk, and report on experiments conducted with a 336-node wide-area deployment of Torsk, demonstrating its efficiency and feasibility. Categories and Subject Descriptors
Jon McLachlan, Andrew Tran, Nicholas Hopper, Yongdae Kim
CCS2