Andi Bergen

dblp:122/3392 · also Andreas Bergen · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0001-6557-2534ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2025 Integrating Small Language Models with Retrieval-Augmented Generation in Computing Education: Key Takeaways, Setup, and Practical Insights
abstract
Leveraging a Large Language Model (LLM) for personalized learning in computing education is promising, yet cloud-based LLMs pose risks around data security and privacy. To address these concerns, we developed and deployed a locally stored Small Language Model (SLM) utilizing Retrieval-Augmented Generation (RAG) methods to support computing students' learning. Previous work has demonstrated that SLMs can match or surpass popular LLMs (gpt-3.5-turbo and gpt-4-32k) in handling conversational data from a CS1 course. We deployed SLMs with RAG (SLM + RAG) in a large course with more than 250 active students, fielding nearly 2,000 student questions, while evaluating data privacy, scalability, and feasibility of local deployments. This paper provides a comprehensive guide for deploying SLM + RAG systems, detailing model selection, vector database choice, embedding methods, and pipeline frameworks. We share practical insights from our deployment, including scalability concerns, accuracy versus context length trade-offs, guardrails and hallucination reduction, as well as data privacy maintenance. We address the "Impossible Triangle" in RAG systems, which states that achieving high accuracy, short context length, and low time consumption simultaneously is not feasible. Furthermore, our novel RAG framework, Intelligence Concentration (IC), categorizes information into multiple layers of abstraction within Milvus collections mitigating trade-offs and enabling educational assistants to deliver more relevant and personalized responses to students quickly.
Zezhu Yu, Suqing Liu, Paul Denny 0001, Andi Bergen, Michael Liut
SIGCSE (1)4
2024 Can Small Language Models With Retrieval-Augmented Generation Replace Large Language Models When Learning Computer Science?
abstract
Leveraging Large Language Models (LLMs) for personalized learning and support is becoming a promising tool in computing education. AI Assistants can help students with programming, problem-solving, converse with them to clarify course content, explain error messages to help with debugging, and much more. However, using cloud-based LLMs poses risks around data security, privacy, but also control of the overarching system.
Suqing Liu, Zezhu Yu, Feiran Huang, Yousef Bulbulia, Andi Bergen, Michael Liut
ITiCSE (1)5
2023 Embedding and Scaling Writing Instruction Across First- and Second-Year Computer Science Courses
abstract
Writing skills are often considered unimportant by computer science students and were under-emphasized in our curriculum. We describe our experience embedding CS-specific writing instruction at scale in most of our large, core, first- and second-year Computer Science courses, each with 300-800+ students. Our approach is to collaborate with a writing specialist and a community of course instructors, centralize the management of writing teaching assistants, and introduce a variety of relevant genres and contexts to help students develop and apply writing skills. We outline the institutional support and organization crucial to a project of this scale. In addition, we report on a survey collecting student perception of the writing instruction/assessment. We reflect on quantitative and qualitative evidence of success, as well as the challenges that we faced. We believe that many of these challenges will be common across institutions, particularly those with large courses.
Lisa Zhang 0003, Bogdan Simion, Michael Kaler, Amna Liaqat, Daniel Dick, Andi Bergen, Michael Miljanovic, Andrew Petersen 0001
SIGCSE (1)6
2017 Documenting and sharing software knowledge using screencasts
Laura MacLeod, Andi Bergen, Margaret-Anne D. Storey
Empir. Softw. Eng.2
2015 Code, camera, action: how software developers document and share program knowledge using YouTube
abstract
Creating documentation is a challenging task in software engineering and most techniques involve the laborious and sometimes tedious job of writing text. This paper explores an alternative to traditional text-based documentation, the screen-cast, which captures a developer's screen while they narrate how a program or software tool works. We conducted a study to investigate how developers produce and share developer-focused screen casts using the You Tube social platform. First, we identified and analyzed a set of development screen casts to determine how developers have adapted to the medium to meet the demands of development-related documentation needs. We also explored the techniques and strategies used for sharing software knowledge. Second, we interviewed screen cast producers to understand their motivations for creating screen casts, and to uncover the perceived benefits and challenges in producing code-focused videos. Our findings reveal that video is a useful medium for communicating program knowledge between developers, and that developers build their online personas and reputation by sharing videos through social channels.
Laura MacLeod, Margaret-Anne D. Storey, Andi Bergen
ICPC3
2013 PALTask Chat: A Personalized Automated Context Aware Web Resources Listing Tool
abstract
With the constant evolution of the Internet, a repetitive and ordinary task such as searching online resources has become more complex due to the amount of web services and formats available (e.g., video, audio, text or images). In order to obtain resources within a specific domain, a user manually performs several tasks, such as navigating through different web services, filtering according to various criteria and selecting the relevant results. However, the insufficient contextual information of the underlying application hampers the user's experience. In this paper, we propose a tool to improve a user's experience by automating the task of retrieving interesting resources in a multi user setting. We apply this approach to a chat scenario where users are exposed to resources that are of common interests by exploiting the users' personal context information.
Pratik Jain, Andi Bergen, Lorena Castaneda, Hausi A. Müller
SERVICES2