VLDB 2026 Research / reviewers in the wild / expert
Kar Way Tan
dblp:71/7126
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-2707-6588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting Doubt in Reflective Learning: A Learning Analytics Study with Large and Small Language ModelsabstractReflective learning enhances understanding, especially when instructors promptly address difficulties raised in student reflections. Automated doubt detection can reduce time for instructors, yet existing classification approaches take substantial time for manual annotation and model training. This paper investigates whether large and small language models (LLMs, SLMs) can automate doubt detection without time-consuming training. Using a dataset of anonymized student reflections, we evaluate zero-shot, few-shot prompting, and multi-step reasoning against prior supervised classification baselines. We show that LLMs (GPT-4o, Claude-4, Gemini-2.5) surpass earlier F1 scores without prompting, while prompting further improves their performance. However, using proprietary LLMs can raise cost and privacy concerns. We also show that selected SLMs (Mistral, Qwen) outperform baselines while addressing these concerns. We extend the analysis with a category-level error study, showing that explicit doubts are detected more reliably, while tentative doubts (softened by cautious language), learning challenge (arising from difficulties in applying concepts), and masked doubts (concealed by positive or polite phrasing) are missed more often. These findings highlight both the promise and limitations of language models for doubt detection and the need to ensure that cautious or polite learners who may not express their doubts explicitly are recognized and supported in learning analytics systems. Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo |
LAK | 2 |
| 2025 | Evaluating ChatGPT to Answer Multi-Modal Exercises in Computer Science EducationabstractThis study investigates ChatGPT-4o's ability to answer multi-modal assessment exercises in computer science (CS) courses. While the use of large language models (LLMs) to answer text-based exercises are extensively researched, their ability to answer exercises involving artifacts of other modalities remains underexplored. To close this gap, we evaluate ChatGPT-4o's answers to 120 multi-modal CS exercises in programming, software design, human-computer interaction, statistical analysis, process analysis, and simulation. The multi-modal artifacts in these exercises include class diagrams, sequence diagrams, user interface images, analytical charts, workflow diagrams and object-flow diagrams. Our comparisons to the expected answers of these exercises show that ChatGPT-4o performs well for exercises with class and sequence diagrams possibly due to the availability of more data for training. The potential for misuse by students highlights these exercises are better suited for closed-book exams or as scaffolding activities. ChatGPT-4o answers better for those multi-modal exercises designed to assess students at the lower levels of Bloom's taxonomy than the higher levels. This discrepancy is possibly due to ChatGPT-4o's lack of understanding underlying design concepts and limited ability to generate new multi-modal artifacts, making exercises requiring higher order of cognitive thinking suitable for take-home assignment. We hope the insights from this study provide a foundation to develop effective multi-modal assessments. Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo, Benjamin Gan |
ITiCSE (1) | 2 |
| 2025 | PromptTutor: Effects of an LLM-Based Chatbot on Learning Outcomes and Motivation in Flipped ClassroomsabstractThis study explores the integration of a Large Language Model (LLM) based chatbot, PromptTutor, into flipped classrooms (FC) for undergraduate Computer Science (CS) education. PromptTutor is designed to provide personalized, immediate feedback to support student learning in FC by incorporating reflective learning and scaffolding strategies. The traditional FC typically lacks this immediate feedback during the pre-class learning phase, risking decreased student motivation according to existing literature. This study examines if students improve in learning outcomes and motivation after using PromptTutor. Through a controlled crossover experiment with 50 students, the study demonstrates statistically significant improvements in students' quiz performance and motivation compared to traditional FC. Our work underscores the potential of LLM-based tools in addressing FC challenges, offering actionable insights for educators and institutional leaders in technology-enhanced learning environments. Eng Lieh Ouh, Adam Ho, Siaw Ling Lo, Kar Way Tan, Feng Lin 0006 |
ITiCSE (1) | 5 |
| 2024 | A Data-Driven Approach for Automated Multi-Site Competitive Facility LocationabstractThis paper addresses the challenge of optimizing large-scale retail expansion in competitive urban environments through a data-driven and automated approach to the Competitive Facility Location (CFL) problem. Traditional CFL methods often face limitations in handling large-scale scenarios, relying on manual pre-selection of candidate sites and imposing restrictions on the number of new locations. Our approach uses Adaptive Large Neighborhood Search (ALNS) enhanced with data enrichment techniques, such as community detection on road networks and population weighting based on mobility data. We developed 2 ALNS variants: Community Geometric Centroid (CGC-ALNS) and Population Weighted Centroid (PWC-ALNS). These methods automate the site selection process, eliminating the need for manual pre-selection and enabling evaluation of a large number of store locations. We benchmarked our approaches against ArcGIS, a widely used commercial software for CFL problems. The results demonstrate notable improvements in performance: CGC-ALNS consistently outperforms ArcGIS with up to a 2% increase in consumer count captured, while PWCALNS achieves even greater gains, with an average increase of 4.6% to 13.1% across various store distribution scenarios. Our key contributions include an automated, data-driven site selection process with no restrictions on the number of new sites, and significant performance improvements over existing commercial solutions. Minghui Tan, Kar Way Tan, Hoong Chuin Lau |
IEEE Big Data | 2 |
| 2023 | Combat COVID-19 at National Level using Risk Stratification with Appropriate InterventionabstractIn the national battle against COVID-19, harnessing population-level big data is imperative, enabling authorities to devise effective care policies, allocate healthcare resources efficiently, and enact targeted interventions. Singapore adopted the Home Recovery Programme (HRP) in September 2021, diverting low-risk COVID-19 patients to home care to ease hospital burdens amid high vaccination rates and mild symptoms. While a patient’s suitability for HRP could be assessed using broad-based criteria, integrating machine learning (ML) model becomes invaluable for identifying high-risk patients prone to severe illness, facilitating early medical assessment. Most prior studies have traditionally depended on clinical and laboratory data, necessitating initial clinic or hospital evaluations. None of these studies incorporated vaccination status, a crucial variable in a well-vaccinated population. This paper proposes a machine learning approach to nationwide risk stratification, offering intervention recommendations by harnessing nationwide datasets. Our best-performing ML model, XGBoost achieves an AUROC of 0.930 utilizing data from multiple data sources including patients’ demographic information, vaccination status and medical history. For broader applicability, we also propose a parsimonious XGBoost model with an AUROC of 0.885 with a selection of five commonly collected variables, namely age, number of vaccine doses taken and number of days since the first, second and booster doses. Importantly, both of our proposed models achieve robust predictive performance without requiring the collection of clinical or laboratory data from patients. We believe that the parsimonious model, leveraging easily attainable data, has the potential for broader adoption across diverse nations, ultimately delivering paramount value to their populations. Xuan Jin, Kar Way Tan |
IEEE Big Data | 2 |
| 2023 | A Big Data Approach to Augmenting the Huff Model with Road Network and Mobility Data for Store Footfall PredictionabstractConventional methodologies for new retail store catchment area and footfall estimation rely on ground surveys which are costly and time-consuming. This study augments existing research in footfall estimation through the innovative integration of mobility data and road network to create population-weighted centroids and delineate residential neighbourhoods via a community detection algorithm. Our findings are then used to enhance Huff Model which is commonly used in site selection and footfall estimation. Our approach demonstrated the vast potential residing within big data where we harness the power of mobility data and road network information, offering a cost-effective and scalable alternative. It obviates the reliance on often outdated census data and government urban planning records, positioning itself as a formidable driver of informed retail strategy. In doing so, our approach is poised to deliver substantial value to the retail industry. Minghui Tan, Kar Way Tan, Hoong Chuin Lau |
IEEE Big Data | 2 |
| 2019 | Do my students understand? Automated identification of doubts from informal reflectionsabstractTraditionally, teaching is usually one directional where the instructor imparts knowledge and there is minimal interaction between learners and instructor. With the focus on learner-centered pedagogy, it can be a challenge to provide timely and relevant guidance to individual learners according to their levels of understanding. One of the options available is to collect reflections from learners after each lesson to extract relevant feedback so that doubts or questions can be addressed in a timely manner. In this paper, we derived an approach to automate the identification of doubts from students’ informal reflections through features analysis, word representation and machine learning. Using reflections as a feedback mechanism and aligning it to the weekly course content can pave the way to a promising approach for learner-centered teaching and personalized learning. Siaw Ling Lo, Kar Way Tan, Eng Lieh Ouh |
ICCE | 2 |
| 2009 | Applying Sanitizable Signature to Web-Service-Enabled Business Processes: Going Beyond Integrity ProtectionabstractThis paper studies the scenario where data in business documents is aggregated by different entities via the use of Web services in streamlined business processes. The documents are transported within the Simple Object Access Protocol (SOAP) messages and travel through multiple intermediary entities, each potentially makes changes to the data in the documents. The WS-security provides integrity protection by allowing portions of a SOAP message to be signed using eXtensible Markup Language (XML) signature scheme. This method however, has not considered the situation where a portion of data may be modified by another entity, therefore a need to allow the originating system to control which intermediary entity is authorized to change which portion of the data. The XML signature scheme also does not provide the final recipient the trust for the intermediary entity that makes the changes. In our paper, we study the security requirements for a streamlined business process, and proposes a novel scheme using sanitizable signature on SOAP messages to complement the XML signature to address not only integrity protection but also control of change as well as establishment of trust for intermediary entities. We show how the proposed scheme can be incorporated into the existing standards and be customizable to achieve flexible use of both the vanilla and sanitizable signatures as required in a business scenario. With the proposed technique, IT systems can be more loosely coupled and reap the benefits of distributed systems, such as delegation of work and encapsulation of business logic. Kar Way Tan, Robert H. Deng |
ICWS | 1 |
| 2009 | Spatial Cloaking Revisited: Distinguishing Information Leakage from Anonymity
Kar Way Tan, Yimin Lin, Kyriakos Mouratidis |
SSTD | 1 |