Shashank Sonkar

dblp:266/1460 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-4409-1341ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
abstract
Student mistakes in mathematics are often systematic: a learner applies a coherent but wrong procedure and repeats it across contexts.We introduce MALRULELIB, a learningscience-grounded framework that translates documented misconceptions into executable procedures, drawing on 67 learning-science and mathematics education sources, and generates step-by-step traces of malrule-consistent student work.We formalize a core studentmodeling problem as Malrule Reasoning Accuracy (MRA): infer a misconception from one worked mistake and predict the student's next answer under cross-template rephrasing.Across nine language models (4B-120B), accuracy drops from 66% on direct problem solving to 40% on cross-template misconception prediction.MALRULELIB encodes 101 malrules over 498 parameterized problem templates and produces paired dual-path traces for both correct reasoning and malrule-consistent student reasoning.Because malrules are executable and templates are parameterizable, MALRULELIB can generate over one million instances, enabling scalable supervision and controlled evaluation.Using MALRULELIB, we observe cross-template degradations of 10-21%, while providing student step traces improves prediction by 3-15%.We release MAL-RULELIB as infrastructure for educational AI that models student procedures across contexts, enabling diagnosis and feedback that targets the underlying misconception.MALRULELIB code is available here.* Equal contribution.Model CRA MRA Forward MRA Llama-3.3-70B70.4% 34.6% 39.6% Qwen3-80B-Think 70.1% 56.4% 53.9% gpt-oss-120b 65.0% 56.9% 48.8%Average 68.5% 49.3% 47.5% Forward MRA MRA (Cross-Template) System: You are simulating a student who has a specific mathematical misconception.Apply the described misconception consistently to solve the problem.System: You are an expert in identifying and understanding student mathematical misconceptions.Given an example of a student's incorrect answer, identify the systematic error and apply it to predict answers for new problems.User: A student has the following misconception: Students distribute square root over addition: √ a 2 + b 2 = a + b Apply this misconception to solve: Evaluate f (x) = √ x 2 + 4 when x = 3.What is f (3)?User: A student solved this problem incorrectly: Problem: Evaluate f (x) = √ x 2 + 25 when x = 8.Student's Answer: 13 Now predict what this same student would answer for: You walk 8 blocks east and 3 blocks north.What is the straight-line distance from your starting point?Expected: 5 (correct: √ 13 ≈ 3.61) Expected: 11 (correct: √ 73 ≈ 8.54)
Xinghe Chen, Naiming Liu, Shashank Sonkar
ACL (1)3
2026 When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
Robinson Ferrer, Damla Turgut, Zhongzhou Chen, Shashank Sonkar
AIED (1)4
2026 Circuit Complexity of Hierarchical Knowledge Tracing
Naiming Liu, Richard G. Baraniuk, Shashank Sonkar
AIED (3)3
2026 Misconception Acquisition Dynamics in Large Language Models
Naiming Liu, Xinghe Chen, Richard G. Baraniuk, Mrinmaya Sachan, Shashank Sonkar
AIED (1)5
2026 FoundationalASSIST: Dataset for Foundational Knowledge Tracing & Pedagogical Grounding of Large Language Models
Eamon Worden, Cristina Heffernan, Neil T. Heffernan, Shashank Sonkar
AIED4
2025 Do LLMs Make Mistakes Like Students? Exploring Natural Alignments Between Language Models and Human Error Patterns
Naiming Liu, Shashank Sonkar, Richard G. Baraniuk
AIED (4)2
2025 Many-Shot Regurgitation Prompting
Shashank Sonkar, Naiming Liu, Richard G. Baraniuk
AIED (5)1
2025 Turing-Like Test for Personalized Educational AI
Shashank Sonkar, Naiming Liu, Xinghe Chen, Richard G. Baraniuk
AIED (6)1
2025 Atomic Learning Objectives and LLMs Labeling: A High-Resolution Approach for Physics Education
Naiming Liu, Shashank Sonkar, Debshila Basu Mallick, Richard G. Baraniuk, Zhongzhou Chen
LAK2
2024 Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
Shashank Sonkar, Naiming Liu, Debshila Basu Mallick, Richard G. Baraniuk
AIED (1)1
2024 Automated Long Answer Grading with RiceChem Dataset
Shashank Sonkar, Kangqi Ni, Lesa Tran Lu, Kristi Kincaid, John S. Hutchinson, Richard G. Baraniuk
AIED (1)1
2024 Leveraging Large Language Models for Next-Generation Educational Technologies
Neil T. Heffernan, Rose E. Wang, Christopher J. MacLellan, Arto Hellas, Chenglu Li, Candace A. Walkington, Joshua Littenberg-Tobias, David Joyner, Steven Moore, Adish Singla, Zachary A. Pardos, Maciej Pankiewicz, Juho Kim 0001, Shashank Sonkar, Clayton Cohn, Anthony Botelho, Andrew S. Lan, Mingyu Feng, Tanja Käser, Eamon Worden
EDM14
2024 Code Soliloquies for Accurate Calculations in Large Language Models
abstract
High-quality conversational datasets are crucial for the successful development of Intelligent Tutoring Systems (ITS) that utilize a Large Language Model (LLM) backend. Synthetic student-teacher dialogues, generated using advanced GPT-4 models, are a common strategy for creating these datasets. However, subjects like physics that entail complex calculations pose a challenge. While GPT-4 presents impressive language processing capabilities, its limitations in fundamental mathematical reasoning curtail its efficacy for such subjects. To tackle this limitation, we introduce in this paper an innovative stateful prompt design. Our design orchestrates a mock conversation where both student and tutorbot roles are simulated by GPT-4. Each student response triggers an internal monologue, or ‘code soliloquy’ in the GPT-tutorbot, which assesses whether its subsequent response would necessitate calculations. If a calculation is deemed necessary, it scripts the relevant Python code and uses the Python output to construct a response to the student. Our approach notably enhances the quality of synthetic conversation datasets, especially for subjects that are calculation-intensive. The preliminary Subject Matter Expert evaluations reveal that our Higgs model, a fine-tuned LLaMA model, effectively uses Python for computations, which significantly enhances the accuracy and computational reliability of Higgs’ responses.
Shashank Sonkar, Xinghe Chen, Myco Le, Naiming Liu, Debshila Basu Mallick, Richard G. Baraniuk
LAK1
2020 Attention Word Embedding
abstract
Word embedding models learn semantically rich vector representations of words and are widely used to initialize natural processing language (NLP) models.The popular continuous bag-ofwords (CBOW) model of word2vec learns a vector embedding by masking a given word in a sentence and then using the other words as a context to predict it.A limitation of CBOW is that it equally weights the context words when making a prediction, which is inefficient, since some words have higher predictive value than others.We tackle this inefficiency by introducing the Attention Word Embedding (AWE) model, which integrates the attention mechanism into the CBOW model.We also propose AWE-S, which incorporates subword information.We demonstrate that AWE and AWE-S outperform the state-of-the-art word embedding models both on a variety of word similarity datasets and when used for initialization of NLP models.
Shashank Sonkar, Andrew E. Waters, Richard G. Baraniuk
COLING1
2020 qDKT: Question-centric Deep Knowledge Tracing
Shashank Sonkar, Andrew S. Lan, Andrew E. Waters, Phillip Grimaldi, Richard G. Baraniuk
EDM1