VLDB 2026 Research / reviewers in the wild / expert
Sullam Jeoung
dblp:321/9906
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0008-8403-5441ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 49% Language models and text generation · 31% Reinforcement learning · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 57% Computational social science and digital humanities · 43% | |
| Databases, data mining, and information retrieval
1 paper |
Data models and query languages · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.5 | 2 | 2025 | Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypes · ICLR 2025 StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language Models · EMNLP 2023 |
Machine learning › Reinforcement learning
multi-turn reinforcement learning |
1.0 | 1 | 2026 | SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL · ACL (1) 2026 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
1.0 | 1 | 2026 | SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL · ACL (1) 2026 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypes · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness › bias in language models
political bias in language models |
0.9 | 1 | 2025 | Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypes · ICLR 2025 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization |
0.9 | 1 | 2025 | A Systematic Survey of Automatic Prompt Optimization Techniques · EMNLP 2025 |
Computing education
learning analytics |
0.9 | 1 | 2025 | Semantic Networks Extracted from Students' Think-Aloud Data are Correlated with Students' Learning Performance · EMNLP 2025 |
Machine learning › Trustworthy machine learning › fairness › bias evaluation
stereotype analysis |
0.7 | 1 | 2023 | StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language Models · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.3 | 1 | 2025 | Semantic Networks Extracted from Students' Think-Aloud Data are Correlated with Students' Learning Performance · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model safety
large language model bias |
0.2 | 1 | 2023 | StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language Models · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.0interleaved feedback · 2.0SPN4RE · 1.7LUKE · 1.7LLM-based extraction · 1.7stereotype content model · 1.3reasoning analysis · 1.3representativeness heuristics · 0.9prompt-based mitigation · 0.9automatic prompt optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQLabstractHarper Hua, Zhen Han, Zhengyuan Shen, Meng-Chieh Lee, Sheng Guan, Qi Zhu, Sullam Jeoung, Yueyan Chen, Yunfei Bai, Shuai Wang, Vassilis N. Ioannidis, Huzefa Rangwala. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Harper Hua, Zhengyuan Shen, Meng-Chieh Lee, Sheng Guan, Qi Zhu 0008, Sullam Jeoung, Yueyan Chen, Vassilis N. Ioannidis, Huzefa Rangwala |
ACL (1) | 7 |
| 2025 | A Systematic Survey of Automatic Prompt Optimization TechniquesabstractKiran Ramnath, Kang Zhou, Sheng Guan, Soumya Smruti Mishra, Xuan Qi, Zhengyuan Shen, Shuai Wang, Sangmin Woo, Sullam Jeoung, Yawei Wang, Haozhu Wang, Han Ding, Yuzhe Lu, Zhichao Xu, Yun Zhou, Balasubramaniam Srinivasan, Qiaojing Yan, Yueyan Chen, Haibo Ding, Panpan Xu, Lin Lee Cheong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Kiran Ramnath, Sheng Guan, Soumya Smruti Mishra, Xuan Qi, Zhengyuan Shen, Sangmin Woo, Sullam Jeoung, Haozhu Wang, Han Ding 0004, Yuzhe Lu, Zhichao Xu 0001, Qiaojing Yan, Yueyan Chen, Haibo Ding, Lin Lee Cheong |
EMNLP | 9 |
| 2025 | Semantic Networks Extracted from Students' Think-Aloud Data are Correlated with Students' Learning PerformanceabstractWhen students reflect on their learning from a textbook via think-aloud processes, network representations can be used to capture the concepts and relations from these data.What can we learn from the resulting network representations about students' learning processes, knowledge acquisition, and learning outcomes?This study brings methods from entity and relation extraction using classic and LLM-based methods to the application domain of educational psychology.We built a ground-truth baseline of relational data that represents relevant (to educational science), textbook-based information as a semantic network.Among the tested models, SPN4RE and LUKE achieved the best performance in extracting concepts and relations from students' verbal data.Network representations of students' verbalizations varied in structure, reflecting different learning processes.Correlating the students' semantic networks with learning outcomes revealed that denser and more interconnected semantic networks were associated with more elaborated knowledge acquisition.Structural features such as the number of edges and surface overlap with textbook networks significantly correlated with students' posttest performance. Pingjing Yang, Sullam Jeoung, Jennifer Cromley, Jana Diesner |
EMNLP | 2 |
| 2025 | Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypesabstractExamining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them.
Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses. Sullam Jeoung, Yubin Ge, Haohan Wang, Jana Diesner |
ICLR | 1 |
| 2024 | Extractive Summarization via Fine-grained Semantic Tuple ExtractionabstractTraditional extractive summarization treats the task as sentence-level classification and requires a fixed number of sentences for extraction.However, this rigid constraint on the number of sentences to extract may hinder model generalization due to varied summary lengths across datasets.In this work, we leverage the interrelation between information extraction (IE) and text summarization, and introduce a fine-grained autoregressive method for extractive summarization through semantic tuple extraction.Specifically, we represent each sentence as a set of semantic tuples, where tuples are predicate-argument structures derived from conducting IE.Then we adopt a Transformerbased autoregressive model to extract the tuples corresponding to the target summary given a source document.In inference, a greedy approach is proposed to select source sentences to cover extracted tuples, eliminating the need for a fixed number.Our experiments on CNN/DM and NYT demonstrate the method's superiority over strong baselines.Through the zero-shot setting for testing the generalization of models to diverse summary lengths across datasets, we further show our method outperforms baselines, including ChatGPT. Yubin Ge, Sullam Jeoung, Jana Diesner |
INLG | 2 |
| 2023 | StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsabstractLarge Language Models (LLMs) have been observed to encode and perpetuate harmful associations present in the training data.We propose a theoretically grounded framework called STEREOMAP to gain insights into their perceptions of how demographic groups have been viewed by society.The framework is grounded in the Stereotype Content Model (SCM); a well-established theory from psychology.According to SCM, stereotypes are not all alike.Instead, the dimensions of Warmth and Competence serve as the factors that delineate the nature of stereotypes.Based on the SCM theory, STEREOMAP maps LLMs' perceptions of social groups (defined by sociodemographic features) using the dimensions of Warmth and Competence.Furthermore, the framework enables the investigation of keywords and verbalizations of reasoning of LLMs' judgments to uncover underlying factors influencing their perceptions.Our results show that LLMs exhibit a diverse range of perceptions towards these groups, characterized by mixed evaluations along the dimensions of Warmth and Competence.Furthermore, analyzing the reasonings of LLMs, our findings indicate that LLMs demonstrate an awareness of social disparities, often stating statistical data and research findings to support their reasoning.This study contributes to the understanding of how LLMs perceive and represent social groups, shedding light on their potential biases and the perpetuation of harmful associations. Sullam Jeoung, Yubin Ge, Jana Diesner |
EMNLP | 1 |