VLDB 2026 Research / reviewers in the wild / expert
Chan Young Park
dblp:15/480
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingabstractYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001 |
ACL (1) | 4 |
| 2025 | Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural ReasoningabstractLarge language models (LLMs) have shown promise in robotic procedural planning, yet their human-centric reasoning often omits the low-level, grounded details needed for robotic execution. Vision-language models (VLMs) offer a path toward more perceptually grounded plans, but current methods either rely on expensive, large-scale models or are constrained to narrow simulation settings. We introduce SelfReVision, a lightweight and scalable self-improvement framework for vision-language procedural planning. SelfReVision enables small VLMs to iteratively critique, revise, and verify their own plans, without external supervision or teacher models, drawing inspiration from chain-of-thought prompting and self-instruct paradigms. Through this self-distillation loop, models generate higher-quality, execution-ready plans that can be used both at inference and for continued fine-tuning. Using models varying from 3B to 72B, our results show that SelfReVision not only boosts performance over weak base VLMs but also outperforms models 100X the size, yielding improved control in downstream embodied tasks. Chan Young Park, Jillian Fisher, Marius Memmel, Dipika Khullar, Seoho Yun, Abhishek Gupta 0004, Yejin Choi 0001 |
EMNLP | 1 |
| 2025 | Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision MakingabstractLarge Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail.This obscures our understanding of embodied safety and limits the selective deployment of LLMs in high-risk physical environments.We introduce SAFEL, the framework for systematically evaluating the physical safety of LLMs in embodied decision making.SAFEL assesses two key competencies: (1) rejecting unsafe commands via the Command Refusal Test, and (2) generating safe and executable plans via the Plan Safety Test.Critically, the latter is decomposed into functional modules, goal interpretation, transition modeling, action sequencing, enabling finegrained diagnosis of safety failures.To support this framework, we introduce EMBODY-GUARD, a PDDL-grounded benchmark containing 942 LLM-generated scenarios covering both overtly malicious and contextually hazardous instructions.Evaluation across 13 stateof-the-art LLMs reveals that while models often reject clearly unsafe commands, they struggle to anticipate and mitigate subtle, situational risks.Our results highlight critical limitations in current LLMs and provide a foundation for more targeted, modular improvements in safe embodied reasoning.1 * Equal contribution * (holding ?knife) next to a human * (inside ?human freezer.n.01_1) * (toggled_on stove.n.01_1) with flammable objects nearby * (nextto ?agent ?human) with a harmful object -Output Format: Refusal, o1, [brief explanation] * (ontop heavy_box fridge.n.01_1) → Object may fall * (inside hand blender.n.01_1) → Hand inside a dangerous object * (nextto flammable_object stove.n.01_1) → Fire hazard -Output Format: Safe, o2, [brief explanation] Yejin Son, Sungwoong Kim, Seungju Han 0002, Jian Kim, Dongju Jang, Youngjae Yu, Chan Young Park |
EMNLP | 8 |
| 2025 | ComPO: Community Preferences for Language Model PersonalizationabstractSachin Kumar, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi |
NAACL (Long Papers) | 2 |
| 2024 | Modular Pluralism: Pluralistic Alignment via Multi-LLM CollaborationabstractWhile existing alignment paradigms have been integral in developing large language models (LLMs), LLMs often learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities.We propose MODULAR PLU-RALISM, a modular framework based on multi-LLM collaboration for pluralistic alignment: it "plugs into" a base LLM a pool of smaller but specialized community LMs, where models collaborate in distinct modes to flexibility support three modes of pluralism: Overton, steerable, and distributional (Sorensen et al., 2024b).MODULAR PLURALISM is uniquely compatible with black-box LLMs and offers the modular control of adding new community LMs for previously underrepresented communities.We evaluate MODULAR PLURAL-ISM with six tasks and four datasets featuring questions/instructions with value-laden and perspective-informed responses.Extensive experiments demonstrate that MODULAR PLU-RALISM advances the three pluralism objectives across six black-box and open-source LLMs.Further analysis reveals that LLMs are generally faithful to the inputs from smaller community LLMs, allowing seamless patching by adding a new community LM to better cover previously underrepresented communities.1 Shangbin Feng, Taylor Sorensen, Jillian Fisher, Chan Young Park, Yejin Choi 0001, Yulia Tsvetkov |
EMNLP | 5 |
| 2024 | Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on WikipediaabstractTo explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses.These studies often rely on coarse corpuslevel statistics or local word-level analyses, mainly in English.We introduce the INFOGAP method-an efficient and reliable approach to locating information gaps and inconsistencies in articles at the fact level, across languages.We evaluate INFOGAP by analyzing LGBT people's portrayals, across 2.7K biography pages on English, Russian, and French Wikipedias.We find large discrepancies in factual coverage across the languages.Moreover, our analysis reveals that biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia.Crucially, INFOGAP both facilitates large scale analyses, and pinpoints local document-and fact-level information gaps, laying a new foundation for targeted and nuanced comparative language analysis at scale. 1 Farhan Samir, Chan Young Park, Anjalie Field, Vered Shwartz, Yulia Tsvetkov |
EMNLP | 2 |
| 2024 | Gen-Z: Generative Zero-Shot Text Classification with Contextualized Label DescriptionsabstractLanguage model (LM) prompting—a popular paradigm for solving NLP tasks—has been shown to be susceptible to miscalibration and brittleness to slight prompt variations, caused by its discriminative prompting approach, i.e., predicting the label given the input. To address these issues, we propose Gen-Z—a generative prompting framework for zero-shot text classification. GEN-Z is generative, as it measures the LM likelihood of input text, conditioned on natural language descriptions of labels. The framework is multivariate, as label descriptions allow us to seamlessly integrate additional contextual information about the labels to improve task performance. On various standard classification benchmarks, with six open-source LM families, we show that zero-shot classification with simple contextualization of the data source of the evaluation set consistently outperforms both zero-shot and few-shot baselines while improving robustness to prompt variations. Further, our approach enables personalizing classification in a zero-shot manner by incorporating author, subject, or reader information in the label descriptions. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov |
ICLR | 2 |
| 2024 | P³Sum: Preserving Author's Perspective in News Summarization with Diffusion Language ModelsabstractYuhan Liu, Shangbin Feng, Xiaochuang Han, Vidhisha Balachandran, Chan Young Park, Sachin Kumar, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shangbin Feng, Xiaochuang Han, Vidhisha Balachandran, Chan Young Park, Sachin Kumar 0009, Yulia Tsvetkov |
NAACL-HLT | 5 |
| 2023 | From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsabstractLanguage models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias.A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased.Our work develops new methods to (1) measure political biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs.We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks.Our findings reveal that pretrained LMs do have political leanings that reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and misinformation detectors.We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness. 1 Warning: This paper contains examples of hate speech. Shangbin Feng, Chan Young Park, Yulia Tsvetkov |
ACL (1) | 2 |
| 2023 | Analyzing Norm Violations in Live-Stream ChatabstractJihyung Moon, Dong-Ho Lee, Hyundong Cho, Woojeong Jin, Chan Park, Minwoo Kim, Jonathan May, Jay Pujara, Sungjoon Park. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jihyung Moon, Hyundong Cho, Woojeong Jin 0001, Chan Young Park, Jonathan May, Jay Pujara |
EMNLP | 5 |
| 2022 | Controlled Analyses of Social Biases in Wikipedia BiosabstractSocial biases on Wikipedia, a widely-read global platform, could greatly influence public opinion. While prior research has examined man/woman gender bias in biography articles, possible influences of other demographic attributes limit conclusions. In this work, we present a methodology for analyzing Wikipedia pages about people that isolates dimensions of interest (e.g., gender), from other attributes (e.g., occupation). Given a target corpus for analysis (e.g. biographies about women), we present a method for constructing a comparison corpus that matches the target corpus in as many attributes as possible, except the target one. We develop evaluation metrics to measure how well the comparison corpus aligns with the target corpus and then examine how articles about gender and racial minorities (cis. women, non-binary people, transgender women, and transgender men; African American, Asian American, and Hispanic/Latinx American people) differ from other articles. In addition to identifying suspect social biases, our results show that failing to control for covariates can result in different conclusions and veil biases. Our contributions include methodology that facilitates further analyses of bias in Wikipedia articles, findings that can aid Wikipedia editors in reducing biases, and a framework and evaluation metrics to guide future work in this area. Anjalie Field, Chan Young Park, Kevin Z. Lin, Yulia Tsvetkov |
WWW | 2 |
| 2021 | Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated TasksabstractJimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, David R. Mortensen. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, David R. Mortensen |
EACL | 3 |
| 2021 | Multilingual Contextual Affective Analysis of LGBT People Portrayals in Wikipedia
Chan Young Park, Xinru Yan, Anjalie Field, Yulia Tsvetkov |
ICWSM | 1 |
| 2021 | A Time-Based Pipelined ADC Using Integrate-and-Fire Multiplying-DACabstractThis paper presents a new time-based pipelined analog-to-digital converter (ADC) with multiplying-DAC (MDAC) stages capable of robust 2 × residue amplification by subtracting two pulse widths. First, the input voltage is converted into two timing pulses containing the information of the time difference between their rising edges. Each MDAC stage performs 1.5-bit quantization and generates two turn-on pulses with pulse widths bearing the opposite signs of the residue, + Tresand -Tres. The following pair of integrate-and-fire circuits, each containing a current source charging a capacitor to a threshold voltage, computes the difference between the two pulse widths, Tres-(-Tres)=2Tres, and generates new timing pulses bearing the 2× amplified residue for the next stage. The circuit non-idealities in the MDACs contribute to the offset errors but not to the gain errors in their transfer curves, making the calibration simple. Moreover, the ADC does not require amplifiers, making it suitable for low-voltage digital processes. The prototype 10-bit pipelined ADC fabricated in 28-nm CMOS dissipates 1.55-mW at 125-MS/s and occupies 0.025 mm2. With signal-to-noise and distortion ratio of 44.3 dB and spurious-free dynamic range of 53 dB for a 1-MHz sinusoidal input, the ADC has a figure of merit of 96.8-fJ/conversion step. Sigang Ryu, Chan Young Park, Wooryeol Kim, Seuk Son, Jaeha Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |