Zichao Wang 0001

dblp:188/0340 · also Jack Z. Wang, Jack Zichao Wang · DBLP profile ↗
← Back
28ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-6974-9939ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive
abstract
The opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and public policy. Analyzing how these interconnected systems simultaneously failed to protect public health requires innovative analytic approaches for exploring the vast amounts of data and documents disclosed in the UCSF-JHU Opioid Industry Documents Archive (OIDA). The complexity, multimodal nature, and specialized characteristics of these healthcare-related legal and corporate documents necessitate more advanced methods and models tailored to specific data types and detailed annotations, ensuring the precision and professionalism in the analysis. In this paper, we tackle this challenge by organizing the original dataset according to document attributes and constructing a benchmark with 400k training documents and 10k for testing. From each document, we extract rich multimodal information—including textual content, visual elements, and layout structures—to capture a comprehensive range of features. Using multiple AI models, we then generate a large-scale dataset comprising 360k training QA pairs and 10k testing QA pairs. Building on this foundation, we develop domain-specific multimodal Large Language Models (LLMs) and explore the impact of multimodal inputs on task performance. To further enhance response accuracy, we incorporate historical QA pairs as contextual grounding for answering current queries. Additionally, we incorporate page references within the answers and introduce an importance-based page classifier, further improving the precision and relevance of the information provided. Preliminary results indicate the improvements with our AI assistant in document information extraction and question-answering tasks.
Xuan Shen, Brian Wingenroth, Zichao Wang 0001, Jason Kuen, Wanrong Zhu, Ruiyi Zhang 0002, Lichun Ma, Anqi Liu 0001, Tong Sun 0005, Kevin S. Hawkins, Kate Tasker, G. Caleb Alexander, Jiuxiang Gu
AAAI3
2026 VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
abstract
While vision-language models (VLMs) have demonstrated remarkable performance across various tasks combining textual and visual information, they continue to struggle with fine-grained visual perception tasks that require detailed pixel-level analysis. Effectively eliciting comprehensive reasoning from VLMs on such intricate visual elements remains an open challenge. In this paper, we present VipAct, an agent framework that enhances VLMs by integrating multi-agent collaboration and vision expert models, enabling more precise visual understanding and comprehensive reasoning. VipAct consists of an orchestrator agent, which manages task requirement analysis, planning, and coordination, along with specialized agents that handle specific tasks such as image captioning and vision expert models that provide high-precision perceptual information. This multi-agent approach allows VLMs to better perform fine-grained visual perception tasks by synergizing planning, reasoning, and tool use. We evaluate VipAct on benchmarks featuring a diverse set of visual perception tasks, with experimental results demonstrating significant performance improvements over state-of-the-art baselines across all tasks. Furthermore, comprehensive ablation studies reveal the critical role of multi-agent collaboration in eliciting more detailed System-2 reasoning and highlight the importance of image input for task planning. Additionally, our error analysis identifies patterns of VLMs' inherent limitations in visual perception, providing insights into potential future improvements. VipAct offers a flexible and extensible framework, paving the way for more advanced visual perception systems across various real-world applications.
Zhehao Zhang 0001, Ryan Rossi, Tong Yu 0001, Franck Dernoncourt, Ruiyi Zhang 0002, Jiuxiang Gu, Sungchul Kim, Xiang Chen 0010, Zichao Wang 0001, Nedim Lipka
AAAI9
2026 Interview-Informed Generative Agents for Product Discovery: A Validation Study
abstract
Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants’ responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are grounded in, yet approximate population-level response distributions. These findings highlight both the potential and the limits of LLM simulation in design research. While unsuitable as a substitute for individual-level insights, simulation may provide value for early-stage concept screening and iteration, where distributional accuracy suffices. We discuss implications for integrating simulation responsibly into product development workflows.
Zichao Wang 0001, Alexa F. Siu
CHI1
2025 Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries
abstract
While large language models (LLMs) are increasingly capable of handling longer contexts, recent work has demonstrated that they exhibit the "lost in the middle" phenomenon (Liu et al., 2024) of unevenly attending to different parts of the provided context.This hinders their ability to cover diverse source material in multidocument summarization, as noted in the DI-VERSESUMM benchmark (Huang et al., 2024).In this work, we contend that principled content selection is a simple way to increase source coverage on this task.As opposed to prompting an LLM to perform the summarization in a single step, we explicitly divide the task into three steps-(1) reducing document collections to atomic key points, (2) using determinantal point processes (DPP) to perform select key points that prioritize diverse content, and (3) rewriting to the final summary.By combining prompting steps, for extraction and rewriting, with principled techniques, for content selection, we consistently improve source coverage on the DIVERSESUMM benchmark across various LLMs.Finally, we also show that by incorporating relevance to a provided user intent into the DPP kernel, we can generate personalized summaries that cover relevant source information while retaining coverage.
Vishakh Padmakumar, Zichao Wang 0001, David T. Arbour, Jennifer A. Healey
ACL (1)2
2025 From Selection to Generation: A Survey of LLM-based Active Learning
abstract
Yu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley
ACL (1)25
2025 Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
abstract
Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan K. Reddy, Ming Jin, Lifu Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mohammad Beigi, Ying Shen 0006, Parshin Shojaee, Qifan Wang 0001, Zichao Wang 0001, Chandan K. Reddy, Ming Jin 0002, Lifu Huang
EMNLP5
2024 Titan: Bringing the Deep Image Prior to Implicit Representations
abstract
We study the interpolation capabilities of implicit neural representations (INRs) of images. In principle, INRs promise a number of advantages, such as continuous derivatives and arbitrary sampling, being freed from the restrictions of a raster grid. However, empirically, INRs have been observed to poorly interpolate between the pixels of the fit image; in other words, they do not inherently possess a suitable prior for natural images. In this paper, we propose to address and improve INRs’ interpolation capabilities by explicitly integrating image prior information into the INR architecture via deep decoder, a specific implementation of the deep image prior (DIP). Our method, which we call TITAN, leverages a residual connection from the input which enables integrating the principles of the grid-based DIP into the grid-free INR. Through super-resolution and computed tomography experiments, we demonstrate that our method significantly improves upon classic INRs, thanks to the induced natural image bias. We also find that by constraining the weights to be sparse, image quality and sharpness are enhanced, increasing the Lipschitz constant.
Lorenzo Luzi, Daniel LeJeune, Ali Siahkoohi, Sina Alemohammad, Vishwanath Saragadam, Hossein Babaei, Naiming Liu, Zichao Wang 0001, Richard G. Baraniuk
ICASSP8
2023 Interpretable Math Word Problem Solution Generation via Step-by-step Planning
abstract
Solutions to math word problems (MWPs) with step-by-step explanations are valuable, especially in education, to help students better comprehend problem-solving strategies.Most existing approaches only focus on obtaining the final correct answer.A few recent approaches leverage intermediate solution steps to improve final answer correctness but often cannot generate coherent steps with a clear solution strategy.Contrary to existing work, we focus on improving the correctness and coherence of the intermediate solutions steps.We propose a step-by-step planning approach for intermediate solution generation, which strategically plans the generation of the next solution step based on the MWP and the previous solution steps.Our approach first plans the next step by predicting the necessary math operation needed to proceed, given history steps, then generates the next step, token-by-token, by prompting a language model with the predicted math operation.Experiments on the GSM8K dataset demonstrate that our approach improves the accuracy and interpretability of the solution on both automatic metrics and human evaluation.
Mengxue Zhang, Zichao Wang 0001, Zhichao Yang 0001, Weiqi Feng, Andrew S. Lan
ACL (1)2
2023 Retrieval-based Controllable Molecule Generation
Zichao Wang 0001, Weili Nie, Jarren Zhuoran Qiao, Chaowei Xiao, Richard G. Baraniuk, Anima Anandkumar
ICLR1
2022 Automated Scoring for Reading Comprehension via In-context BERT Tuning
Nigel Fernandez, Aritra Ghosh 0001, Naiming Liu, Zichao Wang 0001, Benoît Choffin, Richard G. Baraniuk, Andrew S. Lan
AIED (1)4
2022 Towards Human-Like Educational Question Generation with Large Language Models
Zichao Wang 0001, Jakob Valdez, Debshila Basu Mallick, Richard G. Baraniuk
AIED (1)1
2022 Open-ended Knowledge Tracing for Computer Science Education
abstract
In education applications, knowledge tracing refers to the problem of estimating students' time-varying concept/skill mastery level from their past responses to questions and predicting their future performance.One key limitation of most existing knowledge tracing methods is that they treat student responses to questions as binary-valued, i.e., whether they are correct or incorrect.Response correctness analysis/prediction ignores important information on student knowledge contained in the exact content of the responses, especially for open-ended questions.In this paper, we conduct the first exploration into open-ended knowledge tracing (OKT) by studying the new task of predicting students' exact open-ended responses to questions.Our work is grounded in the domain of computer science education with programming questions.We develop an initial solution to the OKT problem, a student knowledge-guided code generation approach, that combines program synthesis methods using language models with student knowledge tracing methods.We also conduct a series of quantitative and qualitative experiments on a real-world student code dataset to validate OKT and demonstrate its promise in educational applications.
Naiming Liu, Zichao Wang 0001, Richard G. Baraniuk, Andrew S. Lan
EMNLP2
2022 DeepHull: Fast Convex Hull Approximation in High Dimensions
abstract
Computing or approximating the convex hull of a dataset plays a role in a wide range of applications, including economics, statistics, and physics, to name just a few. However, convex hull computation and approximation is exponentially complex, in terms of both memory and computation, as the ambient space dimension increases. In this paper, we propose DeepHull, a new convex hull approximation algorithm based on convex deep networks (DNs) with continuous piecewise-affine nonlinearities and nonnegative weights. The idea is that binary classification between true data samples and adversarially generated samples with such a DN naturally induces a polytope decision boundary that approximates the true data convex hull. A range of exploratory experiments demonstrates that DeepHull efficiently produces a meaningful convex hull approximation, even in a high-dimensional ambient space.
Randall Balestriero, Zichao Wang 0001, Richard G. Baraniuk
ICASSP2
2021 Educational Question Mining At Scale: Prediction, Analysis and Personalization
abstract
Online education platforms enable teachers to share a large number of educational resources such as questions to form exercises and quizzes for students. With large volumes of available questions, it is important to have an automated way to quantify their properties and intelligently select them for students, enabling effective and personalized learning experiences. In this work, we propose a framework for mining insights from educational questions at scale. We utilize the state-of-the-art Bayesian deep learning method, in particular partial variational auto-encoders (p-VAE), to analyze real students' answers to a large collection of questions. Based on p-VAE, we propose two novel metrics that quantify question quality and difficulty, respectively, and a personalized strategy to adaptively select questions for students. We apply our proposed framework to a real-world dataset with tens of thousands of questions and tens of millions of answers from an online education platform. Our framework not only demonstrates promising results in terms of statistical metrics but also obtains highly consistent results with domain experts' evaluation.
Zichao Wang 0001, Sebastian Tschiatschek, Simon Woodhead 0002, José Miguel Hernández-Lobato, Simon L. Peyton Jones, Richard G. Baraniuk, Cheng Zhang 0005
AAAI1
2021 Towards Blooms Taxonomy Classification Without Labels
Zichao Wang 0001, Kyle Manning, Debshila Basu Mallick, Richard G. Baraniuk
AIED (1)1
2021 Scientific Formula Retrieval via Tree Embeddings
abstract
Exploiting the ever-growing corpus of scientific content calls for new ways and means to effectively organize, search, and retrieve scientific formulae. We propose a new data-driven framework for retrieving similar scientific formulae via learned formula representations based on tree embeddings. FORTE (for FOrmula Representation learning via Tree Embeddings) leverages operator tree representations of symbolic scientific formulae (such as math equations) to explicitly capture their inherent structural and semantic properties. FORTE employs i) a tree encoder that encodes the formula’s operator tree into an embedding vector and ii) a tree decoder that directly generates a formula’s operator tree from the embedding vector. We also develop a novel tree beam search algorithm that improves the quality of the decoded operator trees. We demonstrate that FORTE (sometimes significantly) outperforms various baseline methods on formula reconstruction and retrieval using a real-world dataset comprising 770k scientific formulae collected on-line.
Zichao Wang 0001, Mengxue Zhang, Richard G. Baraniuk, Andrew S. Lan
IEEE BigData1
2021 Math Operation Embeddings for Open-ended Solution Analysis and Feedback
Mengxue Zhang, Zichao Wang 0001, Richard G. Baraniuk, Andrew S. Lan
EDM2
2021 Math Word Problem Generation with Mathematical Consistency and Problem Context Constraints
abstract
We study the problem of generating arithmetic math word problems (MWPs) given a math equation that specifies the mathematical computation and a context that specifies the problem scenario.Existing approaches are prone to generating MWPs that are either mathematically invalid or have unsatisfactory language quality.They also either ignore the context or require manual specification of a problem template, which compromises the diversity of the generated MWPs.In this paper, we develop a novel MWP generation approach that leverages i) pre-trained language models and a context keyword selection model to improve the language quality of the generated MWPs and ii) an equation consistency constraint for math equations to improve the mathematical validity of the generated MWPs.Extensive quantitative and qualitative experiments on three realworld MWP datasets demonstrate the superior performance of our approach compared to various baselines.
Zichao Wang 0001, Andrew S. Lan, Richard G. Baraniuk
EMNLP (1)1
2021 Wearing A Mask: Compressed Representations of Variable-Length Sequences Using Recurrent Neural Tangent Kernels
abstract
High dimensionality poses many challenges to the use of data, from visualization and interpretation, to prediction and storage for historical preservation. Techniques abound to reduce the dimensionality of fixed-length sequences, yet these methods rarely generalize to variable-length sequences. To address this gap, we extend existing methods that rely on the use of kernels to variable-length sequences via use of the Recurrent Neural Tangent Kernel (RNTK). Since a deep neural network with ReLu activation is a Max-Affine Spline Operator (MASO), we dub our approach Max-Affine Spline Kernel (MASK). We demonstrate how MASK can be used to extend principal components analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE) and apply these new algorithms to separate synthetic time series data sampled from second-order differential equations.
Sina Alemohammad, Hossein Babaei, Randall Balestriero, Matt Y. Cheung, Ahmed Imtiaz Humayun, Daniel LeJeune, Naiming Liu, Lorenzo Luzi, Jasper Tan, Zichao Wang 0001, Richard G. Baraniuk
ICASSP10
2021 The Recurrent Neural Tangent Kernel
Sina Alemohammad, Zichao Wang 0001, Randall Balestriero, Richard G. Baraniuk
ICLR2
2020 VarFA: A Variational Factor Analysis Framework For Efficient Bayesian Learning Analytics
Zichao Wang 0001, Andrew S. Lan, Richard G. Baraniuk
EDM1
2020 Maximizing Welfare with Incentive-Aware Evaluation Mechanisms
abstract
Motivated by applications such as college admission and insurance rate determination, we study a classification problem where the inputs are controlled by strategic individuals who can modify their features at a cost. A learner can only partially observe the features, and aims to classify individuals with respect to a quality score. The goal is to design a classification mechanism that maximizes the overall quality score in the population, taking any strategic updating into account. When scores are linear and mechanisms can assign their own scores to agents, we show that the optimal classifier is an appropriate projection of the quality score. For the more restrictive task of binary classification via linear thresholds, we construct a (1/4)-approximation to the optimal classifier when the underlying feature distribution is sufficiently smooth and admits an oracle for finding dense regions. We extend our results to settings where the prior distribution is unknown and must be learned from samples.
Nika Haghtalab, Nicole Immorlica, Brendan Lucier, Zichao Wang 0001
IJCAI4
2020 Optimal Single-Choice Prophet Inequalities from Samples
abstract
We study the single-choice Prophet Inequality problem when the gambler is given access to samples. We show that the optimal competitive ratio of $1/2$ can be achieved with a single sample from each distribution. When the distributions are identical, we show that for any constant $\varepsilon > 0$, $O(n)$ samples from the distribution suffice to achieve the optimal competitive ratio ($\approx 0.745$) within $(1+\varepsilon)$, resolving an open problem of Correa, Dütting, Fischer, and Schewior.
Aviad Rubinstein, Zichao Wang 0001, S. Matthew Weinberg
ITCS2
2019 Techniques for Automatically Evaluating Machine-Generated Questions
Zichao Wang 0001
EDM1
2019 A Meta-Learning Augmented Bidirectional Transformer Model for Automatic Short Answer Grading
Zichao Wang 0001, Andrew S. Lan, Andrew E. Waters, Phillip Grimaldi, Richard G. Baraniuk
EDM1
2019 A Max-Affine Spline Perspective of Recurrent Neural Networks
Zichao Wang 0001, Randall Balestriero, Richard G. Baraniuk
ICLR (Poster)1
2018 QG-net: a data-driven question generation model for educational content
abstract
The ever growing amount of educational content renders it increasingly difficult to manually generate sufficient practice or quiz questions to accompany it. This paper introduces QG-Net, a recurrent neural network-based model specifically designed for automatically generating quiz questions from educational content such as textbooks. QG-Net, when trained on a publicly available, general-purpose question/answer dataset and without further fine-tuning, is capable of generating high quality questions from textbooks, where the content is significantly different from the training data. Indeed, QG-Net outperforms state-of-the-art neural network-based and rules-based systems for question generation, both when evaluated using standard benchmark datasets and when using human evaluators. QG-Net also scales favorably to applications with large amounts of educational content, since its performance improves with the amount of training data.
Zichao Wang 0001, Andrew S. Lan, Weili Nie, Andrew E. Waters, Phillip Grimaldi, Richard G. Baraniuk
L@S1
2017 A Latent Factor Model For Instructor Content Preference Analysis
Zichao Wang 0001, Andrew S. Lan, Phillip Grimaldi, Richard G. Baraniuk
EDM1