Jin Peng Zhou

dblp:255/1107 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-8407-1110ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 Education
abstract
Large language models hold great promise for transforming K-12 education, but there is an urgent need for systematic evaluation of their core educational capabilities. Existing benchmarks often overlook educational goal cognition and overemphasize answer accuracy, thereby failing to capture deeper subject-level knowledge ability and problem-solving ability. To address this gap, we introduce K-12EduBench: a benchmark for evaluating LLMs’ subject-level knowledge ability, subject-specific problem-solving ability, and educational goal cognition ability in K-12 education. K-12EduBench comprises four components: (1) a dataset of 2,640 objective and 619 subjective questions across nine subjects, annotated with answers, problem-solving processes, and cognitive-level labels; (2) nine Item Response Theory (IRT) models for estimating subject-level knowledge ability; (3) evaluation methods and metrics for assessing multi-step problem-solving ability; and (4) prompts and scoring rubrics for measuring alignment with target cognitive levels. Experiments on advanced LLMs show that education-optimized models consistently outperform general-purpose ones across all three abilities, while under-scaled models lag substantially. We observe a strong positive correlation between subject-level knowledge ability and subject-specific problem-solving ability. Despite gains in educational goal cognition ability, current models—even those tailored for education—still fall short of real-world instructional needs.
Yuqing Ye, Zhifu Chen, Hengnian Gu, Jin Peng Zhou, Dongdai Zhou
AAAI6
2026 Detecting Out-of-Distribution Objects through Class-Conditioned Inpainting
abstract
Recent object detectors have achieved impressive accuracy in identifying objects seen during training. However, real-world deployment often introduces novel and unexpected objects, referred to as out-of-distribution (OOD) objects, posing significant challenges to model trustworthiness. Modern object detectors are typically overconfident, making it unreliable to use their predictions alone for OOD detection. To address this, we propose leveraging an auxiliary model as a complementary solution. Specifically, we utilize an off-the-shelf text-to-image generative model, such as Stable Diffusion, which is trained with objective functions distinct from those of discriminative object detectors. We hypothesize that this fundamental difference enables the detection of OOD objects by measuring inconsistencies between the models. Concretely, for a given detected object bounding box and its predicted in-distribution class label, we perform class-conditioned inpainting on the image with the object removed. If the object is OOD, the inpainted image is likely to deviate significantly from the original, making the reconstruction error a robust indicator of OOD status. Extensive experiments demonstrate that our approach consistently surpasses existing zero-shot and non-zero-shot OOD detection methods, establishing a robust framework for enhancing object detection systems in dynamic environments. Our implementation is available at https://github.com/quanghuy0497/RONIN.
Jin Peng Zhou, Khanh-Huyen Bui, Kilian Q. Weinberger, Wei-Lun Chao, Dung D. Le
WACV2
2025 Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations
abstract
Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and reduce the cost associated with it.But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both?For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains.We show that providing privileged information -such as ground-truth solutions or problem-specific guidelines -improves automated evaluations on such frontier problems.This approach offers two key advantages.First, it expands the range of problems where LMs graders apply.Specifically, weaker models can now rate the predictions of stronger models.Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low.With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models.LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems.
Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding 0002, Kilian Q. Weinberger, Nan Hua, Fei Sha
EMNLP1
2025 Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
abstract
Large language models (LLMs) should undergo rigorous audits to identify potential risks, such as copyright and privacy infringements. Once these risks emerge, timely updates are crucial to remove undesirable responses, ensuring legal and safe model usage. It has spurred recent research into LLM unlearning, focusing on erasing targeted undesirable knowledge without compromising the integrity of other, non-targeted responses. Existing studies have introduced various unlearning objectives to pursue LLM unlearning without necessitating complete retraining. However, each of these objectives has unique properties, and no unified framework is currently available to comprehend them thoroughly. To fill the gap, we propose the metric of the G-effect, quantifying the impacts of unlearning objectives on model performance from a gradient lens. A significant advantage of our metric is its broad ability to detail the unlearning impacts from various aspects across instances, updating steps, and LLM layers. Accordingly, the G-effect offers new insights into identifying drawbacks of existing unlearning objectives, further motivating us to explore a series of candidate solutions for their mitigation and improvements. Finally, we outline promising directions that merit further studies, aiming at contributing to the community to advance this critical field.
Jin Peng Zhou, Zhanke Zhou, Saebyeol Shin, Bo Han 0003, Kilian Q. Weinberger
ICLR2
2025 On Speeding Up Language Model Evaluation
abstract
Developing prompt-based methods with Large Language Models (LLMs) requires making numerous decisions, which give rise to a combinatorial search problem over hyper-parameters. This exhaustive evaluation can be time-consuming and costly. In this paper, we propose an \textit{adaptive} approach to explore this space. We are exploiting the fact that often only few samples are needed to identify clearly superior or inferior settings, and that many evaluation tests are highly correlated. We lean on multi-armed bandits to sequentially identify the next (method, validation sample)-pair to evaluate and utilize low-rank matrix factorization to fill in missing evaluations. We carefully assess the efficacy of our approach on several competitive benchmark problems and show that it can identify the top-performing method using only 5-15% of the typical resources---resulting in 85-95% LLM cost savings. Our code is available at https://github.com/kilian-group/banditeval.
Jin Peng Zhou, Christian K. Belardi, Ruihan Wu, Travis Zhang, Carla P. Gomes, Wen Sun 0002, Kilian Q. Weinberger
ICLR1
2025 Hierarchical Disentanglement of Cognitive States for Enhanced Cognitive Diagnosis
abstract
With the rapid evolution of multimedia technologies and its widespread integration into education, adaptive multimedia learning has gained significant prominence. Cognitive diagnosis (CD) is pivotal in this domain, as it models students' cognitive states using practice data captured by multimedia learning applications. However, existing methods often simplify these states to mere proficiency on knowledge concepts. Constructivism in education emphasizes learning as a continuous cognitive development process, during which students' cognitive states become increasingly complex, involving not only their construction of concepts but also their construction of relations between concepts that have long been overlooked. To this end, we propose the Hierarchical Disentanglement of Cognitive States for Enhanced Cognitive Diagnosis (HDCD). Inspired by the Structure of Observed Learning Outcomes (SOLO) taxonomy, which categorizes cognitive development into core hierarchical levels (Multistructural, Relational, Extended Abstract), we introduce a hierarchical disentanglement strategy to define cognitive states aligned with each SOLO level: Intra-Concept Cognitive States, Relational Cognitive States, and Extended Cognitive States. Specifically, (i) At the multistructural level, intra-concept cognitive states are sampled from student's personalized cognitive distribution, representing the construction of individual concepts. (ii) At the relational level, inter-concept cognitive states are first sampled to represent the construction of relations between concepts. We then employ a hypergraph transformation to collaboratively update both intra-concept and inter-concept cognitive states, forming relational cognitive states. Considering that students' self-constructed knowledge systems involve multiple types of inter-concept relations, relational cognitive states are implemented under both undirected and directed relation views in this work, and then fed into local diagnostic functions, respectively. (iii) At the extended abstract level, outputs from the local diagnostic functions are fused using multi-view attention mechanisms, resulting in extended cognitive states, which integrate information from multiple relational views, are then fed into a global diagnostic function for final prediction. Extensive experiments on real-world datasets demonstrate the superior performance and interpretability of our HDCD.
Hengnian Gu, Zhifu Chen, Jin Peng Zhou, Dongdai Zhou
ACM Multimedia3
2025 Value-Guided Search for Efficient Chain-of-Thought Reasoning
abstract
In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method does not require a fine-grained notion of ``step,'' which is difficult to define for long-context reasoning models. By collecting a dataset of 2.5 million reasoning traces, we train a 1.5B token-level value model and apply it to DeepSeek models for improved performance with test-time compute scaling. We find that block-wise value-guided search (\texttt{VGS}) with a final weighted majority vote achieves better test-time scaling than standard methods such as majority voting or best-of-$n$. Moreover, \texttt{VGS} significantly reduces the inference FLOPs required to achieve the same performance of majority voting. Our dataset, model and codebase are open-sourced at \codeurl.
Jin Peng Zhou, Jonathan D. Chang, Zhaolin Gao, Nathan Kallus, Kianté Brantley, Wen Sun 0002
NeurIPS2
2025 Q#: Provably Optimal Distributional RL for LLM Post-Training
Jin Peng Zhou, Jonathan D. Chang, Zhaolin Gao, Nathan Kallus, Kilian Q. Weinberger, Kianté Brantley, Wen Sun 0002
NeurIPS1
2024 Magnushammer: A Transformer-Based Approach to Premise Selection
abstract
This paper presents a novel approach to premise selection, a crucial reasoning task in automated theorem proving. Traditionally, symbolic methods that rely on extensive domain knowledge and engineering effort are applied to this task. In contrast, this work demonstrates that contrastive training with the transformer architecture can achieve higher-quality retrieval of relevant premises, without the knowledge or feature engineering overhead. Our method, Magnushammer, outperforms the most advanced and widely used automation tool in interactive theorem proving called Sledgehammer. On the PISA and miniF2f benchmarks Magnushammer achieves $59.5\%$ (against $38.3\%$) and $34.0\%$ (against $20.9\%$) success rates, respectively. By combining Magnushammer with a language-model-based automated theorem prover, we further improve the state-of-the-art proof success rate from $57.0\%$ to $71.0\%$ on the PISA benchmark using $4$x fewer parameters. Moreover, we develop and open source a novel dataset for premise selection, containing textual representations of (proof state, relevant premise) pairs. To the best of our knowledge, this is the largest available premise selection dataset, and the first dataset of this kind for the Isabelle proof assistant.
Maciej Mikula, Szymon Tworkowski, Szymon Antoniak, Bartosz Piotrowski, Albert Q. Jiang, Jin Peng Zhou, Christian Szegedy, Lukasz Kucinski, Piotr Milos, Yuhuai Wu
ICLR6
2024 Don't Trust: Verify - Grounding LLM Quantitative Reasoning with Autoformalization
abstract
Large language models (LLM), such as Google's Minerva and OpenAI's GPT families, are becoming increasingly capable of solving mathematical quantitative reasoning problems. However, they still make unjustified logical and computational errors in their reasoning steps and answers. In this paper, we leverage the fact that if the training corpus of LLMs contained sufficiently many examples of formal mathematics (e.g. in Isabelle, a formal theorem proving environment), they can be prompted to translate i.e. autoformalize informal mathematical statements into formal Isabelle code --- which can be verified automatically for internal consistency. This provides a mechanism to automatically reject solutions whose formalized versions are inconsistent within themselves or with the formalized problem statement. We evaluate our method on GSM8K, MATH and MultiArith datasets and demonstrate that our approach provides a consistently better heuristic than vanilla majority voting --- the previously best method to identify correct answers, by more than 12\% on GSM8K. In our experiments it improves results consistently across all datasets and LLM model sizes. The code can be found at https://github.com/jinpz/dtv.
Jin Peng Zhou, Charles Staats, Christian Szegedy, Kilian Q. Weinberger, Yuhuai Wu
ICLR1
2024 REFACTOR: Learning to Extract Theorems from Proofs
abstract
Human mathematicians are often good at recognizing modular and reusable theorems that make complex mathematical results within reach. In this paper, we propose a novel method called theoREm-from-prooF extrACTOR (REFACTOR) for training neural networks to mimic this ability in formal mathematical theorem proving. We show on a set of unseen proofs, REFACTOR is able to extract 19.6\% of the theorems that humans would use to write the proofs. When applying the model to the existing Metamath library, REFACTOR extracted 16 new theorems. With newly extracted theorems, we show that the existing proofs in the MetaMath database can be refactored. The new theorems are used very frequently after refactoring, with an average usage of 733.5 times, and help shorten the proof lengths. Lastly, we demonstrate that the prover trained on the new-theorem refactored dataset proves more test theorems and outperforms state-of-the-art baselines by frequently leveraging a diverse set of newly extracted theorems. Code can be found at https://github.com/jinpz/refactor.
Jin Peng Zhou, Yuhuai Wu, Qiyang Li, Roger B. Grosse
ICLR1
2023 Does Label Differential Privacy Prevent Label Inference Attacks?
abstract
Label differential privacy (label-DP) is a popular framework for training private ML models on datasets with public features and sensitive private labels. Despite its rigorous privacy guarantee, it has been observed that in practice label-DP does not preclude label inference attacks (LIAs): Models trained with label-DP can be evaluated on the public training features to recover, with high accuracy, the very private labels that it was designed to protect. In this work, we argue that this phenomenon is not paradoxical and that label-DP is designed to limit the advantage of an LIA adversary compared to predicting training labels using the Bayes classifier. At label-DP $\epsilon=0$ this advantage is zero, hence the optimal attack is to predict according to the Bayes classifier and is independent of the training labels. Our bound shows the semantic protection conferred by label-DP and gives guidelines on how to choose $\epsilon$ to limit the threat of LIAs below a certain level. Finally, we empirically demonstrate that our result closely captures the behavior of simulated attacks on both synthetic and real world datasets.
Ruihan Wu, Jin Peng Zhou, Kilian Q. Weinberger, Chuan Guo 0001
AISTATS2
2023 Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs
Albert Q. Jiang, Sean Welleck, Jin Peng Zhou, Timothée Lacroix, Jiacheng Liu 0010, Mateja Jamnik, Guillaume Lample, Yuhuai Wu
ICLR3
2023 Unsupervised Out-of-Distribution Detection with Diffusion Inpainting
abstract
Unsupervised out-of-distribution detection (OOD) seeks to identify out-of-domain data by learning only from unlabeled in-domain data. We present a novel approach for this task – Lift, Map, Detect (LMD) – that leverages recent advancement in diffusion models. Diffusion models are one type of generative models. At their core, they learn an iterative denoising process that gradually maps a noisy image closer to their training manifolds. LMD leverages this intuition for OOD detection. Specifically, LMD lifts an image off its original manifold by corrupting it, and maps it towards the in-domain manifold with a diffusion model. For an OOD image, the mapped image would have a large distance away from its original manifold, and LMD would identify it as OOD accordingly. We show through extensive experiments that LMD achieves competitive performance across a broad variety of datasets. Code can be found at https://github.com/zhenzhel/lift_map_detect.
Jin Peng Zhou, Kilian Q. Weinberger
ICML2
2021 Bayesian Preference Elicitation with Keyphrase-Item Coembeddings for Interactive Recommendation
abstract
Interactive (a.k.a. conversational) recommendation systems provide the potential capability to personalize interactions with increasingly prevalent dialog-based AI assistants. In the conversational recommendation setting, a user often has long-term preferences inferred from previous interactions along with ephemeral session-based preferences that need to be efficiently elicited through minimal interaction. Historically, Bayesian preference elicitation methods have proved effective for (i) leveraging prior information to incrementally estimate uncertainty in user preferences as new information is observed, and for (ii) supporting active elicitation of preference feedback to quickly zero in on the best recommendations in a session. Previous work typically focused on eliciting preferences in the space of items or a small set of attributes; in the dialog-based setting, however, we are faced with the task of eliciting preferences in the space of natural language while using this feedback to determine a user’s preferences in item space. To address this task in the era of modern, latent embedding-based recommender systems, we propose a method for coembedding user-item preferences with keyphrase descriptions (i.e., not explicitly known attributes, but rather subjective judgments mined from user reviews or tags) along with a closed-form Bayesian methodology for incrementally estimating uncertainty in user preferences based on elicited keyphrase feedback. We then combine this framework with well-known preference elicitation techniques that can leverage Bayesian posteriors such as Upper Confidence Bounds, Thompson Sampling, and a variety of other methods. Our empirical evaluation on real-world datasets shows that the proposed query selection strategies effectively update user beliefs, leading to high-quality recommendations with a minimal number of keyphrase queries.
Hojin Yang, Scott Sanner, Ga Wu, Jin Peng Zhou
UMAP4
2020 TAFA: Two-headed Attention Fused Autoencoder for Context-Aware Recommendations
abstract
Collaborative filtering with implicit feedback is a ubiquitous class of recommendation problems where only positive interactions such as purchases or clicks are observed. Autoencoder-based recommendation models have shown strong performance on many implicit feedback benchmarks. However, these models tend to suffer from popularity bias making recommendations less personalized. User-generated reviews contain a rich source of preference information, often with specific details that are important to each user, and can help mitigate the popularity bias. Since not all reviews are equally useful, existing work has been exploring various forms of attention to distill relevant information. In the majority of proposed approaches, representations from implicit feedback and review branches are simply concatenated at the end to generate predictions. This can prevent the model from learning deeper correlations between the two modalities and affect prediction accuracy. To address these problems, we propose a novel Two-headed Attention Fused Autoencoder (TAFA) model that jointly learns representations from user reviews and implicit feedback to make recommendations. We apply early and late modality fusion which allows the model to fully correlate and extract relevant information from both input sources. To further combat popularity bias, we leverage the Noise Contrastive Estimation (NCE) objective to “de-popularize” the fused user representation via a two-headed decoder architecture. Empirically, we show that TAFA outperforms leading baselines on multiple real-world benchmarks. Moreover, by tracing attention weights back to reviews we can provide explanations for the generated recommendations and gain further insights into user preferences. Full code for this work is available here: https://github.com/layer6ai-labs/TAFA.
Jin Peng Zhou, Zhaoyue Cheng, Felipe Pérez, Maksims Volkovs
RecSys1