Zhen Hao Wong

dblp:359/3607 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-1757-5322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 64% Efficient and distributed learning · 18% Representation and self-supervised learning · 18%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
mathematical reasoning
1.012026
Let's Verify Math Questions Step by Step · KDD (1) 2026
Data integration and cleaning › data preprocessing
data cleaning
1.012026
Let's Verify Math Questions Step by Step · KDD (1) 2026
Machine learning › Efficient and distributed learning
data-efficient learning
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model training
pretraining data selection
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction
0.312026
Let's Verify Math Questions Step by Step · KDD (1) 2026

Methods — techniques the papers use, named apart from their topics

semantic parsing · 2.0large language model · 2.0consistency check · 2.0multi-actor collaboration · 0.9curriculum learning · 0.9
YearPublicationVenuePosition
2026 Let's Verify Math Questions Step by Step
abstract
Large Language Models (LLMs) have recently achieved remarkable progress in mathematical reasoning. To enable such capabilities, many existing works distill strong reasoning models into long chains of thought or design algorithms to construct high-quality math question-answer (QA) data for training. However, these efforts primarily focus on generating correct reasoning paths and answers, while largely overlooking the correctness of the questions themselves. In this work, we present ValiMath, a benchmark consisting of 2147 human-verified mathematical questions covering a wide range of domains such as arithmetic, algebra, and geometry, which are synthesized and curated from the NuminaMath dataset. Each question is annotated with its logical structure, domain coverage, and question correctness, enabling fine-grained evaluation of question quality. ValiMath serves as a high-quality gold-standard test set for validating mathematical questions in LLM training corpora. Building upon this benchmark, we further propose MathQ-Verify, a pipeline that performs fine-grained parsing of mathematical questions into atomic assumptions and conclusions, and evaluates their semantic soundness through consistency checks. This pipeline achieves high precision in detecting flawed questions and provides a reliable foundation for cleaning noisy mathematical datasets. Experiments show that MathQ-Verify achieves state-of-the-art performance across multiple benchmarks, improving the F1 score by up to 25 percentage points over the direct verification baseline. MathQ-Verify offers a scalable and accurate solution for curating reliable mathematical datasets, reducing label noise and avoiding unnecessary computation on invalid questions. Our code and data are available at the repository https://github.com/OpenDCAI/MathQ-Verify.
Chengyu Shen, Zhen Hao Wong, Runming He, Hao Liang 0017, Meiyi Qiang, Zimo Meng, Zhengyang Zhao 0003, Bohan Zeng, Zhengzhou Zhu, Bin Cui 0001, Wentao Zhang 0001
KDD (1)2
2026 Data Preparation for Large Language Models
Hao Liang 0017, Zhen Hao Wong, Rui-Tong Liu, Yu-Han Wang, Meiyi Qiang, Zhengyang Zhao 0003, Chengyu Shen, Conghui He, Wentao Zhang 0001, Bin Cui 0001
J. Comput. Sci. Technol.2
2026 Robust heterogeneous network representation learning by multifaceted curriculum training
Zhen Hao Wong, Hansi Yang, Quanming Yao, Yaqing Wang 0002
Neural Networks1
2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
abstract
Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection.
Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He
ACL (1)3