VLDB 2026 Research / reviewers in the wild / expert
Mozhi Zhang
dblp:232/3225
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without TrainingabstractWe introduce *Domain2Vec*, a novel approach that decomposes any dataset into a linear combination of several *meta-domains*, a new concept designed to capture the key underlying features of datasets.
*Domain2Vec* maintains a vocabulary of meta-domains and uses a classifier to decompose any given dataset into a domain vector that corresponds to a distribution over this vocabulary.
These domain vectors enable the identification of optimal data mixture for language model (LM) pretraining in a training-free manner under the ***D**istribution **A**lignment **A**ssumption* (DA$^{2}$), which suggests that when the data distribution of the training set and the validation set is more aligned, a lower validation loss is achieved.
Moreover, *Domain2Vec* can be seamlessly integrated into previous works to model the relationship between domain vectors and LM performance, greatly enhancing the efficiency and scalability of previous methods.
Extensive experiments demonstrate that *Domain2Vec* helps find the data mixture that enhances downstream task performance with minimal computational overhead.
Specifically, *Domain2Vec* achieves the same validation loss on Pile-CC using only $51.5$\% of the compute required when training on the original mixture of The Pile Dataset.
Under equivalent compute budget, *Domain2Vec* improves downstream performance by an average of $2.83$\%. Mozhi Zhang, Howe Tissue, Xipeng Qiu |
ICML | 1 |
| 2025 | SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and BeyondabstractRecent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplored. This gap is partly due to the challenge of collecting diverse and verifiable reasoning data suitable for RL.
We hypothesize that logical reasoning is critical for developing general reasoning capabilities, as logic forms a fundamental building block of reasoning. In this work, we present SynLogic, a data synthesis framework and dataset that generates diverse logical reasoning data at scale, encompassing 35 diverse logical reasoning tasks. The SynLogic approach enables controlled synthesis of data with adjustable difficulty and quantity. Importantly, all examples can be verified by simple rules, making them ideally suited for RL with verifiable rewards.
In our experiments, we validate the effectiveness of RL training on the SynLogic dataset based on 7B and 32B models. SynLogic leads to state-of-the-art logical reasoning performance among open-source datasets, surpassing DeepSeek-R1-Distill-Qwen-32B by 6 points on BBEH. Furthermore, mixing SynLogic data with mathematical and coding tasks improves the training efficiency of these domains and significantly enhances reasoning generalization. Notably, our mixed training model outperforms DeepSeek-R1-Zero-Qwen-32B across multiple benchmarks.
These findings position SynLogic as a valuable resource for advancing the broader reasoning capabilities of LLMs. We will open-source both the data synthesis pipeline and the SynLogic dataset. Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Yongyi Hu, Yiqi Shi, Shitong Weng, Aili Chen, Shiqi Chen 0002, Mozhi Zhang, Junxian He |
NeurIPS | 11 |
| 2024 | InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model GuidanceabstractPengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Mozhi Zhang, Ke Ren, Botian Jiang, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Pengyu Wang 0006, Linyang Li, Chenkun Tan, Mozhi Zhang, Botian Jiang, Xipeng Qiu |
EMNLP | 6 |
| 2024 | Calibrating the Confidence of Large Language Models by Eliciting FidelityabstractLarge language models optimized with techniques like RLHF have achieved good alignment in being helpful and harmless.However, post-alignment, these language models often exhibit overconfidence, where the expressed confidence does not accurately calibrate with their correctness rate.In this paper, we decompose the language model confidence into the Uncertainty about the question and the Fidelity to the answer generated by language models.Then, we propose a plug-and-play method, UF Calibration, to estimate the confidence of language models.Our method has shown good calibration performance by conducting experiments with 6 RLHF-LMs on four MCQA datasets.Moreover, we propose two novel metrics, IPR and CE, to evaluate the calibration of the model, and we have conducted a detailed discussion on Truly Well-Calibrated Confidence for large language models.Our method could serve as a strong baseline, and we hope that this work will provide some insights into the model confidence calibration. Mozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo, Yaqian Zhou 0001, Xipeng Qiu |
EMNLP | 1 |
| 2022 | UAMNer: uncertainty-aware multimodal named entity recognition in social media posts
Luping Liu, Mozhi Zhang, Linbo Qing, Xiaohai He |
Appl. Intell. | 3 |
| 2022 | Deep dual-domain semi-blind network for compressed image quality enhancement
Jingbo He, Xiaohai He, Mozhi Zhang, Shuhua Xiong, Honggang Chen |
Knowl. Based Syst. | 3 |
| 2022 | Feature separation and double causal comparison loss for visible and infrared person re-identification
Qiang Liu 0021, Xiaohai He, Mozhi Zhang, Qizhi Teng, Bo Li 0074, Linbo Qing |
Knowl. Based Syst. | 3 |
| 2022 | A video compression artifact reduction approach combined with quantization parameters estimation
Xin Shuai, Linbo Qing, Mozhi Zhang, Weiheng Sun, Xiaohai He |
J. Supercomput. | 3 |
| 2021 | A Dataset and Baselines for Multilingual Reply SuggestionabstractMozhi Zhang, Wei Wang, Budhaditya Deb, Guoqing Zheng, Milad Shokouhi, Ahmed Hassan Awadallah. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mozhi Zhang, Wei Wang 0238, Budhaditya Deb, Guoqing Zheng, Milad Shokouhi, Ahmed Awadallah 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
Keyulu Xu, Mozhi Zhang, Simon S. Du, Ken-ichi Kawarabayashi, Stefanie Jegelka |
ICLR | 2 |
| 2021 | Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthabstractGraph Neural Networks (GNNs) have been studied through the lens of expressive power and generalization. However, their optimization properties are less well understood. We take the first step towards analyzing GNN training by studying the gradient dynamics of GNNs. First, we analyze linearized GNNs and prove that despite the non-convexity of training, convergence to a global minimum at a linear rate is guaranteed under mild assumptions that we validate on real-world graphs. Second, we study what may affect the GNNs’ training speed. Our results show that the training of GNNs is implicitly accelerated by skip connections, more depth, and/or a good label distribution. Empirical results confirm that our theoretical results for linearized GNNs align with the training behavior of nonlinear GNNs. Our results provide the first theoretical support for the success of GNNs with skip connections in terms of optimization, and suggest that deep GNNs with skip connections would be promising in practice. Keyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji Kawaguchi |
ICML | 2 |
| 2021 | How does a Neural Network's Architecture Impact its Robustness to Noisy Labels?abstractNoisy labels are inevitable in large real-world datasets. In this work, we explore an area understudied by previous works --- how the network's architecture impacts its robustness to noisy labels. We provide a formal framework connecting the robustness of a network to the alignments between its architecture and target/noise functions. Our framework measures a network's robustness via the predictive power in its representations --- the test performance of a linear model trained on the learned representations using a small set of clean labels. We hypothesize that a network is more robust to noisy labels if its architecture is more aligned with the target function than the noise. To support our hypothesis, we provide both theoretical and empirical evidence across various neural network architectures and different domains. We also find that when the network is well-aligned with the target function, its predictive power in representations could improve upon state-of-the-art (SOTA) noisy-label-training methods in terms of test accuracy and even outperform sophisticated methods that use clean labels. Mozhi Zhang, Keyulu Xu, John Dickerson 0001, Jimmy Ba |
NeurIPS | 2 |
| 2020 | Exploiting Cross-Lingual Subword Similarities in Low-Resource Document ClassificationabstractText classification must sometimes be applied in a low-resource language with no labeled training data. However, training data may be available in a related language. We investigate whether character-level knowledge transfer from a related language helps text classification. We present a cross-lingual document classification framework (caco) that exploits cross-lingual subword similarity by jointly training a character-based embedder and a word-based classifier. The embedder derives vector representations for input words from their written forms, and the classifier makes predictions based on the word vectors. We use a joint character representation for both the source language and the target language, which allows the embedder to generalize knowledge about source language words to target language words with similar forms. We propose a multi-task objective that can further improve the model if additional cross-lingual or monolingual resources are available. Experiments confirm that character-level knowledge transfer is more data-efficient than word-level transfer between related languages. Mozhi Zhang, Yoshinari Fujinuma, Jordan L. Boyd-Graber |
AAAI | 1 |
| 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to DictionariesabstractCross-lingual word embeddings (CLWE) are often evaluated on bilingual lexicon induction (BLI).Recent CLWE methods use linear projections, which underfit the training dictionary, to generalize on BLI.However, underfitting can hinder generalization to other downstream tasks that rely on words from the training dictionary.We address this limitation by retrofitting CLWE to the training dictionary, which pulls training translation pairs closer in the embedding space and overfits the training dictionary.This simple post-processing step often improves accuracy on two downstream tasks, despite lowering BLI test accuracy.We also retrofit to both the training dictionary and a synthetic dictionary induced from CLWE, which sometimes generalizes even better on downstream tasks.Our results confirm the importance of fully exploiting the training dictionary in downstream tasks and explains why BLI is a flawed CLWE evaluation. Mozhi Zhang, Yoshinari Fujinuma, Michael J. Paul, Jordan L. Boyd-Graber |
ACL | 1 |
| 2020 | Interactive Refinement of Cross-Lingual Word EmbeddingsabstractCross-lingual word embeddings transfer knowledge between languages: models trained on high-resource languages can predict in low-resource languages.We introduce CLIME, an interactive system to quickly refine cross-lingual word embeddings for a given classification problem.First, CLIME ranks words by their salience to the downstream task.Then, users mark similarity between keywords and their nearest neighbors in the embedding space.Finally, CLIME updates the embeddings using the annotations.We evaluate CLIME on identifying health-related text in four low-resource languages: Ilocano, Sinhalese, Tigrinya, and Uyghur.Embeddings refined by CLIME capture more nuanced word semantics and have higher test accuracy than the original embeddings.CLIME often improves accuracy faster than an active learning baseline and can be easily combined with active learning to improve results. Michelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater, Jordan L. Boyd-Graber |
EMNLP (1) | 2 |
| 2020 | What Can Neural Networks Reason About?
Keyulu Xu, Mozhi Zhang, Simon S. Du, Ken-ichi Kawarabayashi, Stefanie Jegelka |
ICLR | 3 |
| 2019 | Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative NormalizationabstractCross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings.However, orthogonal mapping only works on language pairs whose embeddings are naturally isomorphic.For nonisomorphic pairs, our method (Iterative Normalization) transforms monolingual embeddings to make orthogonal alignment easier by simultaneously enforcing that (1) individual word vectors are unit length, and (2) each language's average vector is zero.Iterative Normalization consistently improves word translation accuracy of three CLWE methods, with the largest improvement observed on English-Japanese (from 2% to 44% test accuracy). Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka, Jordan L. Boyd-Graber |
ACL (1) | 1 |