EDBT 2026 Demo / reviewers in the wild / expert
Simiao Zuo
dblp:232/2089
· DBLP profile ↗
18ranked-venue papers
5as first author
16since 2021 · last 2025
0009-0002-8014-3150ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Consistent Natural-Language Explanations via Explanation-Consistency FinetuningabstractLarge language models (LLMs) often generate convincing, fluent explanations. However, different from humans, they often generate inconsistent explanations on different inputs. For example, an LLM may explain “all birds can fly” when answering the question “Can sparrows fly?” but meanwhile answer “no” to the related question “Can penguins fly?”. Explanations should be consistent across related examples so that they allow humans to simulate the LLM’s decision process on multiple examples. We propose explanation-consistency finetuning (EC-finetuning), a method that adapts LLMs to generate more consistent natural-language explanations on related examples. EC-finetuning involves finetuning LLMs on synthetic data that is carefully constructed to contain consistent explanations. Across a variety of question-answering datasets in various domains, EC-finetuning yields a 10.0% relative explanation consistency improvement on 4 finetuning datasets, and generalizes to 7 out-of-distribution datasets not seen during finetuning (+4.5% relative). We will make our code available for reproducibility. Yanda Chen, Chandan Singh, Xiaodong Liu 0003, Simiao Zuo, Bin Yu 0001, He He 0001, Jianfeng Gao 0001 |
COLING | 4 |
| 2024 | Task Oriented In-Domain Data AugmentationabstractLarge Language Models (LLMs) have shown superior performance in various applications and fields.To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on in-domain data.However, existing approaches suffer from two major issues.First, in-domain data are scarce compared with general domainagnostic data.Second, data used for continual pre-training are not task-aware, such that they may not be helpful to downstream applications.We propose TRAIT, a task-oriented in-domain data augmentation framework.Our framework is divided into two parts: in-domain data selection and task-oriented synthetic passage generation.The data selection strategy identifies and selects a large amount of in-domain data from general corpora, and thus significantly enriches domain knowledge in the continual pre-training data.The synthetic passages contain guidance on how to use domain knowledge to answer questions about downstream tasks.By training on such passages, the model aligns with the need of downstream applications.We adapt LLMs to two domains: advertisement and math.On average, TRAIT improves LLM performance by 8% in the advertisement domain and 7.5% in the math domain. Simiao Zuo, Yeyun Gong, Qiang Lou, Yi Liu 0071, Shao-Lun Huang, Jian Jiao 0007 |
EMNLP | 3 |
| 2024 | Evoke: Evoking Critical Thinking Abilities in LLMs via Reviewer-Author Prompt EditingabstractLarge language models (LLMs) have made impressive progress in natural language processing. These models rely on proper human instructions (or prompts) to generate suitable responses. However, the potential of LLMs are not fully harnessed by commonly-used prompting methods: many human-in-the-loop algorithms employ ad-hoc procedures for prompt selection; while auto prompt generation approaches are essentially searching all possible prompts randomly and inefficiently. We propose Evoke, an automatic prompt refinement framework. In Evoke, there are two instances of a same LLM: one as a reviewer (LLM-Reviewer), it scores the current prompt; the other as an author (LLM-Author), it edits the prompt by considering the edit history and the reviewer's feedback. Such an author-reviewer feedback loop ensures that the prompt is refined in each iteration. We further aggregate a data selection approach to Evoke, where only the hard samples are exposed to the LLM. The hard samples are more important because the LLM can develop deeper understanding of the tasks out of them, while the model may already know how to solve the easier cases. Experimental results show that Evoke significantly outperforms existing methods. For instance, in the challenging task of logical fallacy detection, Evoke scores above 80, while all other baseline methods struggle to reach 20. Simiao Zuo, Qiang Lou, Jian Jiao 0007, Denis Charles |
ICLR | 3 |
| 2023 | DeepTagger: Knowledge Enhanced Named Entity Recognition for Web-Based Ads QueriesabstractNamed entity recognition (NER) is a crucial task for online advertisement. State-of-the-art solutions leverage pre-trained language models for this task. However, three major challenges remain unresolved: web queries differ from natural language, on which pre-trained models are trained; web queries are short and lack contextual information; and labeled data for NER is scarce. We propose DeepTagger, a knowledge-enhanced NER model for web-based ads queries. The proposed knowledge enhancement framework leverages both model-free and model-based approaches. For model-free enhancement, we collect unlabeled web queries to augment domain knowledge; and we collect web search results to enrich the information of ads queries. We further leverage effective prompting methods to automatically generate labels using large language models such as ChatGPT. Additionally, we adopt a model-based knowledge enhancement method based on adversarial data augmentation. We employ a three-stage training framework to train DeepTagger models. Simiao Zuo, Qiang Lou, Jian Jiao 0007, Denis Charles |
CIKM | 1 |
| 2023 | Machine Learning Force Fields with Data Cost Aware TrainingabstractMachine learning force fields (MLFF) have been proposed to accelerate molecular dynamics (MD) simulation, which finds widespread applications in chemistry and biomedical research. Even for the most data-efficient MLFFs, reaching chemical accuracy can require hundreds of frames of force and energy labels generated by expensive quantum mechanical algorithms, which may scale as $O(n^3)$ to $O(n^7)$, with $n$ proportional to the number of basis functions. To address this issue, we propose a multi-stage computational framework -- ASTEROID, which lowers the data cost of MLFFs by leveraging a combination of cheap inaccurate data and expensive accurate data. The motivation behind ASTEROID is that inaccurate data, though incurring large bias, can help capture the sophisticated structures of the underlying force field. Therefore, we first train a MLFF model on a large amount of inaccurate training data, employing a bias-aware loss function to prevent the model from overfitting the potential bias of this data. We then fine-tune the obtained model using a small amount of accurate training data, which preserves the knowledge learned from the inaccurate training data while significantly improving the model's accuracy. Moreover, we propose a variant of ASTEROID based on score matching for the setting where the inaccurate training data are unlabeled. Extensive experiments on MD datasets and downstream tasks validate the efficacy of ASTEROID. Our code and data are available at https://github.com/abukharin3/asteroid. Alexander Bukharin, Shengjie Wang 0001, Simiao Zuo, Weihao Gao, Tuo Zhao |
ICML | 4 |
| 2023 | SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes ProcessabstractTransformer Hawkes process models have shown to be successful in modeling event sequence data. However, most of the existing training methods rely on maximizing the likelihood of event sequences, which involves calculating some intractable integral. Moreover, the existing methods fail to provide uncertainty quantification for model predictions, e.g., confidence interval for the predicted event’s arrival time. To address these issues, we propose SMURF-THP, a score-based method for learning Transformer Hawkes process and quantifying prediction uncertainty. Specifically, SMURF-THP learns the score function of the event’s arrival time based on a score-matching objective that avoids the intractable computation. With such a learnt score function, we can sample arrival time of events from the predictive distribution. This naturally allows for the quantification of uncertainty by computing confidence intervals over the generated samples. We conduct extensive experiments in both event type prediction and uncertainty quantification on time of arrival. In all the experiments, SMURF-THP outperforms existing likelihood-based methods in confidence calibration while exhibiting comparable prediction accuracy. Zichong Li, Yanbo Xu, Simiao Zuo, Haoming Jiang, Chao Zhang 0014, Tuo Zhao, Hongyuan Zha |
ICML | 3 |
| 2023 | Less is More: Task-aware Layer-wise Distillation for Language Model CompressionabstractLayer-wise distillation is a powerful tool to compress large models (i.e. teacher models) into small ones (i.e., student models). The student distills knowledge from the teacher by mimicking the hidden representations of the teacher at every intermediate layer. However, layer-wise distillation is difficult. Since the student has a smaller model capacity than the teacher, it is often under-fitted. Furthermore, the hidden representations of the teacher contain redundant information that the student does not necessarily need for the target task's learning. To address these challenges, we propose a novel Task-aware layEr-wise Distillation (TED). TED designs task-aware filters to align the hidden representations of the student and the teacher at each layer. The filters select the knowledge that is useful for the target task from the hidden representations. As such, TED reduces the knowledge gap between the two models and helps the student to fit better on the target task. We evaluate TED in two scenarios: continual pre-training and fine-tuning. TED demonstrates significant and consistent improvements over existing distillation methods in both scenarios. Code is available at https://github.com/cliang1453/task-aware-distillation. Chen Liang 0006, Simiao Zuo, Qingru Zhang, Weizhu Chen, Tuo Zhao |
ICML | 2 |
| 2023 | Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable AlgorithmsabstractMulti-Agent Reinforcement Learning (MARL) has shown promising results across several domains. Despite this promise, MARL policies often lack robustness and are therefore sensitive to small changes in their environment. This presents a serious concern for the real world deployment of MARL algorithms, where the testing environment may slightly differ from the training environment. In this work we show that we can gain robustness by controlling a policy’s Lipschitz constant, and under mild conditions, establish the existence of a Lipschitz and close-to-optimal policy. Motivated by these insights, we propose a new robust MARL framework, ERNIE, that promotes the Lipschitz continuity of the policies with respect to the state observations and actions by adversarial regularization. The ERNIE framework provides robustness against noisy observations, changing transition dynamics, and malicious actions of agents. However, ERNIE’s adversarial regularization may introduce some training instability. To reduce this instability, we reformulate adversarial regularization as a Stackelberg game. We demonstrate the effectiveness of the proposed framework with extensive experiments in traffic light control and particle environments. In addition, we extend ERNIE to mean-field MARL with a formulation based on distributionally robust optimization that outperforms its non-robust counterpart and is of independent interest. Our code is available at https://github.com/abukharin3/ERNIE. Alexander Bukharin, Yue Yu 0001, Qingru Zhang, Zhehui Chen, Simiao Zuo, Chao Zhang 0014, Songan Zhang, Tuo Zhao |
NeurIPS | 6 |
| 2022 | No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang 0006, Haoming Jiang, Simiao Zuo, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao |
ICLR | 3 |
| 2022 | Taming Sparsely Activated Transformer with Stochastic Experts
Simiao Zuo, Xiaodong Liu 0003, Jian Jiao 0007, Young Jin Kim 0006, Hany Hassan, Ruofei Zhang, Jianfeng Gao 0001, Tuo Zhao |
ICLR | 1 |
| 2022 | PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight ImportanceabstractLarge Transformer-based models have exhibited superior performance in various natural language processing and computer vision tasks. However, these models contain enormous amounts of parameters, which restrict their deployment to real-world applications. To reduce the model size, researchers prune these models based on the weights’ importance scores. However, such scores are usually estimated on mini-batches during training, which incurs large variability/uncertainty due to mini-batch sampling and complicated training dynamics. As a result, some crucial weights could be pruned by commonly used pruning methods because of such uncertainty, which makes training unstable and hurts generalization. To resolve this issue, we propose PLATON, which captures the uncertainty of importance scores by upper confidence bound of importance estimation. In particular, for the weights with low importance scores but high uncertainty, PLATON tends to retain them and explores their capacity. We conduct extensive experiments with several Transformer-based models on natural language understanding, question answering and image classification to validate the effectiveness of PLATON. Results demonstrate that PLATON manifests notable improvement under different sparsity levels. Our code is publicly available at https://github.com/QingruZhang/PLATON. Qingru Zhang, Simiao Zuo, Chen Liang 0006, Alexander Bukharin, Weizhu Chen, Tuo Zhao |
ICML | 2 |
| 2022 | MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided AdaptationabstractSimiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Simiao Zuo, Qingru Zhang, Chen Liang 0006, Tuo Zhao, Weizhu Chen |
NAACL-HLT | 1 |
| 2021 | Super Tickets in Pre-Trained Language Models: From Model Compression to Improving GeneralizationabstractChen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Liang 0006, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu 0003, Tuo Zhao, Weizhu Chen |
ACL/IJCNLP (1) | 2 |
| 2021 | Adversarial Regularization as Stackelberg Game: An Unrolled Optimization ApproachabstractAdversarial regularization has been shown to improve the generalization performance of deep learning models in various natural language processing tasks.Existing works usually formulate the method as a zero-sum game, which is solved by alternating gradient descent/ascent algorithms.Such a formulation treats the adversarial and the defending players equally, which is undesirable because only the defending player contributes to the generalization performance.To address this issue, we propose Stackelberg Adversarial Regularization (SALT), which formulates adversarial regularization as a Stackelberg game.This formulation induces a competition between a leader and a follower, where the follower generates perturbations, and the leader trains the model subject to the perturbations.Different from conventional approaches, in SALT, the leader is in an advantageous position.When the leader moves, it recognizes the strategy of the follower and takes the anticipated follower's outcomes into consideration.Such a leader's advantage enables us to improve the model fitting to the unperturbed data.The leader's strategic information is captured by the Stackelberg gradient, which is obtained using an unrolling algorithm.Our experimental results on a set of machine translation and natural language understanding tasks show that SALT outperforms existing adversarial regularization baselines across all tasks.Our code is publicly available. Simiao Zuo, Chen Liang 0006, Haoming Jiang, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao |
EMNLP (1) | 1 |
| 2021 | A Hypergradient Approach to Robust Regression without Correspondence
Yujia Xie, Yixiu Mao, Simiao Zuo, Hongteng Xu, Xiaojing Ye, Tuo Zhao, Hongyuan Zha |
ICLR | 3 |
| 2021 | Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training ApproachabstractYue Yu, Simiao Zuo, Haoming Jiang, Wendi Ren, Tuo Zhao, Chao Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yue Yu 0001, Simiao Zuo, Haoming Jiang, Wendi Ren, Tuo Zhao, Chao Zhang 0014 |
NAACL-HLT | 2 |
| 2020 | Transformer Hawkes ProcessabstractModern data acquisition routinely produce massive amounts of event sequence data in various domains, such as social media, healthcare, and financial markets. These data often exhibit complicated short-term and long-term temporal dependencies. However, most of the existing recurrent neural network based point process models fail to capture such dependencies, and yield unreliable prediction performance. To address this issue, we propose a Transformer Hawkes Process (THP) model, which leverages the self-attention mechanism to capture long-term dependencies and meanwhile enjoys computational efficiency. Numerical experiments on various datasets show that THP outperforms existing models in terms of both likelihood and event prediction accuracy by a notable margin. Moreover, THP is quite general and can incorporate additional structural knowledge. We provide a concrete example, where THP achieves improved prediction performance for learning multiple point processes when incorporating their relational information. Simiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao, Hongyuan Zha |
ICML | 1 |
| 2019 | Tensor maps for synchronizing heterogeneous shape collectionsabstractEstablishing high-quality correspondence maps between geometric shapes has been shown to be the fundamental problem in managing geometric shape collections. Prior work has focused on computing efficient maps between pairs of shapes, and has shown a quantifiable benefit of joint map synchronization, where a collection of shapes are used to improve (denoise) the pairwise maps for consistency and correctness. However, these existing map synchronization techniques place very strong assumptions on the input shapes collection such as all the input shapes fall into the same category and/or the majority of the input pairwise maps are correct. In this paper, we present a multiple map synchronization approach that takes a heterogeneous shape collection as input and simultaneously outputs consistent dense pairwise shape maps. We achieve our goal by using a novel tensor-based representation for map synchronization, which is efficient and robust than all prior matrix-based representations. We demonstrate the usefulness of this approach across a wide range of geometric shape datasets and the applications in shape clustering and shape co-segmentation. Qixing Huang, Zhenxiao Liang, Haoyun Wang, Simiao Zuo, Chandrajit L. Bajaj |
ACM Trans. Graph. | 4 |