VLDB 2026 Research / reviewers in the wild / expert
Zhengyan Zhang
dblp:23/10446
· DBLP profile ↗
38ranked-venue papers
7as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 23 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Generalized Partial-to-Full Point Set Registration With Overlap-Guided Bidirectional Hybrid Mixture Models for Computer-Assisted Orthopedic SurgeryabstractIn computer-assisted orthopedic surgery (CAOS), robust and accurate registration of the preoperative full bone model and the intraoperative partial point set is a prerequisite for reliable surgical navigation, yet remains highly challenging due to partial overlap, noise, and outliers. We propose the Overlap-Guided Bidirectional Hybrid Mixture Registration (OBHMR) framework for robust and accurate partial-to-full registration. First, geometric features (i.e., surface normals) extracted from raw point sets are incorporated in both correspondence estimation and transform computation. Meanwhile, we formulate a hybrid mixture model that jointly represents positions with Gaussian mixtures (GMMs) and normals with von Mises–Fisher (vMF) mixtures across the two generalized point sets. Second, a dual-branch overlap prediction network leverages feature similarity and geometric structure to provide accurate point-wise overlap scores that guide hybrid-mixture construction under partial overlap. Third, a correspondence module integrates rotation-invariant features, multi-level self-attention, and clustering-based refinement to enhance reliability under noise and misalignment. Finally, a bidirectional objective jointly aligns source-to-target and target-to-source mixtures, explicitly accounting for discrepancies induced by noise and outliers in both the preoperative and intraoperative point sets to achieve robust optimization. Extensive experiments on 1,399 femur and 1,301 hip models demonstrate superior performance over state-of-the-art methods across overlap ratios from 5% to 70%, under both isotropic and anisotropic noise and outlier ratios up to 100%, achieving errors as low as 1.27° rotation and 1.18 mm translation at 50% overlap with 2.5mm noise. Additional tests on liver and ModelNet40 confirm strong generalization across medical and non-medical data. Ablation studies further validate the contributions of normals, overlap estimation, and the bidirectional formulation. Xinzhe Du, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | DeepBHMR: Learning Bidirectional Hybrid Mixture Models for Generalized Global Rigid Point Set Registration in Computer-Assisted Orthopedic SurgeryabstractThis paper presents a novel robust and accurate normal-assisted learning-based rigid point set registration approach, i.e., Deep Bi-directional Hybrid Mixture Registration (DeepBHMR), where normal vectors are used in both correspondence and transformation computational stages while the bi-directional registration processes are considered. DeepBHMR consists of three components, (1) the correspondence estimation network that predicts the correspondence probabilities; (2) the posterior estimation module that computes the HMMs parameters; (3) the transformation estimation module that calculates the rigid transformation matrix by utilizing the bidirectional optimization mechanism. DeepBHMR has been extensively validated on various medical data sets, outperforming state-of-the-art registration methods. For femur bones, the mean rotation error value is approximately 1° (i.e, 1.01°) and the translation error is less than 1 mm (i.e., 0.30 mm) respectively, which meets the requirement of computer-assisted orthopedic surgery. Furthermore, even (1) trained with femur data and tested on distinct shapes and (2) under the large transformation, the mean RMSE values of registration are 2.60 mm and 3.05 mm respectively, demonstrating DeepBHMR’s favorable generalizability to different data shapes and great capability to handle global registration. Additionally, the individual significant contributions and computational efficiency of adopting normal vectors and utilizing the bidirectional mechanism have been validated in ablation studies. The results demonstrate the DeepBHMR’s favorable generalizability from femur bones to hip bones and that DeepBHMR can successfully handle the large transformation or partial-to-full registration simultaneously. The code implementation of DeepBHMR has been made publicly available at https://github.com/zzyrobot/DeepBiHMM.git. Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionabstractJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 0002, Liang Zhao 0026, Zhengyan Zhang, Zhenda Xie, Lean Wang, Zhiping Xiao 0001, Chong Ruan, Ming Zhang 0004, Wenfeng Liang, Wangding Zeng |
ACL (1) | 6 |
| 2025 | ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language ModelsabstractActivation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to pursue activation sparsity and acceleration, but few can simultaneously obtain high activation sparsity and comparable model performance. This paper introduces a simple and effective method named “ProSparse” to sparsify LLMs while achieving both targets. Specifically, after introducing ReLU activation, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing for multiple stages. This can enhance activation sparsity and mitigate performance degradation by avoiding radical shifts in activation distributions. With ProSparse, we obtain high sparsity of 89.32% for LLaMA2-7B, 88.80% for LLaMA2-13B, and 87.89% for end-size MiniCPM-1B, respectively, with comparable performance to their original Swish-activated versions. These present the most sparsely activated models among open-source LLaMA versions and competitive end-size models. Inference acceleration experiments further demonstrate the significant practical acceleration potential of LLMs with higher activation sparsity, obtaining up to 4.52x inference speedup. Xu Han 0007, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Zhiyuan Liu 0001, Guangli Li, Maosong Sun 0001 |
COLING | 3 |
| 2025 | Registration After Completion: Towards Sparse and Partial Point Set Registration for Computer-Assisted Orthopedic SurgeryabstractIn computer-assisted orthopedic surgery (CAOS), accurate point set registration is essential for enhancing surgical accuracy. However, the sparse and low-overlap nature of intraoperative point sets presents significant challenges for reliable registration. To deal with these challenges, we propose a novel registration-after-completion framework, where the intraoperative point set is first completed, after which the two full point sets are registered. Our main contributions include the follows. First, we propose a progressive two-stage strategy to progressively complete the sparse and partial intraoperative point set. Second, considering that 1) intra-operative point set contains noise 2) the point completion process is not perfect, and 3) the resolution of preoperative image is limited, we adopt the bidirectional hybrid mixture models (HMMs) to represent the point set pairs and formulate the probabilistic registration network. In the proposed novel correspondence network where a dual-path cross-attention mechanism is adopted for feature fusion and a clustering mechanism is leveraged for calculating point-to-mixture correspondences. Furthermore, the bidirectional registration mechanism is leveraged to compute the transformation based on estimated correspondences. Third, we have extensively validated the proposed approach on various datasets and bone phantoms. Our experiments on 1399 human femur and 1301 hip models demonstrate that our method achieves state-of-the-art performance across overlap rates from 15% to 35% and at various point counts (i.e., 25, 50, and 100 points) under conditions with less than 50% overlap. Additionally, real phantom experiments on femur and hip models validate the method’s performance in simulated surgical scenarios. Experiments on ModelNet40 further confirmed our method’s effectiveness and generalizability. Xinzhe Du, Shixing Ma, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min |
IROS | 3 |
| 2025 | Revisiting 3D Curve to Surface Registration using Tangent and Normal Vectors for Computer-Assisted Orthopedic SurgeryabstractIn this paper, we present a novel curve-to-surface registration method, termed Bi-directional Hybrid Mixture Model Registration based on Dual-constrained Tangent and Normal Vectors (BiHMM-DTN), where two different tangent vectors at the intraoperative point are simultaneously used with the normal vector at the corresponding preoperative point to construct the geometric constraints. While hybrid registration models incorporating tangent and normal vectors (HMM-TN) demonstrate success, their geometric constraints prove inadequate or inappropriate for sparse intraoperative point sets, frequently yielding suboptimal optimization outcomes. By critically revisiting the geometric constraints of HMM-TN, we propose a dual-constraints-based hybrid mixture model registration framework with enhanced intraoperative point set acquisition protocols. To deal with noise and outliers in preoperative and intraoperative point sets—caused by reconstruction inaccuracies and tracking errors, respectively—our approach employs a bi-directional registration mechanism for curve-to-surface registration. We provide rigorous proofs validating the geometric completeness of the dual constraints within this mechanism. The BiHMM-DTN framework is formulated as a maximum likelihood estimation (MLE) problem and optimized using an expectation-maximization (EM) algorithm. Furthermore, to enhance convergence stability and accelerate optimization, the rotation matrix is updated iteratively through successive incremental steps. Extensive experiments on human femur and hip models demonstrate that our method outperforms state-of-the-art approaches, including both traditional optimization and deep learning methods, under various noise and outlier conditions. Furthermore, real-world phantom experiments highlight the potential clinical value of our method for surgical navigation applications. The codes and data are available at https://github.com/sam-zyzhang/BiHMM-DTN.git. Zhengyan Zhang, Xinzhe Du, Rui Song 0002, Max Q.-H. Meng, Zhe Min |
IROS | 1 |
| 2024 | Robust and Scalable Model Editing for Large Language ModelsabstractLarge language models (LLMs) can make predictions using parametric knowledge – knowledge encoded in the model weights – or contextual knowledge – knowledge presented in the context. In many scenarios, a desirable behavior is that LLMs give precedence to contextual knowledge when it conflicts with the parametric knowledge, and fall back to using their parametric knowledge when the context is irrelevant. This enables updating and correcting the model’s knowledge by in-context editing instead of retraining. Previous works have shown that LLMs are inclined to ignore contextual knowledge and fail to reliably fall back to parametric knowledge when presented with irrelevant context. In this work, we discover that, with proper prompting methods, instruction-finetuned LLMs can be highly controllable by contextual knowledge and robust to irrelevant context. Utilizing this feature, we propose EREN (Edit models by REading Notes) to improve the scalability and robustness of LLM editing. To better evaluate the robustness of model editors, we collect a new dataset, that contains irrelevant questions that are more challenging than the ones in existing datasets. Empirical results show that our method outperforms current state-of-the-art methods by a large margin. Unlike existing techniques, it can integrate knowledge from multiple edits, and correctly respond to syntactically similar but semantically unrelated inputs (and vice versa). The source code can be found at https://github.com/thunlp/EREN. Yingfa Chen, Zhengyan Zhang, Xu Han 0007, Chaojun Xiao, Zhiyuan Liu 0001, Kuai Li, Maosong Sun 0001 |
LREC/COLING | 2 |
| 2024 | Exploring the Benefit of Activation Sparsity in Pre-trainingabstractPre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activation has been explored through post-training methods, its potential in pre-training remains untapped. In this work, we first study how activation properties change during pre-training. Our examination reveals that Transformers exhibit sparse activation throughout the majority of the pre-training process while the activation correlation keeps evolving as training progresses. Leveraging this observation, we propose Switchable Sparse-Dense Learning (SSD). SSD adaptively switches between the Mixtures-of-Experts (MoE) based sparse training and the conventional dense training during the pre-training process, leveraging the efficiency of sparse training and avoiding the static activation correlation of sparse training. Compared to dense training, SSD achieves comparable performance with identical model size and reduces pre-training costs. Moreover, the models trained with SSD can be directly used as MoE models for sparse inference and achieve the same performance as dense models with up to $2\times$ faster inference speed. Codes are available at https://github.com/thunlp/moefication. Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Ruobing Xie, Maosong Sun 0001, Jie Zhou 0016 |
ICML | 1 |
| 2024 | OBHMR: Robust Partial-to-full Generalized Point Set Registration with Overlap-guided Bidirectional Hybrid Mixture ModelabstractIn this paper, we introduce a novel overlap-based bidirectional point set registration approach, i.e., Overlap-guided Bidirectional Hybrid Mixture Registration (OBHMR), which incorporates geometric information (i.e., normal vectors) in both the correspondence and transformation stages and formulates the optimization objective of registration in a bidirectional manner. More importantly, to address the issue of partial-to-full registration, OBHMR utilises the predicted point-wise overlap score using networks to formulate the overlap-guided Hybrid Mixture Model consisting of the Gaussian Mixture Model (GMM) and Fisher Mixture Model (FMM). OBHMR contains four components: (1) the overlap-guided correspondence network that estimates the correspondence probabilities and calculates the point-wise overlap score; (2) the learning posterior module that estimates the overlap-guided HMM parameters; (3) the transformation module that computes the rigid transformation by formulating the optimisation objective in a bidirectional registration way, given correspondences and overlap-guided HMM parameters. Experiments using 1457 human femur and 1301 human hip models demonstrate significant improvements in partial-to-full registration performance (p < 0.01) under different overlapping ratios, compared to state-of-the-art registration approaches. Furthermore, individual contributions of three modules (i.e., additional normal vectors, overlap score estimation module and the bidirectional mechanism) in OBHMR have been validated in ablation studies. The results demonstrate OBHMR’s capability of tackling the challenging partial-to-full registration problems in computer-assisted orthopedic surgery. The codes are available at https://github.com/Dxinz/DeepOBHMR. Xinzhe Du, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min |
IROS | 2 |
| 2024 | DeepBHMR: Learning Bidirectional Hybrid Mixture Models for Generalized Rigid Point Set RegistrationabstractIn this paper, we introduce a novel normal-assisted learning-based rigid registration approach, i.e., Deep Bi-directional Hybrid Mixture Registration (DeepBHMR). Our approach utilises helpful normal vectors explicitly in both correspondence and transformation stages and formulates the optimization objective of registration in a bi-directional way that considers noise in both point sets. DeepBHMR consists of three modules: (1) the correspondence network that estimates the correspondence probability relating points within one generalized point set (i.e., positional and normal vectors) with components of Hybrid Mixture Models (HMMs) representing the other generalized point set; (2) the posterior module that computes HMMs parameters; (3) the transformation module that computes the rotation matrix and the translation vector given the estimated generalized-point to hybrid-distribution correspondences and HMMs parameters. DeepBHMR has been validated on 291 human femur and 260 hip models, and extensive experimental results demonstrate that DeepBHMR outperforms the state-of-the-art registration methods (p-value < 0.01). In the circumstance of femur bones, the mean rotation and translation error values are around 1° (i.e., 1.01°) and less than 1 mm (i.e., 0.36mm), respectively. Furthermore, even under the large transformation (i.e., in the range of [0,180]° and [0, 100] mm), the mean RMSE values being 3.05 mm is still satisfactory. Additionally, the results demonstrate the DeepBHMR’s favorable generalizability from femur shapes to hip shapes. We have carefully validated the significant benefits of incorporating normal vectors and the bidirectional mechanism. DeepBHMR can successfully handle the challenging scenario of large transformation and partial registration. The codes are available at https://github.com/zzyrobot/DeepBHMR.git. Zhe Min, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng |
IROS | 2 |
| 2024 | InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context MemoryabstractLarge language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to 1,024K, InfLLM still effectively captures long-distance dependencies. Our code can be found at https://github.com/thunlp/InfLLM. Chaojun Xiao, Pengle Zhang, Xu Han 0007, Guangxuan Xiao, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 6 |
| 2024 | DOA Estimation Method Based on Unsupervised Learning Network With Threshold Capon Spectrum Weighted PenaltyabstractIn complex electronic countermeasure environment, direction-of-arrival (DOA) is very important for targets detection, localization and tracking. However, the power of interference is usually stronger than that of signal, which degrades the DOA estimation performance severely, and even makes DOA estimation failure. To solve this issue, this paper proposes a DOA estimation method based on unsupervised learning network with threshold Capon spectrum weighted penalty. In this work, an unsupervised network is proposed to obtain the DOA estimation spectrum, in which labels are no longer required. Furthermore, deep unfolded layers are introduced to remove the iterative solution of sparse recovery and increase the depth of network. Additionally, loss function contains reconstruction error and penalty term is developed to generate zero traps in direction of interference and signal, overcoming the influence of strong interference. Both numerical simulations and experiments demonstrate the effectiveness of the proposed method. Zhengyan Zhang, Wolin Li, Hongzhe Miao |
IEEE Signal Process. Lett. | 1 |
| 2024 | CGAN: lightweight and feature aggregation network for high-performance interactive image segmentation
Yan Gui, Zhengyan Zhang, Jin Zhang 0018 |
Vis. Comput. | 2 |
| 2023 | READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input NoisesabstractFor many real-world applications, the usergenerated inputs usually contain various noises due to speech recognition errors caused by linguistic variations 1 or typographical errors (typos).Thus, it is crucial to test model performance on data with realistic input noises to ensure robustness and fairness.However, little study has been done to construct such benchmarks for Chinese, where various languagespecific input noises happen in the real world.In order to fill this important gap, we construct READIN: a Chinese multi-task benchmark with REalistic And Diverse Input Noises.READIN contains four diverse tasks and requests annotators to re-enter the original test data with two commonly used Chinese input methods: Pinyin input and speech input.We designed our annotation pipeline to maximize diversity, for example by instructing the annotators to use diverse input method editors (IMEs) for keyboard noises and recruiting speakers from diverse dialectical groups for speech noises.We experiment with a series of strong pretrained language models as well as robust training methods, we find that these models often suffer significant performance drops on READIN even with robustness methods like data augmentation.As the first large-scale attempt in creating a benchmark with noises geared towards user-generated inputs, we believe that READIN serves as an important complement to existing Chinese NLP benchmarks.The source code and dataset can be obtained from https://github.com/ thunlp/READIN. Chenglei Si, Zhengyan Zhang, Yingfa Chen, Xiaozhi Wang, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 2 |
| 2023 | Plug-and-Play Document Modules for Pre-trained ModelsabstractChaojun Xiao, Zhengyan Zhang, Xu Han, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Xiangyang Li, Zhonghua Li, Zhao Cao, Maosong Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chaojun Xiao, Zhengyan Zhang, Xu Han 0007, Chi-Min Chan, Yankai Lin 0001, Zhiyuan Liu 0001, Zhao Cao, Maosong Sun 0001 |
ACL (1) | 2 |
| 2023 | Plug-and-Play Knowledge Injection for Pre-trained Language ModelsabstractZhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang, Deming Ye, Chaojun Xiao, Xu Han, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhengyan Zhang, Yankai Lin 0001, Deming Ye, Chaojun Xiao, Xu Han 0007, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
ACL (1) | 1 |
| 2023 | A Transferable Multi-Agent Reinforcement Learning Method for Distribution Service RestorationabstractThe occurrence of extreme events, which has increased the risk of major outages in the grid, makes the quick and efficient recovery of load in the distribution network become a key issue. The data-driven deep reinforcement learning method has great potential in providing fast decision-making. However, a large number of agents lead to the curse of dimensionality, making it inefficient to obtain effective control strategies. When tackling similar tasks in different power gird, the retraining of multiple agents will bring us great cost. To solve this problem, we propose a transferable multi-agent reinforcement learning framework that employs model reload and buffer reuse methods to transfer control strategies from small-scale simple scenes to large-scale complex scenes. It also utilizes attention mechanisms to aggregate observation features and handles the problem of variable observation dimensions. Finally, the distribution service restoration problem is modeled as a Markov decision process and solved using the QMIX algorithm. The performance of the proposed method has been verified in IEEE 34-node and IEEE 123-node distribution systems. Ruiqi Si, Ji Qiao, Kaixuan Ji, Zhang Jun, Xuanying Pan, Zhengyan Zhang |
SMC | 8 |
| 2023 | Sub-Character Tokenization for Chinese Pretrained Language ModelsabstractAbstract Tokenization is fundamental to pretrained language models (PLMs). Existing tokenization methods for Chinese PLMs typically treat each character as an indivisible token. However, they ignore the unique feature of the Chinese writing system where additional linguistic information exists below the character level, i.e., at the sub-character level. To utilize such information, we propose sub-character (SubChar for short) tokenization. Specifically, we first encode the input text by converting each Chinese character into a short sequence based on its glyph or pronunciation, and then construct the vocabulary based on the encoded text with sub-word segmentation. Experimental results show that SubChar tokenizers have two main advantages over existing tokenizers: 1) They can tokenize inputs into much shorter sequences, thus improving the computational efficiency. 2) Pronunciation-based SubChar tokenizers can encode Chinese homophones into the same transliteration sequences and produce the same tokenization output, hence being robust to homophone typos. At the same time, models trained with SubChar tokenizers perform competitively on downstream tasks. We release our code and models at https://github.com/thunlp/SubCharTokenization to facilitate future work. Chenglei Si, Zhengyan Zhang, Yingfa Chen, Fanchao Qi, Xiaozhi Wang, Zhiyuan Liu 0001, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Anisotropic Generalized Bayesian Coherent Point Drift for Point Set RegistrationabstractRegistration is highly demanded in many real-world scenarios such as robotics and automation. Registration is challenging partly due to the fact that the acquired data is usually noisy and has many outliers. In addition, in many practical applications, one point set (PS) usually only covers a partial region of the other PS. Thus, most existing registration algorithms cannot guarantee theoretical convergence. This article presents a novel, robust, and accurate three-dimensional (3D) rigid point set registration (PSR) method, which is achieved by generalizing the state-of-the-art (SOTA) Bayesian coherent point drift (BCPD) theory to the scenario that high-dimensional point sets (PSs) are aligned and the anisotropic positional noise is considered. The high-dimensional point sets typically consist of the positional vectors and normal vectors. On one hand, with the normal vectors, the proposed method is more robust to noise and outliers, and the point correspondences can be found more accurately. On the other hand, incorporating the registration into the BCPD framework will guarantee the algorithm’s theoretical convergence. Our contributions in this article are three folds. First, the problem of rigidly aligning two general PSs with normal vectors is incorporated into a variational Bayesian inference framework, which is solved by generalizing the BCPD approach while the anisotropic positional noise is considered. Second, the updated parameters during the algorithm’s iterations are given in closed-form or with iterative solutions. Third, extensive experiments have been done to validate the proposed approach and its significant improvements over the BCPD. Note to Practitioners—This paper was motivated by the problem of 3D rigid PSR for computer-assisted surgery (CAS), especially in orthopedic applications. The proposed algorithm is also suitable for other scenarios where the initial coarse registration is conducted. The traditional registration methods are susceptible to noise (especially anisotropic noise), outliers, and incomplete partial data. This paper generalizes the recently proposed BCPD method to the six-dimensional scenario where anisotropic positional noise is considered and normal vectors are incorporated. The proposed noise model is decomposed into three parts to be solved alternately: the membership probability of mixture distributions, the soft correspondence estimation, and the model parameters (i.e., the rotation matrix, translation vector, the covariance matrix with the anisotropic positional error, and the concentration parameter with the estimation of the normal vectors). Especially, the convergence is guaranteed at the theoretical level using the variational inference theory. The experimental results demonstrate the superiority of our algorithm on registration accuracy, convergence speed, and robustness to noise, outliers, and partial data. Zhe Min, Zhengyan Zhang, Xing Yang 0005, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2022 | Automatic Label Sequence Generation for Prompting Sequence-to-sequence ModelsabstractPrompting, which casts downstream applications as language modeling tasks, has shown to be sample efficient compared to standard fine-tuning with pre-trained models. However, one pitfall of prompting is the need of manually-designed patterns, whose outcome can be unintuitive and requires large validation sets to tune. To tackle the challenge, we propose AutoSeq, a fully automatic prompting method: (1) We adopt natural language prompts on sequence-to-sequence models, enabling free-form generation and larger label search space; (2) We propose label sequences – phrases with indefinite lengths to verbalize the labels – which eliminate the need of manual templates and are more expressive than single label words; (3) We use beam search to automatically generate a large amount of label sequence candidates and propose contrastive re-ranking to get the best combinations. AutoSeq significantly outperforms other no-manual-design methods, such as soft prompt tuning, adapter tuning, and automatic search on single label words; the generated label sequences are even better than curated manual ones on a variety of tasks. Our method reveals the potential of sequence-to-sequence models in few-shot learning and sheds light on a path to generic and automatic prompting. The source code of this paper can be obtained from https://github.com/thunlp/Seq2Seq-Prompt. Zichun Yu, Tianyu Gao 0001, Zhengyan Zhang, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
COLING | 3 |
| 2022 | Finding Skill Neurons in Pre-trained Transformer-based Language ModelsabstractTransformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks.However, it remains unclear how the skills required to handle these tasks distribute among model parameters.In this paper, we find that after prompt tuning for specific tasks, the activations of some neurons within pre-trained Transformers 1 are highly predictive of the task labels.We dub these neurons skill neurons and confirm they encode task-specific skills by finding that: (1) Skill neurons are crucial for handling tasks.Performances of pre-trained Transformers on a task significantly drop when corresponding skill neurons are perturbed.(2) Skill neurons are task-specific.Similar tasks tend to have similar distributions of skill neurons.Furthermore, we demonstrate the skill neurons are most likely generated in pre-training rather than finetuning by showing that the skill neurons found with prompt tuning are also crucial for other fine-tuning methods freezing neuron weights, such as the adapter-based tuning and BitFit.We also explore the applications of skill neurons, including accelerating Transformers with network pruning and building better transferability indicators.These findings may promote further research on understanding Transformers.The source code can be obtained from https: //github.com/THU-KEG/Skill-Neuron. Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou 0001, Zhiyuan Liu 0001, Juan-Zi Li |
EMNLP | 3 |
| 2022 | Effective Few-Shot Named Entity Linking by Meta-LearningabstractEntity linking aims to link ambiguous mentions to their corresponding entities in a knowledge base, which is significant and fundamental for various downstream applications, e.g., knowledge base completion, question answering, and information extraction. While great efforts have been devoted to this task, most of these studies follow the assumption that large-scale labeled data is available. However, when the labeled data is insufficient for specific domains due to labor-intensive annotation work, the performance of existing algorithms will suffer an intolerable decline. In this paper, we endeavor to solve the problem of few-shot entity linking, which only requires a minimal amount of in-domain labeled data and is more practical in real situations. Specifically, we firstly propose a novel weak supervision strategy to generate non-trivial synthetic entity-mention pairs based on mention rewriting. Since the quality of the synthetic data has a critical impact on effective model training, we further design a meta-learning mechanism to assign different weights to each synthetic entity-mention pair automatically. Through this way, we can profoundly exploit rich and precious semantic information to derive a well-trained entity linking model under the few-shot setting. The experiments on real-world datasets show that the proposed method can extensively improve the state-of-the-art few-shot entity linking model and achieve impressive performance when only a small amount of labeled data is available. Moreover, we also demonstrate the outstanding ability of the model's transferability. Our code and models will be open-sourced. Xiuxing Li, Zhenyu Li 0008, Zhengyan Zhang, Ning Liu 0014, Haitao Yuan 0002, Wei Zhang 0056, Zhiyuan Liu 0001, Jianyong Wang 0001 |
ICDE | 3 |
| 2022 | Generalized 3D Rigid Point Set Registration with Anisotropic Positional Error Based on Bayesian Coherent Point DriftabstractThis paper presents a novel, robust, and accurate three-dimensional (3D) rigid point set registration (PSR) method, which is achieved by generalizing the state-of-the-art (SOTA) Bayesian coherent point drift (BCPD) theory to the scenario that high-dimensional point sets(PSs) are aligned and that the anisotropic positional noise is considered. Our contributions in this paper are three folds. First, the problem of rigidly aligning two general point sets (PSs) with normal vectors is incorporated into a variational Bayesian inference framework, which is solved by generalizing the BCPD approach while the anisotropic positional noise is considered. Second, the updated parameters during the algorithm's iterations are given in closed-form or iterative solutions. Third, extensive experiments have been done to validate the proposed approach and its significant improvements over the BCPD. Zhe Min, Xing Yang 0005, Zhengyan Zhang, Max Q.-H. Meng |
ICRA | 4 |
| 2022 | MO-Transformer: A Transformer-Based Multi-Object Point Cloud Reconstruction NetworkabstractThis paper proposes a new network for reconstructing multi-object point cloud. Different from previous networks which reconstruct multi-object point cloud as a whole, our network iteratively reconstructs each individual object point cloud from a frame of multi-object point cloud. To achieve this goal, we have designed MO-Transformer, a transformer-based autoregressive network. During training, MO-Transformer takes a frame of multi-object point cloud and individual object point clouds as input. During testing, MO-Transformer iteratively reconstructs individual object point clouds only based on the input multi-object point cloud. To train the proposed MO-Transformer, we design a new loss function called separate Chamfer distance (SCD). In addition, we prove that SCD is an upper bound of the traditional Chamfer distance calculated based on the entire multi-object point cloud. The reconstruction experiment verifies the efficacy of our network in multi-object point cloud reconstruction. Furthermore, the reconstruction experiment also investigates the effect of different dimensions using a series of datasets. The ablation study experiment verifies the necessity of SCD in training MO-Transformer. Erli Lyu, Zhengyan Zhang, Wei Liu 0134, Jiaole Wang, Shuang Song 0002, Max Q.-H. Meng |
IROS | 2 |
| 2022 | Knowledge Inheritance for Pre-trained Language ModelsabstractYujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yujia Qin, Yankai Lin 0001, Jing Yi, Xu Han 0007, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 6 |
| 2022 | Liver function classification based on local direction number and non-local binary pattern
Weijia Huang, Zhengyan Zhang, Caiping Xi, Yuanjiang Li |
Multim. Tools Appl. | 3 |
| 2022 | Generalized Point Set Registration With Fuzzy Correspondences Based on Variational Bayesian InferenceabstractPoint set registration (PSR) is an essential problem in surgical navigation and computer-assisted surgery (CAS). In CAS, PSR can be used to map the intraoperative surgical space with the preoperative volumetric image space. The performances of PSR in real-world surgical scenarios are sensitive to noise and outliers. This article proposes a novel point set registration approach where the additional features (i.e., the normal vectors) extracted from the point sets are utilized and the convergence of the algorithm is guaranteed from the theoretical perspective. More specifically, we formulate the PSR with normal vectors by generalizing the Bayesian coherent point drift (BCPD) into the 6-D scenario. The proposed algorithm is more accurate and robust to noise and outliers, and the theoretical convergence of the proposed approach is guaranteed. Our contributions of this article are summarized as follows. 1) The PSR problem with normal vectors is formally formulated through generalizing the BCPD approach. 2) The formulas for updating the parameters during the algorithm’s iterations are given in closed forms. 3) Extensive experiments have been done to verify the proposed approach and specifically its significant improvements over the BCPD has been validated. Zhe Min, Zhengyan Zhang, Max Q.-H. Meng |
IEEE Trans. Fuzzy Syst. | 3 |
| 2022 | COSINE: Compressive Network Embedding on Large-Scale Information NetworksabstractThere is recently a surge in approaches that learn low-dimensional embeddings of nodes in networks. However, for large-scale real-world networks, it’s inefficient for existing approaches to store amounts of parameters in memory and update them edge by edge. With the knowledge that nodes having similar neighborhoods will be close to each other in the embedding space, we propose COSINE (COmpresSIve Network Embedding) algorithm, which reduces the memory footprint and accelerates the training process by parameter sharing among similar nodes. COSINE applies graph partitioning algorithms to networks and builds parameter sharing dependency of nodes based on the results of partitioning. In this way, COSINE injects prior knowledge about high-order structural information into models, which makes network embedding more efficient and effective. COSINE can be applied to anyembedding lookupmethod and learn high-quality embeddings with limited memory and less training time. We conduct experiments on multi-label classification and link prediction, where baselines and our model have the same memory usage. Experimental results show that COSINE improves baselines by up to 23 percent on classification and 25 percent on link prediction. Moreover, the training time of all representation learning methods using COSINE decreases by 30 to 70 percent. Zhengyan Zhang, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Zhichong Fang, Bo Zhang 0056, Leyu Lin |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | A Unified Understanding of Deep NLP Models for Text ClassificationabstractThe rapid development of deep natural language processing (NLP) models for text classification has led to an urgent need for a unified understanding of these models proposed individually. Existing methods cannot meet the need for understanding different models in one framework due to the lack of a unified measure for explaining both low-level (e.g., words) and high-level (e.g., phrases) features. We have developed a visual analysis tool, DeepNLPVis, to enable a unified understanding of NLP models for text classification. The key idea is a mutual information-based measure, which provides quantitative explanations on how each layer of a model maintains the information of input words in a sample. We model the intra- and inter-word information at each layer measuring the importance of a word to the final prediction as well as the relationships between words, such as the formation of phrases. A multi-level visualization, which consists of a corpus-level, a sample-level, and a word-level visualization, supports the analysis from the overall training set to individual samples. Two case studies on classification tasks and comparison between models demonstrate that DeepNLPVis can help users effectively identify potential problems caused by samples and model architectures and then make informed improvements. Zhen Li 0044, Xiting Wang, Weikai Yang, Jing Wu 0004, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Adversarial Language Games for Advanced Natural Language IntelligenceabstractWe study the problem of adversarial language games, in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research. Yuan Yao 0013, Haoxi Zhong, Zhengyan Zhang, Xu Han 0007, Xiaozhi Wang, Kai Zhang 0033, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
AAAI | 3 |
| 2021 | Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerabstractFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, Maosong Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu 0001, Yasheng Wang, Maosong Sun 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language RepresentationabstractAbstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE models cannot take full advantage of the abundant textual information. In this paper, we propose a unified model for Knowledge Embedding and Pre-trained LanguagERepresentation (KEPLER), which can not only better integrate factual knowledge into PLMs but also produce effective text-enhanced KE with the strong PLMs. In KEPLER, we encode textual entity descriptions with a PLM as their embeddings, and then jointly optimize the KE and language modeling objectives. Experimental results show that KEPLER achieves state-of-the-art performances on various NLP tasks, and also works remarkably well as an inductive KE model on KG link prediction. Furthermore, for pre-training and evaluating KEPLER, we construct Wikidata5M1 , a large-scale KG dataset with aligned entity descriptions, and benchmark state-of-the-art KE methods on it. It shall serve as a new KE benchmark and facilitate the research on large KG, inductive KE, and KG with text. The source code can be obtained from https://github.com/THU-KEG/KEPLER. Xiaozhi Wang, Tianyu Gao 0001, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu 0001, Juan-Zi Li, Jian Tang 0005 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | CSS-LM: A Contrastive Framework for Semi-Supervised Fine-Tuning of Pre-Trained Language ModelsabstractFine-tuning pre-trained language models (PLMs) has demonstrated its effectiveness on various downstream NLP tasks recently. However, in many scenarios with limited supervised data, the conventional fine-tuning strategies cannot sufficiently capture the important semantic features for downstream tasks. To address this issue, we introduce a novel framework (named ‘`CSS-LM’') to improve the fine-tuning phase of PLMs via contrastive semi-supervised learning. Specifically, given a specific task, we retrieve positive and negative instances from large-scale unlabeled corpora according to their domain-level and class-level semantic relatedness to the task. We then perform contrastive semi-supervised learning on both the retrieved unlabeled instances and original labeled instances to help PLMs capture crucial task-related semantic features. The experimental results show that CSS-LM achieves better results than the conventional fine-tuning strategy on a series of downstream tasks with few-shot settings by up to 7.8%, and outperforms the latest supervised contrastive fine-tuning strategy by up to 7.1%. Our datasets and source code will be available to provide more details. Yusheng Su, Xu Han 0007, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Train No Evil: Selective Masking for Task-Guided Pre-TrainingabstractRecently, pre-trained language models mostly follow the pre-train-then-fine-tuning paradigm and have achieved great performance on various downstream tasks.However, since the pretraining stage is typically task-agnostic and the fine-tuning stage usually suffers from insufficient supervised data, the models cannot always well capture the domain-specific and task-specific patterns.In this paper, we propose a three-stage framework by adding a task-guided pre-training stage with selective masking between general pre-training and finetuning.In this stage, the model is trained by masked language modeling on in-domain unsupervised data to learn domain-specific patterns and we propose a novel selective masking strategy to learn task-specific patterns.Specifically, we design a method to measure the importance of each token in sequences and selectively mask the important tokens.Experimental results on two sentiment analysis tasks show that our method can achieve comparable or even better performance with less than 50% of computation cost, which indicates our method is both effective and efficient.The source code of this paper can be obtained from https://github. com/thunlp/SelectiveMasking. Yuxian Gu, Zhengyan Zhang, Xiaozhi Wang, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP (1) | 2 |
| 2019 | ERNIE: Enhanced Language Representation with Informative EntitiesabstractNeural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding.We argue that informative entities in KGs can enhance language representation with external knowledge.In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously.The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks.The source code and experiment details of this paper can be obtained from https:// github.com/thunlp/ERNIE. Zhengyan Zhang, Xu Han 0007, Zhiyuan Liu 0001, Xin Jiang 0002, Maosong Sun 0001, Qun Liu 0001 |
ACL (1) | 1 |
| 2019 | A Unified Framework for Community Detection and Network Representation LearningabstractNetwork representation learning (NRL) aims to learn low-dimensional vectors for vertices in a network. Most existing NRL methods focus on learning representations from local context of vertices (such as their neighbors). Nevertheless, vertices in many complex networks also exhibit significant global patterns widely known as communities. It's intuitive that vertices in the same community tend to connect densely and share common attributes. These patterns are expected to improve NRL and benefit relevant evaluation tasks, such as link prediction and vertex classification. Inspired by the analogy between network representation learning and text modeling, we propose a unified NRL framework by introducing community information of vertices, named as Community-enhanced Network Representation Learning (CNRL). CNRL simultaneously detects community distribution of each vertex and learns embeddings of both vertices and communities. Moreover, the proposed community enhancement mechanism can be applied to various existing NRL models. In experiments, we evaluate our model on vertex classification, link prediction, and community detection using several real-world datasets. The results demonstrate that CNRL significantly and consistently outperforms other state-of-the-art methods while verifying our assumptions on the correlations between vertices and communities. Cunchao Tu, Xiangkai Zeng, Hao Wang 0214, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001, Bo Zhang 0056, Leyu Lin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2017 | TransNet: Translation-Based Network Representation Learning for Social Relation ExtractionabstractConventional network representation learning (NRL) models learn low-dimensional vertex representations by simply regarding each edge as a binary or continuous value. However, there exists rich semantic information on edges and the interactions between vertices usually preserve distinct meanings, which are largely neglected by most existing NRL models. In this work, we present a novel Translation-based NRL model, TransNet, by regarding the interactions between vertices as a translation operation. Moreover, we formalize the task of Social Relation Extraction (SRE) to evaluate the capability of NRL methods on modeling the relations between vertices. Experimental results on SRE demonstrate that TransNet significantly outperforms other baseline methods by 10% to 20% on hits@1. The source code and datasets can be obtained from https://github.com/thunlp/TransNet. Cunchao Tu, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001 |
IJCAI | 2 |
| 2014 | Privacy-area aware dummy generation algorithms for Location-Based ServicesabstractLocation-Based Services (LBSs) have been one of the most popular activities in our daily life. Users can send queries to the LBS server easily to learn their surroundings. However, these location-related queries may result in serious privacy concerns since the un-trusted LBS server has all the information about users and may track them in various ways. In this paper, we propose two dummy-based solutions to achieve k-anonymity for privacy-area aware users in LBSs with considering that side information may be exploited by adversaries. We first choose some candidates based on a virtual circle or grid method, then blur these candidates into the final positions of dummy locations based on the entropy-based privacy metric. Security analysis and evaluation results indicate that the V-circle solution can significantly improve the privacy anonymity level. The V-grid solution can further enlarge the cloaking region while keeping similar privacy level. Ben Niu 0001, Zhengyan Zhang, Xiaoqing Li 0001, Hui Li 0006 |
ICC | 2 |