Zecheng Tang

dblp:326/0272 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context Reasoning
abstract
High-quality multi-hop instruction data is critical for enhancing the reasoning capabilities of large language models (LLMs) in complex long-context scenarios, e.g., long-form reasoning. Nevertheless, there is currently a notable scarcity of such datasets within the community, and existing data synthesis approaches typically fail to provide explicit modeling of intermediate reasoning steps, resulting in unverifiable and potentially erroneous samples. To mitigate above issue, we design the Concept-Graph based Multi-hop Instructions Synthesis (CGMIS) framework, which constructs long-form reasoning paths via concept graph traversal and automatically generates verifiable multi-hop data. The CGMIS framework not only guarantees the accuracy and verifiability of the synthesized data but also enables the construction of high-quality multi-hop instruction datasets from arbitrary corpora. Experiments show that fine-tuning with CGMIS-generated data achieves state-of-the-art performance across 13 long-context reasoning tasks on various models, using only 10% of the data volume required by existing methods.
Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Qiaoming Zhu
AAAI2
2026 DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-Reward
abstract
Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Zhe Zhao, Kehai Chen, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Kehai Chen, Juntao Li 0005, Min Zhang 0005
ACL (1)5
2026 IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking
abstract
Zechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang 0005
ACL (1)3
2026 MicroSSNet: an R package for microbial network construction and analysis at the single-sample and aggregated levels
abstract
BACKGROUND: Network analysis is a fundamental tool for elucidating microbial interactions, which are crucial for understanding the mechanisms that shape ecosystem structure and function. However, aggregated co-abundance/co-occurrence network approaches that infer pairwise relationships among biological entities from large sample collections often overlook sample-specific interaction patterns. To address this limitation, we developed MicroSSNet, an R package designed for analyzing microbial networks, including both aggregated and single-sample networks. RESULTS: We designed MicroSSNet primarily to fill the current gap in bioinformatics tools for constructing single-sample networks (SSNs) from microbiome data, and we evaluated both the performance and limitations of ssPCC-based SSNs using simulated and real datasets. Through Monte Carlo simulations, we assessed the statistical behavior of ssPCC and highlighted scenarios in which ssPCC is less powerful. We then applied MicroSSNet to two distinct datasets: a human gut metagenomic dataset and a soil 16S rRNA gene dataset. In the human gut dataset, SSNs revealed unique edges not detected in the aggregated network. In the soil dataset, SSN features showed some predictive value for group classification. However, SSN-derived patterns should be interpreted cautiously, as they may not exclusively reflect true interaction changes. MicroSSNet additionally implements a full aggregated-network workflow, including bipartite networks and extensive topological property analysis. CONCLUSIONS: Together, MicroSSNet offers a framework for constructing and analyzing both single-sample and aggregated microbial networks. In this work, we also highlight the potential and limitations of single-sample network approaches, supporting their application as exploratory tools in microbiome research across individual and population levels. The package is freely available on GitHub ( https://github.com/TangZecheng622/MicroSSNet ).
Zecheng Tang, Daohua Zhuang, Xinmin Duan, Qingqing Gong, Chen Tian 0019, Peicheng Jiang, Jiangkun Yu, Fangfang Zhao, Guolin Shi, Qinghang Du, Zhiqiang Ye
BMC Bioinform.1
2026 Data Foundations of Long-Context Language Models: A Survey
Zechen Sun, Zhaochen Su, Zecheng Tang, Juntao Li 0005, Wenliang Chen, Min Zhang 0005
Trans. Assoc. Comput. Linguistics4
2025 L-CiteEval: A Suite for Evaluating Fidelity of Long-context Models
abstract
Long-context models (LCMs) have witnessed remarkable advancements in recent years, facilitating real-world tasks like long-document QA.The success of LCMs is founded on the hypothesis that the model demonstrates strong fidelity, enabling it to respond based on the provided long context rather than relying solely on the intrinsic knowledge acquired during pre-training.Yet, in this paper, we find that open-sourced LCMs are not as faithful as expected.We introduce L-CiteEval, an out-of-the-box suite that can assess both generation quality and fidelity in long-context understanding tasks.It covers 11 tasks with context lengths ranging from 8K to 48K and a corresponding automatic evaluation pipeline.Evaluation of 11 cutting-edge closed-source and open-source LCMs indicates that, while there are minor differences in their generation, open-source models significantly lag behind closed-source counterparts in terms of fidelity.Furthermore, we analyze the benefits of citation generation for LCMs from both the perspective of explicit model output and the internal attention mechanism 1 .
Zecheng Tang, Keyan Zhou, Juntao Li 0005, Baibei Ji, Jianye Hou, Min Zhang 0005
ACL (1)1
2025 Enhancing Cryptocurrency Trading Strategies: A Deep Reinforcement Learning Approach Integrating Multi-Source LLM Sentiment Analysis
abstract
Recent advancements in large language models (LLMs) have demonstrated their potential to significantly impact finance trading, particularly through sentiment analysis. The cryptocurrency market, known for its volatility and unpredictability, often renders price-based trading approaches inadequate. This necessitates the adoption of more sophisticated techniques such as market sentiment analysis, which can benefit from the insights provided by LLMs. This study introduces an innovative method that integrates sentiment analysis derived from five distinct LLMs with deep reinforcement learning to devise a cryptocurrency trading strategy. Recognizing that LLM outputs cannot be guaranteed to be infallibly accurate, which contributing to the LLM hallucinations, this paper details the implementation of a stringent outlier detection and removal process. By adopting a “Trust-The-Majority” strategy, the research aims to ensure that trading decisions are informed by reliable sentiment data. In addition, sentiment scores are traditionally timestamped to the publication of news or social media posts. To more accurately reflect the actual impact of such information on market sentiment, this study applies the Ebbinghaus Forgetting Curve to model the waning influence of information over time. This allows for a more nuanced understanding of how news affects market dynamics. The enhanced sentiment scores, in conjunction with traditional market data such as OHLCV (Open, High, Low, Close, Volume), are utilized by a deep reinforcement learning model to make trading decisions. Experimental results demonstrate that the proposed multi-LLM sentiment-driven framework improves trading performance in the fast-paced cryptocurrency market. The methodology outlined in this paper offers a solid foundation for incorporating real-time market sentiment analysis into financial applications.
Nanjiang Du, Yida Zhao, Yicheng Zhu, Siyu Xie, Luyao Yang, Yiru Tong, Shengzhe Xu, Wangying Zhang, Zecheng Tang, Jianfeng Ren, Tianxiang Cui
CIFEr10
2025 OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking
abstract
With the rapid growth of generative AI and its widespread application in image editing, new risks have emerged regarding the authenticity and integrity of digital content. Existing versatile watermarking approaches suffer from tradeoffs between tamper localization precision and visual quality. Constrained by the limited flexibility of previous framework, their localized watermark must remain fixed across all images. Under AIGC-editing, their copyright extraction accuracy is also unsatisfactory. To address these challenges, we propose OmniGuard, a novel augmented versatile watermarking approach that integrates proactive embedding with passive, blind extraction for robust copyright protection and tamper localization. OmniGuard employs a hybrid forensic framework that enables flexible localization watermark selection and introduces a degradation-aware tamper extraction network for precise localization under challenging conditions. Additionally, a lightweight AIGC-editing simulation layer is designed to enhance robustness across global and local editing. Extensive experiments show that OmniGuard achieves superior fidelity, robustness, and flexibility. Compared to the recent state-of-the-art approach EditGuard, our method outperforms it by 4.25dB in PSNR of the container image, 20.7% in F1-Score under noisy conditions, and 14.8% in average bit accuracy.
Xuanyu Zhang 0003, Zecheng Tang, Zhipei Xu, Runyi Li, Youmin Xu, Bin Chen 0006, Jian Zhang 0018
CVPR2
2025 Revealing and Mitigating Over-Attention in Knowledge Editing
abstract
Large Language Models~(LLMs) have demonstrated superior performance across a wide range of tasks, but they still exhibit undesirable errors due to incorrect knowledge learned from the training data. To avoid this, knowledge editing methods emerged to precisely edit the specific model knowledge via efficiently modifying a very small percentage of parameters. However, those methods can lead to the problem of **Specificity Failure**, where the existing knowledge and capabilities are severely degraded due to editing. Our preliminary indicates that Specificity Failure primarily stems from the model's attention heads assigning excessive attention scores to entities related to the edited knowledge, thereby unduly focusing on specific snippets within the context, which we denote as the **Attention Drift** phenomenon. To mitigate such Attention Drift issue, we introduce a simple yet effective method **S**elective **A**ttention **D**rift **R**estriction(**SADR**), which introduces an additional regularization term during the knowledge editing process to restrict changes in the attention weight distribution, thereby preventing undue focus on the edited entity. Experiments on five frequently-used strong LLMs demonstrate the effectiveness of our method, where SADR can significantly mitigate Specificity Failure in the predominant knowledge editing tasks.
Pinzheng Wang, Zecheng Tang, Keyan Zhou, Juntao Li 0005, Qiaoming Zhu, Min Zhang 0005
ICLR2
2025 FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
abstract
The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. Although current image forgery detection and localization (IFDL) methods are generally effective, they tend to face two challenges: \textbf{1)} black-box nature with unknown detection principle, \textbf{2)} limited generalization across diverse tampering methods (e.g., Photoshop, DeepFake, AIGC-Editing). To address these issues, we propose the explainable IFDL task and design FakeShield, a multi-modal framework capable of evaluating image authenticity, generating tampered region masks, and providing a judgment basis based on pixel-level and image-level tampering clues. Additionally, we leverage GPT-4o to enhance existing IFDL datasets, creating the Multi-Modal Tamper Description dataSet (MMTD-Set) for training FakeShield's tampering analysis capabilities. Meanwhile, we incorporate a Domain Tag-guided Explainable Forgery Detection Module (DTE-FDM) and a Multi-modal Forgery Localization Module (MFLM) to address various types of tamper detection interpretation and achieve forgery localization guided by detailed textual descriptions. Extensive experiments demonstrate that FakeShield effectively detects and localizes various tampering techniques, offering an explainable and superior solution compared to previous IFDL methods. The code is available at https://github.com/zhipeixu/FakeShield.
Zhipei Xu, Xuanyu Zhang 0003, Runyi Li, Zecheng Tang, Jian Zhang 0018
ICLR4
2025 LOGO - Long cOntext aliGnment via efficient preference Optimization
abstract
Long-context models (LCMs) have shown great potential in processing long input sequences (even more than 100M tokens) conveniently and effectively. With significant progress, recent research has pointed out that LCMs can accurately locate token-level salient information within the context. Yet, the generation performance of these LCMs is far from satisfactory and might result in misaligned responses, such as hallucinations. To enhance the generation capability of LCMs, existing works have investigated the effects of data size and quality for both pre-training and instruction tuning. Though achieving meaningful improvement, previous methods fall short in either effectiveness or efficiency. In this paper, we introduce LOGO (Long cOntext aliGnment via efficient preference Optimization), a training strategy that first introduces preference optimization for long-context alignment. To overcome the GPU memory-bound issue caused by the long sequence, LOGO employs a reference-free preference optimization strategy and adopts a position synthesis method to construct the training data. By training with only 0.3B data on a single 8 x A800 GPU machine for 16 hours, LOGO allows the Llama-3-8B-Instruct-80K model to achieve comparable performance with GPT-4 in real-world long-context tasks while preserving the model's original capabilities on other tasks, e.g., language modeling and MMLU. Moreover, LOGO can extend the model's context window size while enhancing its generation performance.
Zecheng Tang, Zechen Sun, Juntao Li 0005, Qiaoming Zhu, Min Zhang 0005
ICML1
2025 Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
abstract
Large language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding. However, recent studies indicate that even the best models lack true comprehension of their reasoning processes. In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models. We design a $\textit{\textbf{C}ritic-\textbf{D}iscernment \textbf{G}ame}~(\textbf{CDG})$ in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution. These critiques either aim to assist or mislead the prover. The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback. Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process.
Pinzheng Wang, Juntao Li 0005, Zecheng Tang, Haijia Gui, Min Zhang 0005
ICML3
2025 OpenBA: an open-sourced 15B bilingual asymmetric Seq2Seq model pre-trained from scratch
Juntao Li 0005, Zecheng Tang, Yuyang Ding, Pinzheng Wang, Pei Guo, Wangjie You, Wenliang Chen, Guohong Fu, Qiaoming Zhu, Guodong Zhou 0001, Min Zhang 0005
Sci. China Inf. Sci.2
2025 Robust Reconstructed Neural Network With Spectral Reshaping Activation
abstract
Neural network (NN) is a prominent intelligent model to process information through the connection and activation of multilayer neurons. However, NNs usually encounter with the incorrect activation of neurons because of the excessive coverage for the boundary of compound noises. To address this issue, this article proposes a robust reconstructed NN (RRNN) with spectral reshaping activation (SRA). Primarily, an SRA is designed to replace the original activation of NN, which shrinks the spectrums of the compound noises toward the cluster center through spectral subtraction. It enables RRNN to reshape a concentrated noise space for easy coverage. Then, a hierarchical gradient descent (HGD) algorithm is developed to update the parameters of RRNN. The HGD algorithm establishes a noise-contrastive degree of SRA to penalize the loss function of RRNN, which holds robust performance with different noises. Furthermore, the theoretical proof of RRNN is presented to validate its robustness. Finally, the experimental results confirm the superior robustness of RRNN for tackling noisy samples compared to other methods.
Honggui Han, Zecheng Tang, Hongyan Yang 0001, Junfei Qiao 0001
IEEE Trans. Cybern.2
2025 Robust Reconstructed Neural Network Based on Spectral Elastic Activation
abstract
The presence of sparse samples poses a formidable challenge for neural networks (NNs) in shaping representative patterns due to the limited coverage of concentrative activation range in NNs. To address this issue, a robust reconstructed NN, based on spectral elastic activation (SEA-RRNN), is developed in this article. Primarily, a spectral elastic activation (SEA) is designed to broaden the original activation of NNs, which embeds a spectral increment scaled by the estimated outlier degree on the activation boundary. It enables SEA-RRNN to stretch the boundary of SEA to cover the features of sparse samples. Then, an adaptive robust gradient descent (ARGD) algorithm is introduced to update the parameters of SEA. By tuning the loss between error and correntropy with the estimated outlier degree, the ARGD algorithm establishes two loss functions with complementary distance sensitivity for different parameters of SEA, which alleviates the construction conflict of robust center and precise boundary in SEA-RRNN. Furthermore, the theoretical analysis of SEA-RRNN is provided to validate its convergence and robustness. Finally, the experimental results demonstrate that SEA-RRNN exhibits superior robustness compared to other NN models.
Zecheng Tang, Honggui Han
IEEE Trans. Neural Networks Learn. Syst.1
2024 CMD: a framework for Context-aware Model self-Detoxification
abstract
Text detoxification aims to minimize the risk of language models producing toxic content.However, existing detoxification methods fail to balance the detoxification effectiveness and generation quality.This issue arises from neglecting the constraints imposed by the context: language models are designed to generate output that closely matches the given context, while detoxification methods strive to ensure the safety of the output, even if it deviates semantically from the context.Given this, we introduce a Context-aware Model self-Detoxification (CMD) framework that pays attention to both the context and the detoxification process, i.e., first detoxifying the context and then making the language model generate along the safe context.Specifically, CMD framework involves two phases: utilizing language models to synthesize data and applying these data for training.We also introduce a toxic contrastive loss that encourages the model generation away from the negative toxic samples.Experiments on various LLMs have verified the effectiveness of our MSD framework, which can yield the best performance compared to baselines. 1 Warning: cases in this paper may contain offensive content.
Zecheng Tang, Keyan Zhou, Juntao Li 0005, Yuyang Ding, Pinzheng Wang, Yan Bowen, Renjie Hua, Min Zhang 0005
EMNLP1
2024 LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language Models
abstract
Graphic layout generation, a growing research field, plays a significant role in user engagement and information perception. Existing methods primarily treat layout generation as a numerical optimization task, focusing on quantitative aspects while overlooking the semantic information of layout, such as the relationship between each layout element. In this paper, we propose LayoutNUWA, the first model that treats layout generation as a code generation task to enhance semantic information and harness the hidden layout expertise of large language models~(LLMs). Concretely, we develop a Code Instruct Tuning (CIT) approach comprising three interconnected modules: 1) the Code Initialization (CI) module quantifies the numerical conditions and initializes them as HTML code with strategically placed masks; 2) the Code Completion (CC) module employs the formatting knowledge of LLMs to fill in the masked portions within the HTML code; 3) the Code Rendering (CR) module transforms the completed code into the final layout output, ensuring a highly interpretable and transparent layout generation procedure that directly maps code to a visualized layout. We attain significant state-of-the-art performance (even over 50\% improvements compared to previous works) on multiple datasets, showcasing the strong capabilities of LayoutNUWA.
Zecheng Tang, Chenfei Wu, Nan Duan 0001
ICLR1
2024 StrokeNUWA - Tokenizing Strokes for Vector Graphic Synthesis
abstract
To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model’s ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation "stroke" tokens on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a $94\times$ speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%.
Zecheng Tang, Chenfei Wu, Minheng Ni, Shengming Yin, Zhengyuan Yang, Zicheng Liu 0001, Nan Duan 0001
ICML1
2024 Towards Better Chinese Spelling Check for Search Engines: A New Dataset and Strong Baseline
abstract
Misspellings in search engine queries may prevent search engines from returning accurate results. For Chinese mobile search engines, due to the different input methods (e.g., hand-written and T9 input methods), more types of misspellings exist, making this problem more challenging. As an essential module of search engines, Chinese Spelling Check~(CSC) models aim to detect and correct misspelled Chinese characters from user-issued queries. Despite the great value of CSC to the search engine, there is no CSC benchmark collected from real-world search engine queries. To fill this blank, we construct and release the Alipay Search Engine Query (AlipaySEQ) spelling check dataset. To the best of our knowledge, AlipaySEQ is the first Chinese Spelling Check dataset collected from the real-world scenario of Chinese mobile search engines. It consists of 15,522 high-quality human annotated and 1,175,151 automatically generated samples. To demonstrate the unique challenges of AlipaySEQ in the era of Large Language Models~(LLMs), we conduct a thorough study to analyze the difference between AlipaySEQ and existing SIGHAN benchmarks and compare the performance of various baselines, including existing task-specific methods and LLMs. We observe that all baselines fail to perform satisfactorily due to the over-correction problem. Especially, LLMs exhibit below-par performance on AlipaySEQ, which is rather surprising. Therefore, to alleviate the over-correction problem, we introduce a model-agnostic CSC Self-Refine Framework (SRF) to construct a strong baseline. Comprehensive experiments demonstrate that our proposed SRF, though more effective against existing models on both the AlipaySEQ and SIGHAN15, is still far from achieving satisfactory performance on our real-world dataset. With the newly collected real-world dataset and strong baseline, we hope more progress can be achieved on such a challenging and valuable task.
Yue Wang 0039, Zilong Zheng, Zecheng Tang, Juntao Li 0005, Kunlong Chen, Jinxiong Chang, Qishen Zhang, Zhongyi Liu 0001, Min Zhang 0005
WSDM3
2024 Self-organizing broad network with frequency-domain analysis
Honggui Han, Zecheng Tang, Hongyan Yang 0001, Junfei Qiao 0001
Eng. Appl. Artif. Intell.2
2024 Robust Modeling for Industrial Process Based on Frequency Reconstructed Fuzzy Neural Network
abstract
The model bias caused by input outliers is a dramatic obstacle to the application of models in industrial processes. To cope with this problem, this article proposes a robust modeling method based on frequency reconstructed fuzzy neural network (FRFNN) for industrial process. The robust modeling consists of two parts: One is feature extraction, where a Fourier-based filter is developed with input data denoising. It enables the model to suppress high-frequency input noises and burst outliers. The other one is feature representation that is realized with a FRFNN. The soft margins of membership functions of FRFNN are designed with Fourier estimation of outliers, which have the capability of outlier-tolerant for filtered-residual outliers. Moreover, an adaptive gradient descent algorithm is introduced to update the model parameters. Based on the adaptive learning rate decaying with outliers, this algorithm is insensitive to the bias effect of outliers and also maintains convergence. Finally, the proposed robust modeling method is tested on two real-world industrial datasets with input outliers. The experimental results demonstrate that the proposed robust modeling method can strengthen robustness and achieve superior performance over other previous methods.
Honggui Han, Zecheng Tang, Hongyan Yang 0001, Junfei Qiao 0001
IEEE Trans. Fuzzy Syst.2
2024 Time-Aware Fuzzy Neural Network Based on Frequency-Enhanced Modulation Mechanism
abstract
Fuzzy neural network (FNN) is regarded as a prominent approach in application of time-series modeling. With the capability of fuzzy reasoning, FNN can capture temporal patterns from the time-series samples. However, the existing FNNs may suffer from the temporal pattern distortion because possibly multiscale features cannot be explored sufficiently. To address this problem, a time-aware fuzzy neural network, based on the frequency-enhanced modulation mechanism (FEM-TAFNN), is developed for time-series prediction in this article. First, a Fourier-based decoder is established to extract the multiscale features. This decoder employs the frequency-domain model to orthogonally separate the time-scale features with different frequencies into independent temporal patterns based on the Fourier basis, which prevents the overlap of temporal patterns using time-domain analysis. Second, a frequency-enhanced modulation mechanism is designed to shape fuzzy rules of FNN based on the contribution of different temporal patterns in the frequency spectrum. It enables FEM-TAFNN to modulate out the realistic multiscale temporal patterns. Finally, the proposed FEM-TAFNN is tested on four multiscale time-series datasets. The empirical results confirm its superior prediction performance than other methods.
Honggui Han, Zecheng Tang, Hongyan Yang 0001, Junfei Qiao 0001
IEEE Trans. Fuzzy Syst.2
2023 Open-ended Long Text Generation via Masked Language Modeling
abstract
Pre-trained autoregressive (AR) language models such as BART and GPTs have dominated Open-ended Long Text Generation (Open-LTG).However, the AR nature will decrease the inference efficiency along with the increase of generation length, which hinder their application in Open-LTG.To improve inference efficiency, we alternatively explore the potential of the pre-trained masked language models (MLMs) along with a representative iterative non-autoregressive (NAR) decoding strategy for Open-LTG.Our preliminary study shows that pre-trained MLMs can merely generate short text and will collapse for long text modeling.To enhance the long text generation capability of MLMs, we introduce two simple yet effective strategies for the iterative NAR model: dynamic sliding window attention (DSWA) and linear temperature decay (LTD).It can alleviate long-distance collapse problems and achieve longer text generation with a flexible trade-off between performance and inference speedup.Experiments on the storytelling and multi-paragraph opinionated article writing tasks show that pre-trained MLMs can achieve more than 3 × → 13 × speedup with better performance than strong AR models.Our code is available at GitHub * .
Xiaobo Liang, Zecheng Tang, Juntao Li 0005, Min Zhang 0005
ACL (1)2
2023 Beyond Hard Samples: Robust and Effective Grammatical Error Correction with Cycle Self-Augmenting
Kaiqi Feng, Zecheng Tang, Juntao Li 0005, Min Zhang 0005
NLPCC (2)2
2022 Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change
abstract
Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., language model pre-trained on static data from past years performs worse over time on emerging data.Existing methods mainly perform continual training to mitigate such a misalignment.While effective to some extent but is far from being addressed on both the language modeling and downstream tasks.In this paper, we empirically observe that temporal generalization is closely affiliated with lexical semantic change, which is one of the essential phenomena of natural languages.Based on this observation, we propose a simple yet effective lexical-level masking strategy to post-train a converged language model.Experiments on two pre-trained language models, two different classification tasks, and four benchmark datasets demonstrate the effectiveness of our proposed method over existing temporal adaptation methods, i.e., continual training with new data.Our code is available at https: //github.com/zhaochen0110/LMLM.
Zhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu 0003, Min Zhang 0005, Juntao Li 0005
EMNLP2
2022 General High-Pass Convolution: A Novel Convolutional Layer for Image Manipulation Detection
Zecheng Tang, Yang Liu 0132
PRCV (1)1