Yazheng Yang

dblp:222/9478 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-1627-8341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cognitive Alpha Mining via LLM-Driven Code-Based Evolution
abstract
Fengyuan Liu, Yi Huang, Sichun Luo, Yuqi Wang, Yazheng Yang, Xinye Li, Zefa Hu, Junlan Feng, Qi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yi Huang 0017, Sichun Luo, Yazheng Yang, Zefa Hu, Junlan Feng
ACL (1)5
2026 Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science With Table-Specific Pretraining
abstract
In data science, predictive tasks such as classification, regression, and missing value imputation are fundamental challenges in tabular data analysis. This research investigates the application of Large Language Models (LLMs) to these tasks. While LLMs excel in natural language understanding, their effectiveness on structured tabular data remains limited due to minimal exposure during pretraining. To address this gap, we construct a large-scale corpus of annotated tables and introduce a tailored pretraining framework. Our trained model achieves significant improvements over baselines, with an average gain of 8.9% in classification and 10.7% in regression tasks. We further evaluate its performance in zero-shot and few-shot prediction, as well as in-context learning scenarios. Extensive experiments demonstrate substantial gains over existing benchmarks, highlighting the potential of LLMs for tabular data processing. Additionally, we apply our approach across multiple open-source LLMs and demonstrate its generalizability. This work establishes a new benchmark for enhancing tabular intelligence through LLM-based pretraining.
Yazheng Yang, Yuqi Wang 0003, Yaxuan Li 0002, Sankalok Sen, Lei Li 0039, Qi Liu 0049
IEEE Trans. Knowl. Data Eng.1
2024 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
abstract
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049
EMNLP7
2024 Improving Long Text Understanding with Knowledge Distilled from Summarization Model
abstract
Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its gist and sometimes can be regarded as noise. With recent advances of abstractive summarization, we propose our Gist Detector to leverage the gist detection ability of a summarization model and integrate the extracted gist into downstream models to enhance their long text understanding ability. Specifically, Gist Detector first learns the gist detection knowledge distilled from a summarization model, and then produces gist-aware representations to augment downstream models. We evaluate our method on three different tasks: long document classification, distantly supervised open-domain question answering, and non-parallel text style transfer. The experimental results show that our method can significantly improve the performance of baseline models on all tasks.
Yazheng Yang, Xiaokang Chen
ICASSP2
2024 UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
abstract
Recent advancements in Natural Language Processing (NLP) have witnessed the groundbreaking impact of pretrained models, yielding impressive outcomes across various tasks. This study seeks to extend the power of pretraining methodologies to facilitating the prediction over tables in data science, a domain traditionally overlooked, yet inherently challenging due to the plethora of table schemas intrinsic to different tasks. The primary research questions underpinning this work revolve around the establishment of a universal pretraining protocol for tables with varied structures, the generalizability and transferability of learned knowledge across tasks, the adaptation to diverse downstream applications, and the incorporation of incremental columns over time. In response to these challenges, we introduce UniTabE, a straightforward yet effective method designed to process tables in a uniform manner, devoid of constraints imposed by specific table structures. UniTabE's core concept relies on representing each basic table element with a module, termed TabUnit. This is subsequently followed by a Transformer encoder to refine the representation. Moreover, our model is designed to facilitate pretraining and finetuning through the utilization of free-form prompts. In order to implement the pretraining phase, we curated an expansive tabular dataset comprising approximately 13 billion samples, meticulously gathered from the Kaggle platform. This research primarily centers on classification and regression tasks involving tabular data, and conducts rigorous experimental testing and analyses to validate the effectiveness of our methodology. The experimental results demonstrate UniTabE's superior performance against several baseline models across a multitude of benchmark datasets. This, therefore, underscores UniTabE's potential to significantly enhance the semantic representation of tabular data, thereby marking a significant stride for tabular data analysis.
Yazheng Yang, Yuqi Wang 0003, Guang Liu 0006, Ledell Wu, Qi Liu 0049
ICLR1
2023 MSSRNet: Manipulating Sequential Style Representation for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer task aims to rewrite a text into target style while preserving its main content. Traditional methods rely on the use of a fixed-sized vector to regulate text style, which is difficult to accurately convey the style strength for each individual token. In fact, each token of a text contains different style intensity and makes different contribution to the overall style. Our proposed method addresses this issue by assigning individual style vector to each token in a text, allowing for fine-grained control and manipulation of the style strength. Additionally, an adversarial training framework integrated with teacher-student learning is introduced to enhance training stability and reduce the complexity of high-dimensional optimization. The results of our experiments demonstrate the efficacy of our method in terms of clearly improved style transfer accuracy and content preservation in both two-style transfer and multi-style transfer settings.
Yazheng Yang, Zhou Zhao 0001, Qi Liu 0049
KDD1
2022 MPII: Multi-Level Mutual Promotion for Inference and Interpretation
abstract
In order to better understand the rationale behind model behavior, recent works have exploited providing interpretation to support the inference prediction.However, existing methods tend to provide human-unfriendly interpretation, and are prone to sub-optimal performance due to one-side promotion, i.e. either inference promotion with interpretation or vice versa.In this paper, we propose a multi-level Mutual Promotion mechanism for self-evolved Inference and sentence-level Interpretation (MPII).Specifically, from the model-level, we propose a Step-wise Integration Mechanism to jointly perform and deeply integrate inference and interpretation in an autoregressive manner.From the optimizationlevel, we propose an Adversarial Fidelity Regularization to improve the fidelity between inference and interpretation with the Adversarial Mutual Information training strategy.Extensive experiments on NLI and CQA tasks reveal that the proposed MPII approach can significantly outperform baseline models for both the inference performance and the interpretation quality. 1
Sanyuan Chen, Yazheng Yang, Qi Dai 0001
ACL (1)3
2022 GIFT: Graph-guIded Feature Transfer for Cold-Start Video Click-Through Rate Prediction
abstract
Short video has witnessed rapid growth in the past few years in e-commerce platforms like Taobao. To ensure the freshness of the content, platforms need to release a large number of new videos every day, making conventional click-through rate (CTR) prediction methods suffer from the item cold-start problem. In this paper, we propose GIFT, an efficient Graph-guIded Feature Transfer system, to fully take advantages of the rich information of warmed-up videos to compensate for the cold-start ones. Specifically, we establish a heterogeneous graph that contains physical and semantic linkages to guide the feature transfer process from warmed-up video to cold-start videos.Specifically, we establish a heterogeneous graph that contains physical and semantic linkages to guide the feature transfer process. The physical linkages consist of the explicit relationships (e.g., produced by the same author, or showcasing the same product etc.), and the semantic linkages measure the proximity of multi-modal representations of two videos. We elaborately design the feature transfer function to make aware of different parts of transferred features (e.g., id representations and historical statistics) from different types of nodes and edges along the metapath on the graph. We conduct extensive experiments on a large real-world dataset, and the results show that our GIFT system outperforms SOTA methods significantly and brings a 6.82% lift on CTR in the homepage of Taobao App.
Sihao Hu, Zhao Li 0007, Yazheng Yang, Qingwen Liu 0002, Shouling Ji
CIKM5
2021 TopNet: Learning from Neural Topic Model to Generate Long Stories
abstract
Long story generation (LSG) is one of the coveted goals in natural language processing. Different from most text generation tasks, LSG requires to output a long story of rich content based on a much shorter text input, and often suffers from information sparsity. In this paper, we propose TopNet to alleviate this problem, by leveraging the recent advances in neural topic modeling to obtain high-quality skeleton words to complement the short input. In particular, instead of directly generating a story, we first learn to map the short text input to a low-dimensional topic distribution (which is pre-assigned by a topic model). Based on this latent topic distribution, we can use the reconstruction decoder of the topic model to sample a sequence of inter-related words as a skeleton for the story. Experiments on two benchmark datasets show that our proposed framework is highly effective in skeleton word selection and significantly outperforms the state-of-the-art models in both automatic evaluation and human evaluation.
Yazheng Yang, Boyuan Pan, Deng Cai 0001, Huan Sun 0001
KDD1
2021 Self-supervised attention flow for dialogue state tracking
Boyuan Pan, Yazheng Yang, Bo Li 0026, Deng Cai 0001
Neurocomputing2
2020 Adversarial Mutual Information for Text Generation
abstract
Recent advances in maximizing mutual information (MI) between the source and target have demonstrated its effectiveness in text generation. However, previous works paid little attention to modeling the backward network of MI (i.e., dependency from the target to the source), which is crucial to the tightness of the variational information maximization lower bound. In this paper, we propose Adversarial Mutual Information (AMI): a text generation framework which is formed as a novel saddle point (min-max) optimization aiming to identify joint interactions between the source and target. Within this framework, the forward and backward networks are able to iteratively promote or demote each other’s generated instances by comparing the real and synthetic data distributions. We also develop a latent noise sampling strategy that leverages random variations at the high-level semantic space to enhance the long term dependency in the generation process. Extensive experiments based on different text generation tasks demonstrate that the proposed AMI framework can significantly outperform several strong baselines, and we also show that AMI has potential to lead to a tighter lower bound of maximum mutual information for the variational information maximization problem.
Boyuan Pan, Yazheng Yang, Kaizhao Liang, Bhavya Kailkhura, Zhongming Jin 0001, Xian-Sheng Hua 0001, Deng Cai 0001, Bo Li 0026
ICML2
2020 Bi-Decoder Augmented Network for Neural Machine Translation
Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001
Neurocomputing2
2018 Discourse Marker Augmented Network with Reinforcement Learning for Natural Language Inference
abstract
Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is one of the most important problems in natural language processing.It requires to infer the logical relationship between two given sentences.While current approaches mostly focus on the interaction architectures of the sentences, in this paper, we propose to transfer knowledge from some important discourse markers to augment the quality of the NLI model.We observe that people usually use some discourse markers such as "so" or "but" to represent the logical relationship between two sentences.These words potentially have deep connections with the meanings of the sentences, thus can be utilized to help improve the representations of them.Moreover, we use reinforcement learning to optimize a new objective function with a reward defined by the property of the NLI datasets to make full use of the labels information.Experiments show that our method achieves the state-of-the-art performance on several large-scale datasets.
Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001
ACL (1)2
2018 MacNet: Transferring Knowledge from Machine Comprehension to Sequence-to-Sequence Models
abstract
Machine Comprehension (MC) is one of the core problems in natural language processing, requiring both understanding of the natural language and knowledge about the world. Rapid progress has been made since the release of several benchmark datasets, and recently the state-of-the-art models even surpass human performance on the well-known SQuAD evaluation. In this paper, we transfer knowledge learned from machine comprehension to the sequence-to-sequence tasks to deepen the understanding of the text. We propose MacNet: a novel encoder-decoder supplementary architecture to the widely used attention-based sequence-to-sequence models. Experiments on neural machine translation (NMT) and abstractive text summarization show that our proposed framework can significantly improve the performance of the baseline models, and our method for the abstractive text summarization achieves the state-of-the-art results on the Gigaword dataset.
Boyuan Pan, Yazheng Yang, Hao Li 0009, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001
NeurIPS2