Feng Wang 0023

dblp:90/4225-23 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-6975-4335ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
abstract
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for improving the complex reasoning abilities of large language models (LLMs).However, current RLVR methods face two significant challenges: the near-miss reward problem, where a small mistake can invalidate an otherwise correct reasoning process, greatly hindering training efficiency; and exploration stagnation, where models tend to focus on solutions within their "comfort zone," lacking the motivation to explore potentially more effective alternatives.To address these challenges, we propose StepHint, a novel RLVR algorithm that utilizes multi-level stepwise hints to help models explore the solution space more effectively.StepHint partitions valid reasoning chains into reasoning steps using our proposed adaptive partitioning method.The initial few steps are used as hints, and simultaneously, multiple-level hints (each comprising a different number of steps) are provided to the model.This approach directs the model's exploration toward a promising solution subspace while preserving its flexibility for independent exploration.By providing hints, StepHint mitigates the near-miss reward problem, thereby improving training efficiency.Additionally, the external reasoning pathways help the model develop better reasoning abilities, enabling it to move beyond its "comfort zone" and mitigate exploration stagnation.StepHint outperforms competitive RLVR enhancement methods across six mathematical benchmarks and two out-of-domain benchmarks. 1
Ang Lv, Jinpeng Li 0003, Feng Wang 0023, Haoyuan Hu, Rui Yan 0001
ACL (1)5
2025 FinS-Pilot: A Benchmark for Online Financial RAG System
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across various professional domains, with their performance typically evaluated through standardized benchmarks. In the financial field, the stringent demands for professional accuracy and real-time data processing often necessitate the use of retrieval-augmented generation (RAG) techniques. However, the development of financial RAG benchmarks has been constrained by data confidentiality issues and the lack of dynamic data integration. To address this issue, we introduce FinS-Pilot, a novel benchmark for evaluating RAG systems in online financial applications. Constructed from real-world financial assistant interactions, our benchmark incorporates both real-time API data and text data, organized through an intent classification framework covering critical financial domains. The benchmark enables comprehensive evaluation of financial assistants' capabilities in handling both static knowledge and time-sensitive market information.Through systematic experiments with multiple Chinese leading LLMs, we demonstrate FinS-Pilot's effectiveness in identifying models suitable for financial applications while addressing the current gap in specialized evaluation tools for the financial domain. Our work contributes both a practical evaluation framework and a curated dataset to advance research in financial NLP systems. The code and dataset are accessible on GitHub.
Feng Wang 0023, Jiaxin Mao, Danqing Xu
CIKM1
2025 PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
abstract
Large language models (LLMs) enhanced with retrieval-augmented generation (RAG) have introduced a new paradigm for web search. However, the limited context awareness of LLMs degrades their performance on RAG tasks. Existing methods to enhance context awareness are often inefficient, incurring time or memory overhead during inference, and many are tailored to specific position embeddings. In this paper, we propose Position-Embedding-Agnostic attention Re-weighting (PEAR), which enhances the context awareness of LLMs with zero inference overhead. Specifically, on a proxy task focused on context copying, we first detect heads which suppress the models' context awareness, thereby diminishing RAG performance. To weaken the impact of these heads, we re-weight their outputs with learnable coefficients. The LLM (with frozen parameters) is optimized by adjusting these coefficients to minimize loss on the proxy task. During inference, the optimized coefficients are fixed to re-weight these heads, regardless of the specific task at hand. Our proposed PEAR offers two major advantages over previous approaches: (1) It introduces zero additional inference overhead in terms of memory usage or inference time, while outperforming competitive baselines in accuracy and efficiency across various RAG tasks. (2) It is independent of position embedding algorithms, ensuring broader applicability. Our code is available at https://github.com/TTArch/PEAR-RAG.
Tao Tan 0005, Ang Lv, Hongzhan Lin 0002, Songhao Wu, Feng Wang 0023, Jingtong Wu, Rui Yan 0001
WWW7
2025 Personalized Review Summarization by Using Graph-Based Retrieval Augmemted Generation
abstract
Review summarization aims to provide a summary that covers the main aspect of the product review and reflects personal preference. Existing methods employ the historical reviews of customer and product to provide useful clues for the target summary generation. However, most of the existing methods indiscriminately model the historical reviews of customer and product. Since the historicalcustomerreviews provide the personal information while the historicalproductreviews provide the commonly focused aspect of the product, these two types of heterogeneous information should be separately modeled. Moreover, the review rating of the historical reviews can be seen as a high-level abstraction of the customer preference and product which have been ignored by most of the existing methods. In this paper, we propose the Heterogeneous Historical Review aware Review Summarization (HHRRS) which separately models the two types of historical reviews with the rating information by a graph reasoning module with a contrastive loss. We employ a multi-task paradigm that conducts the review sentiment classification and summarization (GRARS) to model the two types of heterogeneous information in a fine-grained manner. We conduct extensive experiments on four benchmark datasets, and demonstrate the superiority of HHRRS on both tasks.
Shuo Shang, Xin Cheng 0002, Yiren Xiong, Shen Gao, Xiuying Chen, Feng Wang 0023, Dongyan Zhao 0001, Rui Yan 0001
IEEE Trans. Knowl. Data Eng.7
2025 Unified Multi-Scenario Summarization Evaluation and Explanation
abstract
Summarization quality evaluation is a non-trivial task in text summarization. Contemporary methods can be mainly categorized into two scenarios: (1)reference-based:evaluating with human-labeled reference summary; (2)reference-free:evaluating the summary consistency of the document. Recent studies mainly focus on one of these scenarios and explore training neural models to align with human criteria and finally give a numeric score. However, the models from different scenarios are optimized individually, which may result in sub-optimal performance since they neglect the shared knowledge across different scenarios. Besides, designing individual models for each scenario caused inconvenience to the user. Moreover, only providing the numeric quality evaluation score for users cannot help users to improve the summarization model, since they do not know why the score is low. Inspired by this, we proposeUnifiedMulti-scenarioSummarizationEvaluator (UMSE) andMulti-AgentSummarizationEvaluationExplainer (MASEE). More specifically, we propose a perturbed prefix tuning method to share cross-scenario knowledge between scenarios and use a self-supervised training paradigm to optimize the model without extra human labeling. Our UMSE is the first unified summarization evaluation framework engaged with the ability to be used in three evaluation scenarios. We propose a multi-agent summary evaluation explanation method MASEE, which employs several LLM-based agents to generate detailed natural language explanations in four different aspects. Experimental results across three typical scenarios on the benchmark dataset SummEval indicate that our UMSE can achieve comparable performance with several existing strong methods that are specifically designed for each scenario. And intensive quantitative and qualitative experiments also demonstrate the effectiveness of our proposed explanation method, which can generate consistent and accurate explanations.
Shuo Shang, Zhitao Yao, Chongyang Tao, Xiuying Chen, Feng Wang 0023, Zhaochun Ren, Shen Gao
IEEE Trans. Knowl. Data Eng.6
2024 An Integrated Data Processing Framework for Pretraining Foundation Models
abstract
The ability of the foundation models heavily relies on large-scale, diverse, and high-quality pretraining data. In order to improve data quality, researchers and practitioners often have to manually curate datasets from difference sources and develop dedicated data cleansing pipeline for each data repository. Lacking a unified data processing framework, this process is repetitive and cumbersome. To mitigate this issue, we propose a data processing framework that integrates a Processing Module which consists of a series of operators at different granularity levels, and an Analyzing Module which supports probing and evaluation of the refined data. The proposed framework is easy to use and highly flexible. In this demo paper, we first introduce how to use this framework with some example use cases and then demonstrate its effectiveness in improving the data quality with an automated evaluation with ChatGPT and an end-to-end evaluation in pretraining the GPT-2 model. The code and demonstration video are accessible on GitHub.
Feng Wang 0023, Yutao Zhu 0001, Wayne Xin Zhao, Jiaxin Mao
SIGIR2
2023 Fragility Index: A New Approach for Binary Classification
abstract
In binary classification problems, many performance metrics evaluate the probability that some error exceeds a threshold. Nevertheless, they focus more on the probability and fail to capture the magnitude of the error, which evaluates how large this error exceeds the threshold. Capturing the magnitude of error is desired in many applications. For example, in detecting disease and predicting credit default, the magnitude of error illustrates the confidence in making the wrong prediction. We propose a novel metric, the Fragility Index (FI), to evaluate the performance of binary classifiers by capturing the magnitude of the error. FI alleviates the risk of misclassification by penalizing the large error greatly, which is seldom considered by standard metrics. Moreover, to strengthen the generalization ability and handle unseen samples, we adopt the framework of distributionally robust optimization and robust satisficing, which allows us to derive and control the maximum degree of fragility of the classifier when the distribution of samples shifts. We show that FI can be easily calculated and optimized for common probabilistic distance measures. Experiments with real datasets demonstrate the new insights brought by FI and the advantages of classifiers selected under FI, which always improve the robustness and reduce the risk of large errors as compared to classifiers selected by alternative metrics.
Chen Yang 0025, Bo Cao 0007, Daniel Zhuoyu Long, Feng Wang 0023, Ruohan Zhan
KDD9
2022 Keywords and Instances: A Hierarchical Contrastive Learning Framework Unifying Hybrid Granularities for Text Generation
abstract
Mingzhe Li, XieXiong Lin, Xiuying Chen, Jinxiong Chang, Qishen Zhang, Feng Wang, Taifeng Wang, Zhongyi Liu, Wei Chu, Dongyan Zhao, Rui Yan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Mingzhe Li 0001, Xiexiong Lin, Xiuying Chen, Jinxiong Chang, Qishen Zhang, Feng Wang 0023, Taifeng Wang, Zhongyi Liu 0001, Dongyan Zhao 0001, Rui Yan 0001
ACL (1)6
2019 Adapting Translation Models for Transcript Disfluency Detection
abstract
Transcript disfluency detection (TDD) is an important component of the real-time speech translation system, which arouses more and more interests in recent years. This paper presents our study on adapting neural machine translation (NMT) models for TDD. We propose a general training framework for adapting NMT models to TDD task rapidly. In this framework, the main structure of the model is implemented similar to the NMT model. Additionally, several extended modules and training techniques which are independent of the NMT model are proposed to improve the performance, such as the constrained decoding, denoising autoencoder initialization and a TDD-specific training object. With the proposed training framework, we achieve significant improvement. However, it is too slow in decoding to be practical. To build a feasible and production-ready solution for TDD, we propose a fast non-autoregressive TDD model following the non-autoregressive NMT model emerged recently. Even we do not assume the specific architecture of the NMT model, we build our TDD model on the basis of Transformer, which is the state-of-the-art NMT model. We conduct extensive experiments on the publicly available set, Switchboard, and in-house Chinese set. Experimental results show that the proposed model significantly outperforms previous state-ofthe-art models.
Qianqian Dong, Feng Wang 0023, Zhen Yang 0007, Wei Chen 0048, Bo Xu 0002
AAAI2
2019 Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping
abstract
Self-attention network, an attention-based feedforward neural network, has recently shown the potential to replace recurrent neural networks (RNNs) in a variety of NLP tasks. However, it is not clear if the self-attention network could be a good alternative of RNNs in automatic speech recognition (ASR), which processes the longer speech sequences and may have online recognition requirements. In this paper, we present a RNN-free end-to-end model: self-attention aligner (SAA), which applies the self-attention networks to a simplified recurrent neural aligner (RNA) framework. We also propose a chunk-hopping mechanism, which enables the SAA model to encode on segmented frame chunks one after another to support online recognition. Experiments on two Mandarin ASR datasets show the replacement of RNNs by the self-attention networks yields a 8.4%-10.2% relative character error rate (CER) reduction. In addition, the chunk-hopping mechanism allows the SAA to have only a 2.5% relative CER degradation with a 320ms latency. After jointly training with a self-attention network language model, our SAA model obtains further error rate reduction on multiple datasets. Especially, it achieves 24.12% CER on the Mandarin ASR benchmark (HKUST), exceeding the best end-to-end model by over 2% absolute CER.
Linhao Dong, Feng Wang 0023, Bo Xu 0002
ICASSP2
2019 Hybrid Attention for Chinese Character-Level Neural Machine Translation
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
Neurocomputing1
2019 Effectively training neural machine translation models with monolingual data
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
Neurocomputing3
2018 Unsupervised Neural Machine Translation with Weight Sharing
abstract
Unsupervised neural machine translation (NMT) is a recently proposed approach for machine translation which aims to train the model without using any labeled data.The models proposed for unsupervised NMT often use only one shared encoder to map the pairs of sentences from different languages to a shared-latent space, which is weak in keeping the unique and internal characteristics of each language, such as the style, terminology, and sentence structure.To address this issue, we introduce an extension by utilizing two independent encoders but sharing some partial weights which are responsible for extracting high-level representations of the input sentences.Besides, two different generative adversarial networks (GANs), namely the local GAN and global GAN, are proposed to enhance the cross-language translation.With this new approach, we achieve significant improvements on English-German, English-French and Chinese-to-English translation tasks.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
ACL (1)3
2018 Semi-Supervised Disfluency Detection
abstract
While the disfluency detection has achieved notable success in the past years, it still severely suffers from the data scarcity. To tackle this problem, we propose a novel semi-supervised approach which can utilize large amounts of unlabelled data. In this work, a light-weight neural net is proposed to extract the hidden features based solely on self-attention without any Recurrent Neural Network (RNN) or Convolutional Neural Network (CNN). In addition, we use the unlabelled corpus to enhance the performance. Besides, the Generative Adversarial Network (GAN) training is applied to enforce the similar distribution between the labelled and unlabelled data. The experimental results show that our approach achieves significant improvements over strong baselines.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Qianqian Dong, Bo Xu 0002
COLING1
2018 Cascaded Mutual Modulation for Visual Reasoning
abstract
Visual reasoning is a special visual question answering problem that is multi-step and compositional by nature, and also requires intensive text-vision interactions.We propose CMM: Cascaded Mutual Modulation as a novel end-to-end visual reasoning model.CMM includes a multi-step comprehension process for both question and image.In each step, we use a Feature-wise Linear Modulation (FiLM) technique to enable textual/visual pipeline to mutually control each other.Experiments show that CMM significantly outperforms most related models, and reach stateof-the-arts on two visual reasoning benchmarks: CLEVR and NLVR, collected from both synthetic and natural languages.Ablation studies confirm that both our multistep framework and our visual-guided language modulation are critical to the task.Our code is available at https://github. com/FlamingHorizon/CMM-VR.
Yiqun Yao, Jiaming Xu 0001, Feng Wang 0023, Bo Xu 0002
EMNLP3
2018 Self-Attention Based Network for Punctuation Restoration
abstract
Inserting proper punctuation into Automatic Speech Recognizer(ASR) transcription is a challenging and promising task in real-time Spoken Language Translation(SLT). Traditional methods built on the sequence labelling framework are weak in handling the joint punctuation. To tackle this problem, we propose a novel self-attention based network, which can solve the aforementioned problem very well. In this work, a light-weight neural net is proposed to extract the hidden features based solely on self-attention without any Recurrent Neural Nets(RNN) and Convolutional Neural Nets(CNN). We conduct extensive experiments on complex punctuation tasks. The experimental results show that the proposed model achieves significant improvements on joint punctuation task while being superior to traditional methods on simple punctuation task as well.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
ICPR1
2018 Unsupervised Domain Adaptation for Neural Machine Translation
abstract
Impressive neural machine translation (NMT) results are achieved in domains with large-scale, high quality bilingual training corpora. However, transferring to a target domain with significant domain shifts but no bilingual training corpora remains largely unexplored. To address the aforementioned setting of unsupervised domain adaptation, we propose a novel adversarial training procedure for NMT to leverage the widespread monolingual data in target domain. Two discriminative networks, namely the domain discriminator and pair discriminator, are introduced to guide the translation model. The domain discriminator evaluates whether the sentences generated by the translation model are indistinguishable from the ones in target domain. The pair discriminator assesses whether the generated sentences are paired with the source-side sentences. The translation model acts as an adversary to the two discriminators, which aims to generate sentences uneasily discriminated by the discriminators. We tested our approach on Chinese-English and English-German translation tasks. Experimental results show that our approaches achieve great success in unsupervised domain adaptation for NMT.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
ICPR3
2018 Improving Neural Machine Translation with Conditional Sequence Generative Adversarial Nets
abstract
Zhen Yang, Wei Chen, Feng Wang, Bo Xu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
NAACL-HLT3
2018 Generative adversarial training for neural machine translation
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
Neurocomputing3
2017 Joint Extraction of Entities and Relations Based on a Novel Tagging Scheme
abstract
Joint extraction of entities and relations is an important task in information extraction.To tackle this problem, we firstly propose a novel tagging scheme that can convert the joint extraction task to a tagging problem.Then, based on our tagging scheme, we study different end-toend models to extract entities and their relations directly, without identifying entities and relations separately.We conduct experiments on a public dataset produced by distant supervision method and the experimental results show that the tagging based methods are better than most of the existing pipelined and joint learning methods.What's more, the end-to-end model proposed in this paper, achieves the best results on the public dataset.
Suncong Zheng, Feng Wang 0023, Hongyun Bao, Yuexing Hao, Peng Zhou 0009, Bo Xu 0002
ACL (1)2
2017 Towards Compact and Fast Neural Machine Translation Using a Combined Method
abstract
Neural Machine Translation (NMT) lays intensive burden on computation and memory cost.It is a challenge to deploy NMT models on the devices with limited computation and memory budgets.This paper presents a four stage pipeline to compress model and speed up the decoding for NMT.Our method first introduces a compact architecture based on convolutional encoder and weight shared embeddings.Then weight pruning is applied to obtain a sparse model.Next, we propose a fast sequence interpolation approach which enables the greedy decoding to achieve performance on par with the beam search.Hence, the time-consuming beam search can be replaced by simple greedy decoding.Finally, vocabulary selection is used to reduce the computation of softmax layer.Our final model achieves 10× speedup, 17× parameters reduction, <35MB storage size and comparable performance compared to the baseline model.
Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
EMNLP3
2017 A class-specific copy network for handling the rare word problem in neural machine translation
abstract
Neural machine translation (NMT) has shown promising results and rapidly gained adoption in many large-scale settings. With the NMT model being widely used in empirical productions, its long-standing weakness in handling the rare and out of vocabulary words has been amplified a lot. In order to release the model from the stress of “understanding” the rare words, copy mechanism has been proposed to deal with the rare and unseen words for the neural network models using attention. However the negative side of the copy mechanism is that the model is only able to decide whether to copy or not. It is unable to detect which class should the rare word be copied to, such as person, location, and organization. This paper deeply investigates this limitation of the NMT model. As a result, we propose a new NMT model by novelly incorporating a class-specific copy network. With the network, the proposed NMT model is able to decide which class the words in the target belong to and which class in the source should be copied to. Experimental results on Chinese-English translation tasks show that the proposed model outperforms the traditional NMT model with a large margin especially for sentences containing the rare words.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
IJCNN1
2017 Multi-sense based neural machine translation
abstract
Attention mechanism advances the neural machine translation (NMT) by reducing the confusion introduced by irrelevant words in long sentences. However, the confusion caused by ambiguous words hasn't been handled yet and it may be a bottleneck for the NMT model. This paper validates the hypothesis and proposes a simple and flexible framework, which enables the NMT model to only focus on the relevant sense type of the input word in current context. Experiments show that the proposed model achieves substantial improvements on every test set over competitive baselines. Our contributions come from twofold. Firstly, to the best of our knowledge, this is the first effort to introduce the multi-sense representation, which represents each sense type of the word with a sense-specific embedding, into NMT. Secondly, We propose a sense search module which can detect the sense type of the word automatically. Flexibility and versatility are the most attractive characteristic of the proposed sense search module. It can be applied to any other semantic related NLP tasks with little modification.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
IJCNN3
2016 A Character-Aware Encoder for Neural Machine Translation
abstract
This article proposes a novel character-aware neural machine translation (NMT) model that views the input sequences as sequences of characters rather than words. On the use of row convolution (Amodei et al., 2015), the encoder of the proposed model composes word-level information from the input sequences of characters automatically. Since our model doesn’t rely on the boundaries between each word (as the whitespace boundaries in English), it is also applied to languages without explicit word segmentations (like Chinese). Experimental results on Chinese-English translation tasks show that the proposed character-aware NMT model can achieve comparable translation performance with the traditional word based NMT models. Despite the target side is still word based, the proposed model is able to generate much less unknown words.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
COLING3