Fan Yang 0087

dblp:29/3081-87 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-7113-3300ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
abstract
The use of denoising diffusion models is becoming increasingly popular in the field of image editing. However, current approaches often rely on either image-guided methods, which provide a visual reference but lack control over semantic consistency, or text-guided methods, which ensure alignment with the text guidance but compromise visual quality. To resolve this issue, we propose a framework that integrates a fusion of generated visual references and text guidance into the semantic latent space of a frozen pre-trained diffusion model. Using only a tiny neural network, our framework provides control over diverse content and attributes, driven intuitively by the simple prompt. Compared to state-of-the-art methods, the framework generates images of higher quality while providing realistic editing effects across various benchmark datasets. The code is available at https://github.com/SadAngelF/Editing-via-Step-Wise-Alignment.
Zhanbo Feng, Zenan Ling, Ci Gong, Feng Zhou 0011, Wugedele Bao, Jie Li 0002, Fan Yang 0087, Robert C. Qiu
ICASSP8
2025 Robust and Communication-Efficient Federated Domain Adaptation via Random Features
abstract
Modern machine learning (ML) models have grown to a scale where training them on a single machine becomes impractical. As a result, there is a growing trend to leverage federated learning (FL) techniques to train large ML models in a distributed and collaborative manner. These models, however, when deployed on new devices, might struggle to generalize well due to domain shifts. In this context, federated domain adaptation (FDA) emerges as a powerful approach to address this challenge. Most existing FDA approaches typically focus on aligning the distributions between source and target domains by minimizing their (e.g., MMD) distance. Such strategies, however, inevitably introduce high communication overheads and can be highly sensitive to network reliability. In this paper, we introduce RF-TCA, an enhancement to the standard Transfer Component Analysis approach that significantly accelerates computation without compromising theoretical and empirical performance. Leveraging the computational advantage of RF-TCA, we further extend it to FDA setting with FedRF-TCA. The proposed FedRF-TCA protocol boasts communication complexity that isindependentof the sample size, while maintaining performance that is either comparable to or even surpasses state-of-the-art FDA methods. We present extensive experiments to showcase the superior performance and robustness (to network condition) of FedRF-TCA.
Zhanbo Feng, Yuanjie Wang, Jie Li 0002, Fan Yang 0087, Jiong Lou, Tiebin Mi, Robert C. Qiu, Zhenyu Liao 0001
IEEE Trans. Knowl. Data Eng.4
2023 Ambiguous Learning from Retrieval: Towards Zero-shot Semantic Parsing
abstract
Shan Wu, Chunlei Xin, Hongyu Lin, Xianpei Han, Cao Liu, Jiansong Chen, Fan Yang, Guanglu Wan, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chunlei Xin, Xianpei Han, Cao Liu, Jiansong Chen, Fan Yang 0087, Guanglu Wan, Le Sun 0001
ACL (1)7
2022 Confidence Calibration for Intent Detection via Hyperspherical Space and Rebalanced Accuracy-Uncertainty Loss
abstract
Data-driven methods have achieved notable performance on intent detection, which is a task to comprehend user queries. Nonetheless, they are controversial for over-confident predictions. In some scenarios, users do not only care about the accuracy but also the confidence of model. Unfortunately, mainstream neural networks are poorly calibrated, with a large gap between accuracy and confidence. To handle this problem defined as confidence calibration, we propose a model using the hyperspherical space and rebalanced accuracy-uncertainty loss. Specifically, we project the label vector onto hyperspherical space uniformly to generate a dense label representation matrix, which mitigates over-confident predictions due to overfitting sparse one-hot label matrix. Besides, we rebalance samples of different accuracy and uncertainty to better guide model training. Experiments on the open datasets verify that our model outperforms the existing calibration methods and achieves a significant improvement on the calibration metric.
Yantao Gong, Cao Liu, Fan Yang 0087, Guanglu Wan, Jiansong Chen, Houfeng Wang
AAAI3
2022 DESED: Dialogue-based Explanation for Sentence-level Event Detection
abstract
Many recent sentence-level event detection efforts focus on enriching sentence semantics, e.g., via multi-task or prompt-based learning. Despite the promising performance, these methods commonly depend on label-extensive manual annotations or require domain expertise to design sophisticated templates and rules. This paper proposes a new paradigm, named dialogue-based explanation, to enhance sentence semantics for event detection. By saying dialogue-based explanation of an event, we mean explaining it through a consistent information-intensive dialogue, with the original event description as the start utterance. We propose three simple dialogue generation methods, whose outputs are then fed into a hybrid attention mechanism to characterize the complementary event semantics. Extensive experimental results on two event detection datasets verify the effectiveness of our method and suggest promising research opportunities in the dialogue-based explanation paradigm.
Yinyi Wei, Shuaipeng Liu, Jianwei Lv, Xiangyu Xi, Hailei Yan, Wei Ye 0004, Tong Mo, Fan Yang 0087, Guanglu Wan
COLING8
2022 MUSIED: A Benchmark for Event Detection from Multi-Source Heterogeneous Informal Texts
abstract
Event detection (ED) identifies and classifies event triggers from unstructured texts, serving as a fundamental task for information extraction.Despite the remarkable progress achieved in the past several years, most research efforts focus on detecting events from formal texts (e.g., news articles, Wikipedia documents, financial announcements).Moreover, the texts in each dataset are either from a single source or multiple yet relatively homogeneous sources.With massive amounts of user-generated text accumulating on the Web and inside enterprises, identifying meaningful events in these informal texts, usually from multiple heterogeneous sources, has become a problem of significant practical value.As a pioneering exploration that expands event detection to the scenarios involving informal and heterogeneous texts, we propose a new large-scale Chinese event detection dataset based on user reviews, text conversations, and phone conversations in a leading e-commerce platform for food service.We carefully investigate the proposed dataset's textual informality and multi-source heterogeneity characteristics by inspecting data samples quantitatively and qualitatively.Extensive experiments with state-of-the-art event detection methods verify the unique challenges posed by these characteristics, indicating that multisource informal event detection remains an open problem and requires further efforts.Our benchmark and code are released at https: //github.com/myeclipse/MUSIED.
Xiangyu Xi, Jianwei Lv, Shuaipeng Liu, Wei Ye 0004, Fan Yang 0087, Guanglu Wan
EMNLP5
2022 A Region-based Document VQA
abstract
Practical Document Visual Question Answering (DocVQA) needs not only to recognize and extract the document contents, but also reason on them for answering questions. However, previous DocVQA data mainly focuses on in-line questions, where the answers could be directly extracted after locating keywords in the documents, which needs less reasoning. This paper therefore builds a large-scale dataset named Region-based Document VQA (RDVQA), which includes more practical questions for DocVQA. We then propose a novel Reason-over-In-region-Question-answering (ReIQ) model for addressing the problems. It is a pre-training-based model, where a Spatial-Token Pre-trained Model (STPM) is employed as the backbone. Two novel pre-training tasks, Masked Text Box Regression and Shuffled Triplet Reconstruction, are proposed to learn the entailment relationship between text blocks and tokens as well as contextual information, respectively. Moreover, a DocVQA State Tracking Module (DocST) is also proposed to track the DocVQA state in the fine-tuning stage. Experimental results show that our model improves the performance onRDVQA significantly, although more work should be done for practical DocVQA as shown inRDVQA.
Xinya Wu, Duo Zheng, Jiashen Sun, Minzhen Hu, Fangxiang Feng, Xiaojie Wang 0006, Huixing Jiang, Fan Yang 0087
ACM Multimedia9
2022 A Low-Cost, Controllable and Interpretable Task-Oriented Chatbot: With Real-World After-Sale Services as Example
abstract
Though widely used in industry, traditional task-oriented dialogue systems suffer from three bottlenecks: (i) difficult ontology construction (e.g., intents and slots); (ii) poor controllability and interpretability; (iii) annotation-hungry. In this paper, we propose to represent utterance with a simpler concept named Dialogue Action, upon which we construct a tree-structured TaskFlow and further build task-oriented chatbot with TaskFlow as core component. A framework is presented to automatically construct TaskFlow from large-scale dialogues and deploy online. Our experiments on real-world after-sale customer services show TaskFlow can satisfy the major needs, as well as reduce the developer burden effectively.
Xiangyu Xi, Chenxu Lv, Yuncheng Hua, Wei Ye 0004, Chaobo Sun, Shuaipeng Liu, Fan Yang 0087, Guanglu Wan
SIGIR7
2022 Dialogue Topic Segmentation via Parallel Extraction Network with Neighbor Smoothing
abstract
Dialogue topic segmentation is a challenging task in which dialogues are split into segments with pre-defined topics. Existing works on topic segmentation adopt a two-stage paradigm, including text segmentation and segment labeling. However, such methods tend to focus on the local context in segmentation, and the inter-segment dependency is not well captured. Besides, the ambiguity and labeling noise in dialogue segment bounds bring further challenges to existing models. In this work, we propose the Parallel Extraction Network with Neighbor Smoothing (PEN-NS) to address the above issues. Specifically, we propose the parallel extraction network to perform segment extractions, optimizing the bipartite matching cost of segments to capture inter-segment dependency. Furthermore, we propose neighbor smoothing to handle the segment-bound noise and ambiguity. Experiments on a dialogue-based and a document-based topic segmentation dataset show that PEN-NS outperforms state-the-of-art models significantly.
Jinxiong Xia, Cao Liu, Jiansong Chen, Fan Yang 0087, Guanglu Wan, Houfeng Wang
SIGIR5
2021 From Paraphrasing to Semantic Parsing: Unsupervised Semantic Parsing via Synchronous Semantic Decoding
abstract
Shan Wu, Bo Chen, Chunlei Xin, Xianpei Han, Le Sun, Weipeng Zhang, Jiansong Chen, Fan Yang, Xunliang Cai. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bo Chen 0020, Chunlei Xin, Xianpei Han, Le Sun 0001, Jiansong Chen, Fan Yang 0087
ACL/IJCNLP (1)8
2021 Density-Based Dynamic Curriculum Learning for Intent Detection
abstract
Pre-trained language models have achieved noticeable performance on the intent detection task. However, due to assigning an identical weight to each sample, they suffer from the overfitting of simple samples and the failure to learn complex samples well. To handle this problem, we propose a density-based dynamic curriculum learning model. Our model defines the sample's difficulty level according to their eigenvectors' density. In this way, we exploit the overall distribution of all samples' eigenvectors simultaneously. Then we apply a dynamic curriculum learning strategy, which pays distinct attention to samples of various difficulty levels and alters the proportion of samples during the training process. Through the above operation, simple samples are well-trained, and complex samples are enhanced. Experiments on three open datasets verify that the proposed density-based algorithm can distinguish simple and complex samples significantly. Besides, our model obtains obvious improvement over the strong baselines.
Yantao Gong, Cao Liu, Jiazhen Yuan, Fan Yang 0087, Guanglu Wan, Jiansong Chen, Ruiyao Niu, Houfeng Wang
CIKM4
2021 Domain-Lifelong Learning for Dialogue State Tracking via Knowledge Preservation Networks
abstract
Dialogue state tracking (DST), which estimates user goals given a dialogue context, is an essential component of task-oriented dialogue systems.Conventional DST models are usually trained offline, which requires a fixed dataset prepared in advance.This paradigm is often impractical in real-world applications since online dialogue systems usually involve continually emerging new data and domains.Therefore, this paper explores Domain-Lifelong Learning for Dialogue State Tracking (DLL-DST), which aims to continually train a DST model on new data to learn incessantly emerging new domains while avoiding catastrophically forgetting old learned domains.To this end, we propose a novel domainlifelong learning method, called Knowledge Preservation Networks (KPN), which consists of multi-prototype enhanced retrospection and multi-strategy knowledge distillation, to solve the problems of expression diversity and combinatorial explosion in the DLL-DST task.Experimental results show that KPN effectively alleviates catastrophic forgetting and outperforms previous state-of-the-art lifelong learning methods by 4.25% and 8.27% of whole joint goal accuracy on the MultiWOZ benchmark and the SGD benchmark, respectively.
Qingbin Liu, Cao Liu, Jiansong Chen, Fan Yang 0087, Shizhu He, Kang Liu 0001, Jun Zhao 0001
EMNLP (1)6
2021 Distant Supervision based Machine Reading Comprehension for Extractive Summarization in Customer Service
abstract
Given a long text, the summarization system aims to obtain a shorter highlight while keeping important information on the original text. For customer service, the summaries of most dialogues between an agent and a user focus on several fixed key points, such as user's question, user's purpose, the agent's solution, and so on. Traditional extractive methods are difficult to extract all predefined key points exactly. Furthermore, there is a lack of large-scale and high-quality extractive summarization datasets containing key points. In order to solve the above challenges, we propose a Distant Supervision based Machine Reading Comprehension model for extractive Summarization (DSMRC-S). DSMRC-S transforms the summarization task into the machine reading comprehension problem, to fetch key points from the original text exactly according to the predefined questions. In addition, a distant supervision method is proposed to alleviate the lack of eligible extractive summarization datasets. We conduct experiments on a large-scale summarization dataset collected in customer service scenarios, and the results show that the proposed DSMRC-S outperforms the strong baseline methods by 4 points on ROUGE-L.
Cao Liu, Jingyu Wang 0001, Shujie Hu, Fan Yang 0087, Guanglu Wan, Jiansong Chen, Jianxin Liao
SIGIR5
2020 Logic-guided Semantic Representation Learning for Zero-Shot Relation Classification
abstract
Relation classification aims to extract semantic relations between entity pairs from the sentences.However, most existing methods can only identify seen relation classes that occurred during training.To recognize unseen relations at test time, we explore the problem of zero-shot relation classification.Previous work regards the problem as reading comprehension or textual entailment, which have to rely on artificial descriptive information to improve the understandability of relation types.Thus, rich semantic knowledge of the relation labels is ignored.In this paper, we propose a novel logic-guided semantic representation learning model for zero-shot relation classification.Our approach builds connections between seen and unseen relations via implicit and explicit semantic representations with knowledge graph embeddings and logic rules.Extensive experimental results demonstrate that our method can generalize to unseen relation types and achieve promising improvements.
Juan Li 0010, Ruoxu Wang, Ningyu Zhang 0001, Wen Zhang 0015, Fan Yang 0087, Huajun Chen
COLING5
2009 A Chinese-English Organization Name Translation System Using Heuristic Web Mining and Asymmetric Alignment
Fan Yang 0087, Jun Zhao 0001, Kang Liu 0001
ACL/IJCNLP1
2008 Chinese-English Backward Transliteration Assisted with Mining Monolingual Web Pages
Fan Yang 0087, Jun Zhao 0001, Bo Zou 0002, Kang Liu 0001
ACL1
2008 CRFs-Based Named Entity Recognition Incorporated with Heuristic Entity List Searching
Fan Yang 0087, Jun Zhao 0001, Bo Zou 0002
IJCNLP1