Bin Sun 0004

dblp:01/5401-4 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 MODE+: A benchmark and a probe into multimodal open-domain dialogue evaluation
Hang Yin 0007, Xinglin Wang, Pinren Lu, Bin Sun 0004, Peiwen Yuan, Kan Li 0001
Neurocomputing5
2025 Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
abstract
Direction reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remains underexplored. To address this, we propose the Compass Direction Reasoning (CDR) benchmark, designed to evaluate the direction reasoning capabilities of multimodal language models (MLMs). CDR includes three types images to test spatial (up, down, left, right) and compass (north, south, east, west) directions. Our evaluation reveals that most MLMs struggle with direction reasoning, often performing at random guessing levels. Experiments show that training directly with CDR data yields limited improvements, as it requires an understanding of real-world physical rules. We explore the impact of mixdata and CoT fine-tuning methods, which significantly enhance MLM performance in compass direction reasoning by incorporating diverse data and step-by-step reasoning, improving the model’s ability to understand direction relationships.
Hang Yin 0007, Zhifeng Lin, Xin Liu 0132, Bin Sun 0004, Kan Li 0001
ICASSP4
2024 Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data
abstract
Large Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In some cases, however, LLMs may produce incorrect reasoning chains, especially when facing complex mathematical problems. Previous studies only transfer knowledge from positive samples and drop the synthesized data with wrong answers. In this work, we illustrate the merit of negative data and propose a model specialization framework to distill LLMs with negative samples besides positive ones. The framework consists of three progressive steps, covering from training to inference stages, to absorb knowledge from negative data. We conduct extensive experiments across arithmetic reasoning tasks to demonstrate the role of negative data in distillation from LLM.
Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Bin Sun 0004, Xinglin Wang, Heda Wang, Kan Li 0001
AAAI5
2024 Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
abstract
Self-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost method that requires multiple sampling with the preset size. In this paper, we propose a simple and scalable sampling process, Early-Stopping Self-Consistency (ESC), to greatly reduce the cost of SC without sacrificing performance. On this basis, one control scheme for ESC is further derivated to dynamically choose the performance-cost balance for different tasks and models. To demonstrate ESC's effectiveness, we conducted extensive experiments on three popular categories of reasoning tasks: arithmetic, commonsense and symbolic reasoning over language models with varying scales. The empirical results show that ESC reduces the average number of sampling of chain-of-thought reasoning by a significant margin on six benchmarks, including MATH (-33.8%), GSM8K (-80.1%), StrategyQA (-76.8%), CommonsenseQA (-78.5%), Coin Flip (-84.2%) and Last Letters (-67.4%), while attaining comparable performances.
Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun 0004, Heda Wang, Kan Li 0001
ICLR6
2024 MODE: a multimodal open-domain dialogue dataset with explanation
Hang Yin 0007, Pinren Lu, Bin Sun 0004, Kan Li 0001
Appl. Intell.4
2023 Heterogeneous-Branch Collaborative Learning for Dialogue Generation
abstract
With the development of deep learning, advanced dialogue generation methods usually require a greater amount of computational resources. One promising approach to obtaining a high-performance and lightweight model is knowledge distillation, which relies heavily on the pre-trained powerful teacher. Collaborative learning, also known as online knowledge distillation, is an effective way to conduct one-stage group distillation in the absence of a well-trained large teacher model. However, previous work has a severe branch homogeneity problem due to the same training objective and the independent identical training sets. To alleviate this problem, we consider the dialogue attributes in the training of network branches. Each branch learns the attribute-related features based on the selected subset. Furthermore, we propose a dual group-based knowledge distillation method, consisting of positive distillation and negative distillation, to further diversify the features of different branches in a steadily and interpretable way. The proposed approach significantly improves branch heterogeneity and outperforms state-of-the-art collaborative learning methods on two widely used open-domain dialogue datasets.
Yiwei Li 0001, Shaoxiong Feng, Bin Sun 0004, Kan Li 0001
AAAI3
2023 Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent Variables
abstract
Conditional variational models, using either continuous or discrete latent variables, are powerful for open-domain dialogue response generation. However, previous works show that continuous latent variables tend to reduce the coherence of generated responses. In this paper, we also found that discrete latent variables have difficulty capturing more diverse expressions. To tackle these problems, we combine the merits of both continuous and discrete latent variables and propose a Hybrid Latent Variable (HLV) method. Specifically, HLV constrains the global semantics of responses through discrete latent variables and enriches responses with continuous latent variables. Thus, we diversify the generated responses while maintaining relevance and coherence. In addition, we propose Conditional Hybrid Variational Transformer (CHVT) to construct and to utilize HLV with transformers for dialogue generation. Through fine-grained symbolic-level semantic information and additive Gaussian mixing, we construct the distribution of continuous variables, prompting the generation of diverse expressions. Meanwhile, to maintain the relevance and coherence, the discrete latent variable is optimized by self-separation training. Experimental results on two dialogue generation datasets (DailyDialog and Opensubtitles) show that CHVT is superior to traditional transformer-based variational mechanism w.r.t. diversity, relevance and coherence metrics. Moreover, we also demonstrate the benefit of applying HLV to fine-tuning two pre-trained dialogue models (PLATO and BART-base).
Bin Sun 0004, Fei Mi, Weichao Wang, Yiwei Li 0001, Kan Li 0001
AAAI1
2023 Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue Evaluation
abstract
Turn-level dialogue evaluation models (TDEMs), using self-supervised learning (SSL) framework, have achieved state-of-the-art performance in open-domain dialogue evaluation. However, these models inevitably face two potential problems. First, they have low correlations with humans on medium coherence samples as the SSL framework often brings training data with unbalanced coherence distribution. Second, the SSL framework leads TDEM to nonuniform score distribution. There is a danger that the nonuniform score distribution will weaken the robustness of TDEM through our theoretical analysis. To tackle these problems, we propose Better Correlation and Robustness (BCR), a distribution-balanced self-supervised learning framework for TDEM. Given a dialogue dataset, BCR offers an effective training set reconstructing method to provide coherence-balanced training signals and further facilitate balanced evaluating abilities of TDEM. To get a uniform score distribution, a novel loss function is proposed, which can adjust adaptively according to the uniformity of score distribution estimated by kernel density estimation. Comprehensive experiments on 17 benchmark datasets show that vanilla BERT-base using BCR outperforms SOTA methods significantly by 11.3% on average. BCR also demonstrates strong generalization ability as it can lead multiple SOTA methods to attain better correlation and robustness.
Peiwen Yuan, Xinglin Wang, Bin Sun 0004, Yiwei Li 0001
NeurIPS4
2023 Stop filtering: Multi-view attribute-enhanced dialogue learning
Yiwei Li 0001, Bin Sun 0004, Shaoxiong Feng, Kan Li 0001
Knowl. Based Syst.2
2022 Diversifying Neural Dialogue Generation via Negative Distillation
abstract
Generative dialogue models suffer badly from the generic response problem, limiting their applications to a few toy scenarios.Recently, an interesting approach, namely negative training, has been proposed to alleviate this problem by reminding the model not to generate high-frequency responses during training.However, its performance is hindered by two issues, ignoring low-frequency but generic responses and bringing low-frequency but meaningless responses.In this paper, we propose a novel negative training paradigm, called negative distillation, to keep the model away from the undesirable generic responses while avoiding the above problems.First, we introduce a negative teacher model that can produce querywise generic responses, and then the student model is required to maximize the distance with multi-level negative knowledge.Empirical results show that our method outperforms previous negative training methods significantly. 1
Yiwei Li 0001, Shaoxiong Feng, Bin Sun 0004, Kan Li 0001
NAACL-HLT3
2022 THINK: A novel conversation model for generating grammatically correct and coherent responses
Bin Sun 0004, Shaoxiong Feng, Yiwei Li 0001, Jiamou Liu, Kan Li 0001
Knowl. Based Syst.1
2021 Generating Relevant and Coherent Dialogue Responses using Self-Separated Conditional Variational AutoEncoders
abstract
Bin Sun, Shaoxiong Feng, Yiwei Li, Jiamou Liu, Kan Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bin Sun 0004, Shaoxiong Feng, Yiwei Li 0001, Jiamou Liu, Kan Li 0001
ACL/IJCNLP (1)1
2020 Regularizing Dialogue Generation by Imitating Implicit Scenarios
abstract
Human dialogues are scenario-based and appropriate responses generally relate to the latent context knowledge entailed by the specific scenario.To enable responses that are more meaningful and context-specific, we propose to improve generative dialogue systems from the scenario perspective, where both dialogue history and future conversation are taken into account to implicitly reconstruct the scenario knowledge.More importantly, the conversation scenarios are further internalized using imitation learning framework, where the conventional dialogue model that has no access to future conversations is effectively regularized by transferring the scenario knowledge contained in hierarchical supervising signals from the scenario-based dialogue model, so that the future conversation is not required in actual inference.Extensive evaluations show that our approach significantly outperforms state-of-theart baselines on diversity and relevance, and expresses scenario-specific knowledge.
Shaoxiong Feng, Xuancheng Ren, Hongshen Chen, Bin Sun 0004, Kan Li 0001, Xu Sun 0001
EMNLP (1)4