Huashan Sun

dblp:49/3852 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Simulated Rewards, Skewed Strategies: Tracing the Acquired Preference Bias in LLM-Based Dialogue Planners
abstract
Large language models have enabled sophisticated dialogue planning policy, but their reliance on LLM-generated simulation and feedback for policy optimization may introduce systematic preference bias. We present the first comprehensive analysis of preference bias in LLM-based dialogue planners, evaluating four state-of-the-art planning policies across three dialogue domains using multiple LLM families at varying scales. Our investigation reveals that all tested planners exhibit significant preference bias, systematically favoring narrow strategy sets rather than maintaining balanced distributions. User simulation emerges as the primary bias driver, while diverse persona simulation fails as an effective mitigation strategy. Most concerning, preference bias drives planners toward ethically problematic strategies that achieve short-term success while undermining real-world effectiveness and ethical standards. Our findings establish fundamental challenges for responsible deployment of LLM-based dialogue systems and provide crucial insights for developing more reliable and ethically-aligned planning approaches.
Heyan Huang, Yizhe Yang, Huashan Sun, Jiawei Li 0020, Yang Gao 0016
AAAI3
2026 Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning Models
abstract
While Large Reasoning Models (LRMs) have demonstrated remarkable capabilities through explicit Chain-of-Thought (CoT) generation, they frequently suffer from "overthinking".In this work, we bridge this gap by introducing Token-level Marginal Utility, which quantifies the per-token log-probability gain of the ground-truth answer.Leveraging this dense supervision signal, we propose MUTO (Marginal Utility Guided Thinking Optimization), a unified training framework designed to synthesize concise reasoning chains.Rather than relying only on coarse trajectory-level length control, MUTO identifies tokens that reduce the model's likelihood of the correct answer and penalizes such negative-utility reasoning, yielding concise yet effective CoT trajectories.Experiments on DeepSeek-R1-Distill-Qwen backbones (1.5B and 7B) across six math reasoning benchmarks show that MUTO yields a markedly better efficiency-accuracy Pareto frontier.It reduces average token usage by 87.1% at 1.5B while improving accuracy by 2.3%, and cuts tokens by 80.2% at 7B with only -0.1% accuracy change, achieving the best length-normalized accuracy among baselines.
Jiawei Li 0020, Yang Gao 0016, Huashan Sun, Chong Feng 0001
ACL (1)3
2026 EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
abstract
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yu Bai 0018, Huashan Sun, Yiguan Lin, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Yang Gao 0016, Heyan Huang
ACL (1)3
2024 Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A Survey
abstract
Jiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou, Yinghao Li, Huashan Sun, Yuhang Liu, Xingpeng Si, Yuhao Ye, Yixiao Wu, Yiguan Lin, Bin Xu, Bowen Ren, Chong Feng, Yang Gao, Heyan Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jiawei Li 0020, Yizhe Yang, Yu Bai 0018, Xiaofeng Zhou 0004, Huashan Sun, Xingpeng Si, Yuhao Ye, Yixiao Wu, Yiguan Lin, Ren Bowen, Chong Feng 0001, Yang Gao 0016, Heyan Huang
ACL (1)6
2024 Temporal context video compression with flow-guided feature prediction
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun, Zhuang Miao
Expert Syst. Appl.4
2024 A survey of feature matching methods
abstract
Abstract Feature matching plays a crucial role in computer vision, with applications in visual localization, simultaneous localization and mapping (SLAM), image stitching, and more. It establishes correspondences between sets of feature points from multiple images, enabling various tasks. Over the years, feature matching has witnessed significant development, with an increasing number of methods being applied. However, different methods exhibit different degrees of applicability in different scenarios and requirements due to their different rationales. To cope with these issues, a comprehensive analysis and comparison of matching methods are essential. Existing reviews often lack coverage of deep learning models and focus more on feature detection and description, neglecting the matching process. This survey investigates feature detection, description, and matching techniques within the feature‐based image‐matching pipeline. Representative methods, their mechanisms, and application scenarios are also briefly introduced. In addition, comprehensive evaluations of classical and state‐of‐the‐art methods are conducted through extensive experiments on representative datasets. Particularly, matching‐based applications are compared to fully demonstrate the advantages of the methods. Lastly, this survey highlights current problems and development directions in matching methods, serving as a reference for researchers in the field.
Qian Huang 0008, Yiming Wang 0008, Huashan Sun
IET Image Process.4
2024 Ship detection based on YOLO algorithm for visible images
abstract
Abstract Ship detection is a crucial task for waterway surveillance and channel optimization, especially in close proximity to the shore. However, detecting ship in visible image‐based detection remains a challenge due to the limited nature of visible image datasets. To address this issue, the Inland Ships Data Set (ISDS) is constructed to facilitate research on ship identification. On the other hand, most detection methods struggle to accurately identify ships that are small in size. Therefore, a visible image‐based ship detection model is proposed that employs a multi‐scale weighted feature fusion structure with the YOLOv4 detection model to improve the efficacy of small ship detection. Specifically, the YOLOv4 model is improved through fusing multi‐scale feature, redesigning priori frame, and enhancing loss function. The model, named YOLOv4‐MSW (i.e. YOLOv4 based on Multi‐Scale Weighted feature fusion), exhibits improved performance on ship detection in experiments conducted on the ISDS dataset, outperforming the original YOLOv4 model by improving the average precision (AP) by 4.87% and the recall rate by 10.03%. Meanwhile, the model achieve better detection accuracy and improve the average precision rate by at least 0.86% compared to existing learned object detection methods. The code related to this work are released at https://github.com/Sunhuashan/YOLOv4‐MSW . The whole dataset is available at https://drive.google.com/drive/folders/1fzJ2fcqiko6lFwqIEGghMceoQgv‐8jBy .
Qian Huang 0008, Huashan Sun, Yiming Wang 0008
IET Image Process.2
2023 FGC-VC: Flow-Guided Context Video Compression
abstract
Deep video compression has attracted more and more attention in recent years. Previous works rely on feature space operations, which may cause the offset maps overflow degrading reconstructed frame quality. In this work, we propose a flow- guided module to guide the offset maps learning explicitly and alleviate offset maps overflow. Moreover, we introduce a context scheme to explore the temporal prior and fuse the hyper prior model to improve the compression ratio. For coding speed, we drop the time-consuming auto regressive module. Experimental results demonstrate that our method out-performs the previous learning-based schemes and traditional codecs. Compared to x265 with medium preset, our approach brings average 38.53% and 54.67% bit rate savings in PSNR and MS-SSIM metrics, respectively.
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun
ICIP4
2023 End-to-End Variable-Rate Image Compression with Bi-Resolution Spatial-Channel Context Aggregation
abstract
Recently, neural network-based image compression techniques have demonstrated remarkable compression performance. The use of context-adaptive entropy models greatly enhances the rate-distortion (R-D) performance by effectively capturing spatial redundancy in latent representations. However, latent representations still contain some spatial correlations(e.g. same spatial structure), it needs to be eliminated by further processing. And many compression models are single-rate model, which is difficult to cover a big range of bitrate. In order to address this issue, we propose a novel variable-rate image compression algorithm that efficiently leverages bi-resolution spatial-channel information through learned mechanisms. In this paper, we first proposed a BRP network to divide our latent representations and side information into HR and LR components, eliminating the spatial redundancy in same location. Combining the spatial-channel context, we proposed a BSC context model, including a decreasing-granularity checkerboard pattern and channel grouping based on cosine slicing strategy. To cover a wide range of bitrate, we take a weight map as input to control bit allocation, achieving multiple compression rates. Our experimental results show that our method provides a better rate-distortion trade-off than BPG, JPEG and other recent image compression methods based on deep learning.
Qian Huang 0008, Yiming Wang 0008, Huashan Sun
MMAsia4
2023 Optical Flow based Feature Prediction and Decomposed Context for Video Compression
abstract
In recent years, there have been a growing interest in developing end-to-end neural video codecs. Previous works generally use a past decoded frame as reference directly, utilizing the motion information between it and the input frame to reduce temporal redundancy. However, this approach may lead to high bit rate consumption of the motion and fails to take advantage of the prior information in other reconstructed frames. In this work, We propose a learned video coding framework with optical flow based feature prediction module and decomposed context module. Specifically, we employ the previous optical flow to generate a warped frame, and along with other reconstructions, they are used for a more accurate reference forecasting, thereby reducing the bit rate required for motion compression. Moreover, based on the conditional coding framework, our decomposed context module explores conditional context in past decoded frames and further reduces additional spatiotemporal correlations. Experimental results demonstrate that our approach yields better performance than previous learned video compression methods and traditional standard codecs. For example, our neural codec achieves 28.94% coding gain over HEVC in PSNR metric and about 2.00% coding gain over VVC in MS-SSIM metric.
Huashan Sun, Qian Huang 0008, Yiming Wang 0008, Ruoyu Hao
MMAsia1
2007 Evolutionary Neural Networks Applied to Land-cover Classification in Zhaoyuan, China
abstract
This paper proposes a method for the classification of land cover in remote sensing imagery using evolutionary artificial neural networks (EANN) compared against multilayer perceptrons (MLP) with backpropagation algorithm. Evolutionary neural networks have combined the features of artificial neural networks (ANN) and evolutionary algorithms (EA) in the way that simultaneously evolving ANN architecture and weights. The parsimony of evolved ANN is encouraged by preferring node mutation and connection mutation. This enables consistent reductions of mean square errors of spectral classification with respect to sample pixels. Land-cover classification experiments were carried out by EANN-based classifiers and MLP-based classifiers in a 300times300 pixels Landsat-7 Enhanced Thematic Mapper plus (ETM+) high-resolution image of Zhaoyuan in Shandong province in eastern China. We found that the use of evolutionary algorithms for finding the optimal ANN results mainly in improvements in overall accuracy of an ANN with backpropagation algorithm and produce more compact ANN with good generalization ability in comparison with MLP. It is observed that classification accuracy of up to 90% is achievable for Landsat data produced by EANN.
Lishan Kang, Fujiang Liu, Huashan Sun, Linlu Mei
CIDM4