Xingchun Diao

dblp:175/6585 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0007-2554-8899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AIM: Manifold-based Data Filtering for Representation Finetuning
abstract
Representation Finetuning (ReFT) has recently emerged as an efficient paradigm for adapting pretrained language models by editing hidden representations rather than model weights. However, our preliminary experiments reveal that ReFT is notably more sensitive to training data quality compared to traditional parameter-efficient finetuning methods, particularly to samples with incorrect labels, which can severely degrade performance. Inspired by prior work demonstrating that the hidden representations of generalizable neural networks exhibit low-dimensional manifold structures, we hypothesize that effective generalization in ReFT requires geometrically structured transformations between pre- and post-intervention representations. This implies that the intervention vectors representing these transformations should form a low-dimensional manifold, rendering the inconsistent transformations induced by label noise as detectable geometric outliers. To leverage this insight, we introduce Aligning Interventions on a learned Manifold (AIM), a representation-based data filtering method for ReFT, which identifies high-quality training samples by measuring the geometric consistency of their intervention vectors with respect to a robust reference manifold derived via principal component analysis on trusted data. Extensive experiments on both commonsense and arithmetic reasoning tasks confirm the effectiveness of AIM, showing consistent improvements over strong data selection baselines across multiple model scales.
Qibin Zheng, Xingchun Diao
AAAI4
2025 Don't Stop Pre-Training Small Language Models for Continual Enhancement of Reasoning
abstract
We investigate the continual enhancement of mathematical reasoning abilities in small language models (SLMs). While large language models (LLMs) demonstrate impressive reasoning performance, their deployment is often constrained by substantial computational costs. Existing approaches to improving SLMs mainly rely on knowledge distillation from costly teacher LLMs, which typically improves mathematical reasoning at the expense of general capabilities. In this work, we show that continual pre-training (CPT) has strong potential to enhance the mathematical reasoning ability of SLMs without relying on large teacher models. We also find that its effectiveness critically depends on the quality of the training data. To maximize efficiency and performance, we propose Dual-Metric Selection for Continual Pre-training (DRIFT), a novel data selection strategy that identifies optimal training data through task-aligned loss differences and distributional regularization. To further enhance task-specific reasoning while preserving general capabilities, we introduce a metadata-aware data mixture that integrates diverse sources during CPT. Extensive experiments on multiple arithmetic reasoning benchmarks demonstrate the effectiveness of DRIFT: SLMs trained with DRIFT achieve substantial gains in reasoning performance, surpassing larger models on specific tasks, while largely preserving general capabilities.
Qibin Zheng, Yi Liu 0043, Xingchun Diao
ECAI4
2024 G-ETI: Incorporating Graph Information for Improved Unsupervised Event Type Induction
Lingzhi Xiang, Xingchun Diao
WISE (2)4
2022 A Trusted Storage System for Digital Object in the Human-Cyber-Physical Environment
Xiang Jing, Yueyang Hu, Chaoran Luo, Xingchun Diao, Gang Huang 0001, Haiou Jiang
BlockSys4
2022 DataAttest: A Framework to Attest Off-Chain Data Authenticity
Ying Zhang 0012, Xiang Jing, Xingchun Diao, Gang Huang 0001
BlockSys4
2022 Fission: Autonomous, Scalable Sharding for IoT Blockchain
abstract
IoT blockchain suffers heavy performance issues because of the massive transactions generated by various IoT nodes. By dividing nodes into different shards, sharding can produce blocks in parallel and hence improve the throughput of the blockchain system. Unlike the traditional blockchain system, IoT blockchain mainly consists of smart devices and the transactions are usually generated from the real world, such as the sensor data, photos taken by cameras, and so on. In IoT blockchain, closer nodes usually share a lower network latency and the transactions they generate are more related. Therefore, location-based sharding is an effective approach to improve the performance of IoT blockchain. Traditionally, IoT nodes are di-vided into different shards based on geographical locations or the connected edge server. However, the key challenge of sharding in IoT blockchain is how to guarantee the equality of shards division as to the unpredictable distribution and the dynamic behavior of IoT nodes. On one hand, shards can not be pre-divided because we can not predict the number or the distribution of the IoT nodes. On the other hand, nodes continuously joining or quitting shards will also break the equality of the shards division. In this paper, we propose Fission, a sharding mechanism designed for IoT blockchain. Fission divide shards based on the Voronoi diagram without any preknowledge about the nodes distribution, and support dynamic, autonomous sharding adjustment based on distributed Delaunay Triangulation. In addition, Fission uses a new diffusion-based consensus algorithm to achieve the linear scalability of throughput. The experimental results show that Fission can construct and adjust shards at a very low cost and can execute in a decentralized manner. The throughput can reach 1900tps in 500 nodes with only 5M bps bandwidth, and can scale linearly as the nodes increase.
Chaoran Luo, Yueyang Hu, Ying Zhang 0012, Yi Liu 0014, Xingchun Diao, Gang Huang 0001
COMPSAC6
2020 From Whole to Part: Reference-Based Representation for Clustering Categorical Data
abstract
Dissimilarity measures play a crucial role in clustering and, are directly related to the performance of clustering algorithms. However, effectively measuring the dissimilarity is not easy, especially for categorical data. The main difficulty of the dissimilarity measurement for categorical data is that its representation lacks a clear space structure. Therefore, the space structure-based representation has been proposed to provide the categorical data with a clear linear representation space. This representation improves the clustering performance obviously but only applies to small data sets because its dimensionality increases rapidly with the size of the data set. In this paper, we investigate the possibility of reducing the dimensionality of the space structure-based representation while maintaining the same representation ability. A lightweight representation scheme is proposed by taking a set of representative objects as the reference system (called the reference set) to position other objects in the Euclidean space. Moreover, a preclustering-based strategy is designed to select an appropriate reference set quickly. Finally, the representation scheme together with the k -means algorithm provides an efficient method to cluster the categorical data. The theoretical and the experimental analysis shows that the proposed method outperforms state-of-the-art methods in terms of both accuracy and efficiency.
Qibin Zheng, Xingchun Diao, Jianjun Cao, Yi Liu 0043, Junnan Yao, Chen Chang, Guojun Lv
IEEE Trans. Neural Networks Learn. Syst.2
2019 Tag-aware recommendation based on Bayesian personalized ranking and feature mapping
abstract
Collaborative filtering recommendation with implicit feedbacks (i.e., clicks, views, check-ins) has been gaining increasing attention in various real applications. Tagging information is the common resource to complement implicit feedbacks to assist collaborative filtering recommendation. However, existing tag-aware recommendation methods still suffer from the problem of high dimension and sparsity of tagging information. They also fail to realize that recommendation is inherent a ranking-oriented optimization task. To this end, we propose a novel tag-aware recommendation framework by incorporating tag mapping scheme into ranking-based collaborative filtering model, to boost ranking-oriented personalized recommendation performance. We first build ranking-oriented optimization model based on Bayesian personalized ranking optimization criterion with matrix factorization, by leveraging implicit feedbacks to learn the latent feature vectors of users and items. Then, we propose an explicit-to-implicit feature mapping scheme, mapping the high-dimensional and sparse explicit tags (i.e., user-tag weighting matrix and item-tag weighting matrix) to low-dimensional and compact implicit features of uses and items. This could serve as the regularization constraint of latent features derived from implicit feedbacks. To further enhance recommendation performance, we also introduce users’ neighbor relationships to regularize user latent features based on manifold learning. Experiments on real-world recommendation datasets show that the proposed recommendation method outperformed competing methods on ranking-oriented recommendation performance.
Xingchun Diao, Jianjun Cao, Lei Zhang 0126, Qin Feng
Intell. Data Anal.2