Qiao Xiao

dblp:182/7575 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
abstract
Hong Ting Tsang, Jiaxin Bai, Haoyu Huang, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hong Ting Tsang, Jiaxin Bai, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu 0001, Yangqiu Song
ACL (1)4
2025 Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption Robustness
abstract
It is generally perceived that Dynamic Sparse Training opens the door to a new era of scalability and efficiency for artificial neural networks at, perhaps, some costs in accuracy performance for the classification task. At the same time, Dense Training is widely accepted as being the "de facto" approach to train artificial neural networks if one would like to maximize their robustness against image corruption. In this paper, we question this general practice. Consequently, \textit{we claim that}, contrary to what is commonly thought, the Dynamic Sparse Training methods can consistently outperform Dense Training in terms of robustness accuracy, particularly if the efficiency aspect is not considered as a main objective (i.e., sparsity levels between 10\% and up to 50\%), without adding (or even reducing) resource cost. We validate our claim on two types of data, images and videos, using several traditional and modern deep learning architectures for computer vision and three widely studied Dynamic Sparse Training algorithms. Our findings reveal a new yet-unknown benefit of Dynamic Sparse Training and open new possibilities in improving deep learning robustness beyond the current state of the art.
Boqian Wu, Qiao Xiao, Shunxin Wang, Nicola Strisciuglio, Mykola Pechenizkiy, Maurice van Keulen, Decebal Constantin Mocanu, Elena Mocanu
ICLR2
2025 MKGM: Multimodal knowledge-guided joint recognition of bridge defect-structural information
Jianxi Yang, Hao Li 0118, Qiao Xiao, Shixin Jiang
Adv. Eng. Informatics5
2025 Few-shot machine reading comprehension for bridge inspection via domain-specific and task-aware pre-tuning approach
Luyi Zhang, Qiao Xiao, Jianxi Yang, Shixin Jiang
Eng. Appl. Artif. Intell.3
2025 A Question-Aware Few-Shot Text-to-SQL Neural Model for Industrial Databases
abstract
Intelligent question answering over industrial databases is a challenging task due to the multicolumn context and complex questions. The existing methods need to be improved in terms of SQL generation accuracy. In this paper, we propose a question‐aware few‐shot Text‐to‐SQL approach based on the SDCUP pretrained model. Specifically, an attention‐based filtering approach is proposed to reduce the redundant information from multiple columns in the industrial database scenario. We further propose an operator semantics enhancement method to improve the ability of identifying complex conditions in queries. Experimental results on the industrial benchmarks in the fields of electric energy and structural inspection show that the proposed model outperforms the baseline models across all few‐shot settings.
Jianxi Yang, Qiao Xiao, Shixin Jiang
Int. J. Intell. Syst.5
2024 Are Sparse Neural Networks Better Hard Sample Learners?
Qiao Xiao, Boqian Wu, Lu Yin 0006, Christopher Neil Gadzinski, Tianjin Huang, Mykola Pechenizkiy, Decebal Constantin Mocanu
BMVC1
2024 Open-Set Semi-Supervised Learning by Distribution Alignment
abstract
Semi-Supervised Learning (SSL) has been shown to be effective in the closed-set case where the label spaces in labeled and unlabeled data are the same. However, in open-set SSL, its performance is seriously degraded since unlabeled data contains some classes not seen in the labeled data, leading to the distribution mismatch between labeled and unlabeled data. To solve this problem, we propose a Distribution Aligned Openset SSL (DAOSSL) method, which aims to explicitly reduce the empirical distribution mismatch between the labeled and unlabeled data. Specifically, we first introduce a progressive separation mechanism that utilizes a coarse-to-fine pipeline to weigh the unlabeled data. Based on this weighting strategy, we then propose a weighted distribution alignment approach to minimize the distribution discrepancy between the labeled and unlabeled data. These two strategies can be easily integrated into existing deep SSL approaches for open-set SSL tasks. The effectiveness of the proposed DAOSSL method is demonstrated through empirical studies, which show that the method is able to successfully reduce the distribution mismatch between labeled and unlabeled data, resulting in performance improvement in open-set SSL tasks.
Qiao Xiao, Jinjing Zhu, Boqian Wu
IJCNN1
2024 MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma 0001, Lu Yin 0006, Qiao Xiao, Stavros Petridis, Shiwei Liu 0003, Maja Pantic
INTERSPEECH5
2024 Dynamic Data Pruning for Automatic Speech Recognition
abstract
The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss.
Qiao Xiao, Pingchuan Ma 0001, Adriana Fernandez-Lopez, Boqian Wu, Lu Yin 0006, Stavros Petridis, Mykola Pechenizkiy, Maja Pantic, Decebal Constantin Mocanu, Shiwei Liu 0003
INTERSPEECH1
2024 E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation
abstract
Deep neural networks have evolved as the leading approach in 3D medical image segmentation due to their outstanding performance. However, the ever-increasing model size and computational cost of deep neural networks have become the primary barriers to deploying them on real-world, resource-limited hardware. To achieve both segmentation accuracy and efficiency, we propose a 3D medical image segmentation model called Efficient to Efficient Network (E2ENet), which incorporates two parametrically and computationally efficient designs. i. Dynamic sparse feature fusion (DSFF) mechanism: it adaptively learns to fuse informative multi-scale features while reducing redundancy. ii. Restricted depth-shift in 3D convolution: it leverages the 3D spatial information while keeping the model and computational complexity as 2D-based methods. We conduct extensive experiments on AMOS, Brain Tumor Segmentation and BTCV Challenge, demonstrating that E2ENet consistently achieves a superior trade-off between accuracy and efficiency than prior arts across various resource constraints. %In particular, with a single model and single scale, E2ENet achieves comparable accuracy on the large-scale challenge AMOS-CT, while saving over 69% parameter count and 27% FLOPs in the inference phase, compared with the previous best-performing method. Our code has been made available at: https://github.com/boqian333/E2ENet-Medical.
Boqian Wu, Qiao Xiao, Shiwei Liu 0003, Lu Yin 0006, Mykola Pechenizkiy, Decebal Constantin Mocanu, Maurice van Keulen, Elena Mocanu
NeurIPS2
2024 TPKE-QA: A gapless few-shot extractive question answering approach via task-aware post-training and knowledge enhancement
Qiao Xiao, Jianxi Yang, Yu Chen 0076, Shixin Jiang
Expert Syst. Appl.1
2024 Selective Random Walk for Transfer Learning in Heterogeneous Label Spaces
abstract
Transfer learning has been widely used in different scenarios, especially in those lacking enough labeled data. However, most of the existing transfer learning methods are based on the assumption that the source and target domains should share the label space entirely or partially, which greatly limits their application scopes. In this article, a Selective Random Walk (SRW) method for transfer learning in heterogeneous label spaces is proposed to make full use of unlabeled auxiliary data, which acts as a bridge for knowledge transfer from the source domain to the target domain. The proposed SRW method can explicitly identify transfer sequences between source and target instances via auxiliary instances based on random walk techniques. Since not all of the transfer sequences generated by random walk are credible for the target task, the SRW method can learn to weight transfer sequences adaptively. Based on the weights of the transfer sequences, the SRW method leverages knowledge by forcing adjacent data points in the transfer sequence to be similar and making the target data point in the sequence represented by other data points in the same sequence. Experiments show that the SRW method outperforms state-of-the-art models in plenty of transfer learning tasks with heterogeneous label spaces constructed within and across several benchmark datasets.
Qiao Xiao, Yu Zhang 0006, Qiang Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 MRC-PASCL: A Few-Shot Machine Reading Comprehension Approach via Post-Training and Answer Span-Oriented Contrastive Learning
abstract
The rapid development of pre-trained language models (PLMs) has significantly enhanced the performance of machine reading comprehension (MRC). Nevertheless, the traditional fine-tuning approaches necessitate extensive labeled data. MRC remains a challenging task in the few-shot settings or low-resource scenarios. This study proposes a novel few-shot MRC approach via post-training and answer span-oriented contrastive learning, termed MRC-PASCL. Specifically, in the post-training module, a novel noun-entity-aware data selection and generation strategy is proposed according to characteristics of MRC task and data, focusing on masking nouns and named entities in the context. In terms of fine-tuning, the proposed answer span-oriented contrastive learning manner selects spans around the golden answers as negative examples, and performs multi-task learning together with the standard MRC answer prediction task. Experimental results show that MRC-PASCL outperforms the PLMs-based baseline models and the 7B and 13B large language models (LLMs) cross most MRQA 2019 datasets. Further analyses show that our approach achieves better inference efficiency with lower computational resource requirement. The analysis results also indicate that the proposed method can better adapt to the domain-specific scenarios.
Qiao Xiao, Jianxi Yang, Luyi Zhang, Yu Chen 0076
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 A Versatile Framework for Unsupervised Domain Adaptation Based on Instance Weighting
abstract
Despite the progress made in domain adaptation, solving Unsupervised Domain Adaptation (UDA) problems with a general method under complex conditions caused by label shifts between domains remains a challenging task. In this work, we comprehensively investigate four distinct UDA settings including closed set domain adaptation, partial domain adaptation, open set domain adaptation, and universal domain adaptation, where shared common classes between source and target domains coexist alongside domain-specific private classes. The prominent challenges inherent in diverse UDA settings center around the discrimination of common/private classes and the precise measurement of domain discrepancy. To surmount these challenges effectively, we propose a novel yet effective method called Learning Instance Weighting for Unsupervised Domain Adaptation (LIWUDA), which caters to various UDA settings. Specifically, the proposed LIWUDA method constructs a weight network to assign weights to each instance based on its probability of belonging to common classes, and designs Weighted Optimal Transport (WOT) for domain alignment by leveraging instance weights. Additionally, the proposed LIWUDA method devises a Separate and Align (SA) loss to separate instances with low similarities and align instances with high similarities. To guide the learning of the weight network, Intra-domain Optimal Transport (IOT) is proposed to enforce the weights of instances in common classes to follow a uniform distribution. Through the integration of those three components, the proposed LIWUDA method demonstrates its capability to address all four UDA settings in a unified manner. Experimental evaluations conducted on four benchmark datasets substantiate the effectiveness of the proposed LIWUDA method. The code is available at https://github.com/JinjingZhu/LIWUDA.
Jinjing Zhu, Feiyang Ye 0001, Qiao Xiao, Pengxin Guo 0001, Yu Zhang 0006, Qiang Yang 0001
IEEE Trans. Image Process.3
2023 More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity
Shiwei Liu 0003, Tianlong Chen 0001, Xiaohan Chen 0001, Xuxi Chen, Qiao Xiao, Boqian Wu, Tommi Kärkkäinen, Mykola Pechenizkiy, Decebal Constantin Mocanu, Zhangyang Wang
ICLR5
2023 Few-Shot Relation Extraction via the Entity Feature Enhancement and Attention-Based Prototypical Network
abstract
In order to overcome the problems that the feature representation and classification effect of the existing methods need to be improved in complex contexts, this paper presents a novel few‐shot relation extraction approach via the entity feature enhancement and attention‐based prototypical network. The proposed model uses the pretrained RoBERTa model as the encoder while using the BiLSTM module for directional feature extraction. We further incorporate the entity feature enhancement module to improve the feature representation ability of the model. At last, the attention‐based prototypical network is used to predict relations. The experimental results show that the proposed method not only outperforms the baseline models on the datasets from the bridge inspection and health domains but also achieves competitive results on the FewRel dataset in the general domain.
Qiao Xiao, Jianxi Yang, Hao Ren 0009, Yu Chen 0076
Int. J. Intell. Syst.2
2022 Dynamic Sparse Network for Time Series Classification: Learning What to "See"
abstract
The receptive field (RF), which determines the region of time series to be “seen” and used, is critical to improve the performance for time series classification (TSC). However, the variation of signal scales across and within time series data, makes it challenging to decide on proper RF sizes for TSC. In this paper, we propose a dynamic sparse network (DSN) with sparse connections for TSC, which can learn to cover various RF without cumbersome hyper-parameters tuning. The kernels in each sparse layer are sparse and can be explored under the constraint regions by dynamic sparse training, which makes it possible to reduce the resource cost. The experimental results show that the proposed DSN model can achieve state-of-art performance on both univariate and multivariate TSC datasets with less than 50% computational cost compared with recent baseline methods, opening the path towards more accurate resource-aware methods for time series analyses. Our code is publicly available at: https://github.com/QiaoXiao7282/DSN.
Qiao Xiao, Boqian Wu, Yu Zhang 0006, Shiwei Liu 0003, Mykola Pechenizkiy, Elena Mocanu, Decebal Constantin Mocanu
NeurIPS1
2021 Distant Transfer Learning via Deep Random Walk
abstract
Transfer learning, which is to improve the learning performance in the target domain by leveraging useful knowledge from the source domain, often requires that those two domains are very close, which limits its application scope. Recently, distant transfer learning has been studied to transfer knowledge between two distant or even totally unrelated domains via unlabeled auxiliary domains that act as a bridge in the spirit of human transitive inference that two completely unrelated concepts can be connected through gradual knowledge transfer. In this paper, we study distant transfer learning by proposing a DeEp Random Walk basEd distaNt Transfer (DERWENT) method. Different from existing distant transfer learning models that implicitly identify the path of knowledge transfer between the source and target instances through auxiliary instances, the proposed DERWENT model can explicitly learn such paths via the deep random walk technique. Specifically, based on sequences identified by the random walk technique on a data graph where source and target data have no direct connection, the proposed DERWENT model enforces adjacent data points in a sequence to be similar, makes the ending data point be represented by other data points in the same sequence, and considers weighted classification losses of source data. Empirical studies on several benchmark datasets demonstrate that the proposed DERWENT algorithm yields the state-of-the-art performance.
Qiao Xiao, Yu Zhang 0006
AAAI1
2021 Multi-Objective Meta Learning
abstract
Meta learning with multiple objectives has been attracted much attention recently since many applications need to consider multiple factors when designing learning models. Existing gradient-based works on meta learning with multiple objectives mainly combine multiple objectives into a single objective in a weighted sum manner. This simple strategy usually works but it requires to tune the weights associated with all the objectives, which could be time consuming. Different from those works, in this paper, we propose a gradient-based Multi-Objective Meta Learning (MOML) framework without manually tuning weights. Specifically, MOML formulates the objective function of meta learning with multiple objectives as a Multi-Objective Bi-Level optimization Problem (MOBLP) where the upper-level subproblem is to solve several possibly conflicting objectives for the meta learner. To solve the MOBLP, we devise the first gradient-based optimization algorithm by alternatively solving the lower-level and upper-level subproblems via the gradient descent method and the gradient-based multi-objective optimization method, respectively. Theoretically, we prove the convergence properties of the proposed gradient-based optimization algorithm. Empirically, we show the effectiveness of the proposed MOML framework in several meta learning problems, including few-shot learning, domain adaptation, multi-task learning, and neural architecture search. The source code of MOML is available at https://github.com/Baijiong-Lin/MOML.
Feiyang Ye 0001, Baijiong Lin, Zhixiong Yue, Pengxin Guo 0001, Qiao Xiao, Yu Zhang 0006
NeurIPS5