Ting Xiao 0002

dblp:169/4621-2 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0003-3155-7664ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MSCFNet: mixed-scale context fusion network for medical image segmentation
Dongfang Tang, Ting Xiao 0002, Hao Wang 0234, Zhe Wang 0002, Wen Gao 0001
Appl. Intell.3
2026 Improving multi-instance learning with hierarchical attention and frequency-domain hard sample distillation
Ting Xiao 0002, Minqian Sun, Yiqing Xia, Hai Yang 0002, Zhe Wang 0002, Peng Liu 0008
Knowl. Based Syst.1
2026 Bamboo: A Novel Session-Aware Framework With Equiangular Tight Frame Prototypes for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) presents a greater challenge compared with few-shot task-incremental learning (FSTIL) due to the need to classify all previous classes without prior knowledge of the session identifier (session-ID). To address this, we propose Bamboo, a novel framework for FSCIL that introduces a cascading inference mechanism to explicitly infer the session-ID for each sample. This mechanism is enabled by a novel, session-specific equiangular tight frame prototype (ETF-P) classifier. By adaptively fusing session-agnostic and session-specific semantics, the ETF-P classifier reliably determines if a sample belongs to its associated session, which is the core decision required at each step of the cascade. Considering the incremental nature of the learning process, which resembles the continuous growth of bamboo, we treat the base session classifier as the foundational bamboo node and progressively add new session classifiers as additional nodes on top. During the testing phase, each sample flows sequentially through the bamboo nodes, from top to bottom, to determine its session-ID and to be classified accordingly. Overall, the Bamboo framework is capable of perceiving session-ID without prior knowledge and classifying each sample within the correct session, leading to state-of-the-art performance on multiple benchmark datasets.
Xuehan Lu, Zhe Wang 0002, Zhiling Fu, Xinlei Xu, Qian Zhang 0017, Ting Xiao 0002, Wenli Du
IEEE Trans. Neural Networks Learn. Syst.6
2026 Unsupervised Skill Discovery Through Skill Regions Differentiation
abstract
Unsupervised reinforcement learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based exploration or empowerment-driven skill learning. However, entropy-based exploration struggles in large-scale state spaces (e.g., images), and empowerment-based methods with mutual information (MI) estimations have limitations in state exploration. To address these challenges, we propose a novel skill discovery objective that maximizes the deviation of the state density of one skill from the explored regions of other skills, encouraging inter-skill state diversity similar to the initial MI objective. For state-density estimation, we construct a novel conditional autoencoder with soft modularization for different skill policies in high-dimensional space. Meanwhile, to incentivize intra-skill exploration, we formulate an intrinsic reward based on the learned autoencoder that resembles count-based exploration in a compact latent space. Through extensive experiments in challenging state and image-based tasks, we find our method learns meaningful skills and achieves superior performance in various downstream tasks.
Ting Xiao 0002, Jiakun Zheng, Rushuai Yang, Qiaosheng Zhang 0002, Peng Liu 0008, Zhe Wang 0002, Chenjia Bai
IEEE Trans. Neural Networks Learn. Syst.1
2025 Radiology Report Generation via Multi-objective Preference Optimization
abstract
Automatic Radiology Report Generation (RRG) is an important topic for alleviating the substantial workload of radiologists. Existing RRG approaches rely on supervised regression based on different architectures or additional knowledge injection, while the generated report may not align optimally with radiologists’ preferences. Especially, since the preferences of radiologists are inherently heterogeneous and multi-dimensional, e.g., some may prioritize report fluency, while others emphasize clinical accuracy. To address this problem, we propose a new RRG method via Multi-objective Preference Optimization (MPO) to align the pre-trained RRG model with multiple human preferences, which can be formulated by multi-dimensional reward functions and optimized by multi-objective reinforcement learning (RL). Specifically, we use a preference vector to represent the weight of preferences and use it as a condition for the RRG model. Then, a linearly weighed reward is obtained via a dot product between the preference vector and multi-dimensional reward. Next, the RRG model is optimized to align with the preference vector by optimizing such a reward via RL. In the training stage, we randomly sample diverse preference vectors from the preference space and align the model by optimizing the weighted multi-objective rewards, which leads to an optimal policy on the entire preference space. When inference, our model can generate reports aligned with specific preferences without further fine-tuning. Extensive experiments on two public datasets show the proposed method can generate reports that cater to different preferences in a single model and achieve state-of-the-art performance.
Ting Xiao 0002, Lei Shi 0004, Peng Liu 0008, Zhe Wang 0002, Chenjia Bai
AAAI1
2025 Online Iterative Self-Alignment for Radiology Report Generation
abstract
Radiology Report Generation (RRG) is an important research topic for relieving radiologists' heavy workload.Existing RRG models mainly rely on supervised fine-tuning (SFT) based on different model architectures using data pairs of radiological images and corresponding radiologist-annotated reports.Recent research has shifted focus to post-training improvements, aligning RRG model outputs with human preferences using reinforcement learning (RL).However, the limited data coverage of high-quality annotated data poses risks of overfitting and generalization.This paper proposes a novel Online Iterative Self-Alignment (OISA) method for RRG that consists of four stages: self-generation of diverse data, selfevaluation for multi-objective preference data, self-alignment for multi-objective optimization and self-iteration for further improvement.Our approach allows for generating varied reports tailored to specific clinical objectives, enhancing the overall performance of the RRG model iteratively.Unlike existing methods, our framework significantly increases data quality and optimizes performance through iterative multiobjective optimization.Experimental results demonstrate that our method surpasses previous approaches, achieving state-of-the-art performance across multiple evaluation metrics.
Ting Xiao 0002, Lei Shi 0004, Yang Zhang 0072, HaoFeng Yang, Zhe Wang 0002, Chenjia Bai
ACL (1)1
2025 Ipvar: Advancing Pathogenicity Prediction Via Hierarchical Fusion of Structure Foundation Models Alphafold 3 and Esm C
abstract
Interpreting the functional consequences of coding variants remains a central challenge in human genetics, particularly given the clinical importance of distinguishing pathogenic mutations from benign variation. Here we present IPVAR, a deep learning framework that uniquely integrates tertiary protein structural features predicted by both ESM C and AlphaFold 3, alongside established conservation metrics, to advance variant pathogenicity prediction. IPVAR leverages a hierarchical crossattention mechanism to capture both global and fine-grained structural dependencies between complementary representations, and incorporates an adaptive modality weighting strategy to dynamically balance information from each protein structure model. Comprehensive benchmarking demonstrates that IPVAR substantially outperforms state-of-the-art methods, including those based solely on sequence annotations or individual structural predictors, achieving an area under the ROC curve (AUC) of 0.9838 on the ClinVar dataset and 0.9202 on an independent Mendelian disease variant cohort. Ablation studies further confirm that both the multi-model integration and advanced fusion modules are critical to the model's superior performance. These findings establish IPVAR as a new benchmark for the functional interpretation of genomic variants, and highlight the value of integrating diverse structural foundation models to improve clinical variant assessment.
Hanwen Huang, Ziquan Bao, Yingzhuo Wang, Qin Zhou 0002, Ting Xiao 0002, Qian Zhang 0068, Dongdong Li 0003, Hai Yang 0002
BIBM6
2025 MMMNet: Multimodal Feature Fusion and Multilevel Representation Merging for Pulmonary Nodule Classification
abstract
In early lung cancer screening, precise classification of benign and malignant pulmonary nodules is of critical importance to clinical decision making and individualized treatment. Although early methods have achieved considerable progress, two main problems remain: existing diagnostic methods have limitations in utilizing multimodal data and capturing semantic information, and traditional multimodal approaches relying on late-stage feature fusion fail to facilitate the valuable information from internal model layers. To overcome these challenges, we propose Multimodal and Multilevel Merging Net (MMMNet), a multimodal architecture for pulmonary nodule classification that includes two innovative modules: a Multimodal Feature Fusion Module that combines computed tomography scans and text annotations parallelly in multiple layers to construct a feature pyramid, and a Multilevel Feature Merge Module that recursively merges the fused features to utilize both high-level semantic information and low-level visual characteristics. The approach also integrates Focal Loss to tackle the imbalanced classification and Contrastive Loss to align the multimodal features. The proposed approach is evaluated on the LIDC-IDRI dataset, yielding an accuracy of$\mathbf{9 1. 7 3 \%}$and a specificity of$\mathbf{9 5. 5 2 \%}$. Experiment results show that the approach enhances the indepth multimodal feature mining and has a promising potential in medical image analysis and clinical application.
Haihua Huang, Dongfang Tang, Ting Xiao 0002, Hai Yang 0002, Zhe Wang 0002, Wen Gao 0001
BIBM4
2025 Few-Shot Adaptive Diffusion with Semantic Injection and Parameter Smoothing
Yunjie Cai, Ting Xiao 0002, Yanbing Zhang, Zhe Wang 0002
ICMR2
2025 Dual-Prototype Learning in Multiple Instance Learning for Histopathology Image Classification
abstract
Existing prototype learning-based Multiple Instance Learning (MIL) methods mainly focus on learning a single set of prototypes for each class or generating a generic prototype from the overall data distribution. This design forces the model to compress general and heterogeneous features into identical prototype embeddings, prioritizing general features over subtle but discriminative features. Additionally, these methods often guide prototype updates by jointly optimizing attention score distributions and the distances between instances and prototypes, resulting in prototype biases due to over-concentration. To address these issues, we propose a dual prototype learning MIL (DP-MIL) framework that introduces two distinct sets of prototypes: primary prototypes, which capture general WSI features, and boundary prototypes, which capture discriminative features near the decision boundary. The DP-MIL framework employs three prototype-tailored losses: an alienation loss to encourage primary prototypes to be distant from decision boundaries, an affinity loss to anchor boundary prototypes near these boundaries, and a distance loss to enforce separation between the two prototype sets. To mitigate prototype semantic drift during training, we introduce a prototype joint updating and refinement strategy: for each prototype, we use its corresponding global token to filter out the most similar instances to momentum update the corresponding prototype set, while the boundary prototype set is refined with the mean pooled feature of hard samples. Extensive experiments on four datasets demonstrate the effectiveness of our DP-MIL framework and prototype updating strategy.
Ting Xiao 0002, Minqian Sun, Yiqing Xia, Zhe Wang 0002
ACM Multimedia1
2025 HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning
abstract
For robotic manipulation, existing robotics datasets and simulation benchmarks predominantly cater to robot-arm platforms. However, for humanoid robots equipped with dual arms and dexterous hands, simulation tasks and high-quality demonstrations are notably lacking. Bimanual dexterous manipulation is inherently more complex, as it requires coordinated arm movements and hand operations, making autonomous data collection challenging. This paper presents HumanoidGen, an automated task creation and demonstration collection framework that leverages atomic dexterous operations and LLM reasoning to generate relational constraints. Specifically, we provide spatial annotations for both assets and dexterous hands based on the atomic operations, and perform an LLM planner to generate a chain of actionable spatial constraints for arm movements based on object affordances and scenes. To further improve planning ability, we employ a variant of Monte Carlo tree search to enhance LLM reasoning for long-horizon tasks and insufficient annotation. In experiments, we create a novel benchmark with augmented scenarios to evaluate the quality of the collected data. The results show that the performance of the 2D and 3D diffusion policies can scale with the generated dataset. Project page is https://openhumanoidgen.github.io.
Zhi Jing 0004, Jicong Ao, Ting Xiao 0002, Yu-Gang Jiang 0001, Chenjia Bai
NeurIPS4
2025 VSNet: classification of pulmonary nodules in 3D using vision transformer and sequence spatial attention mechanism
Dongfang Tang, Ting Xiao 0002, Conghao Zhang, Zhe Wang 0002, Wen Gao 0001
Multim. Tools Appl.2
2025 Low-Rank Representation with Empirical Kernel Space Embedding of Manifolds
Wenyi Feng, Zhe Wang 0002, Ting Xiao 0002
Neural Networks3
2025 Semantic Mask Reconstruction and Category Semantic Learning for few-shot image generation
Ting Xiao 0002, Yunjie Cai, Jiaoyan Guan, Zhe Wang 0002
Neural Networks1
2025 Fusion of global and adaptive local information for few-shot image classification
Ting Xiao 0002, Yiqing Xia, Ruiqi Tang, Wenli Du, Zhe Wang 0002
Pattern Recognit.1
2024 Constrained Ensemble Exploration for Unsupervised Skill Discovery
abstract
Unsupervised Reinforcement Learning (RL) provides a promising paradigm for learning useful behaviors via reward-free per-training. Existing methods for unsupervised RL mainly conduct empowerment-driven skill discovery or entropy-based exploration. However, empowerment often leads to static skills, and pure exploration only maximizes the state coverage rather than learning useful behaviors. In this paper, we propose a novel unsupervised RL framework via an ensemble of skills, where each skill performs partition exploration based on the state prototypes. Thus, each skill can explore the clustered area locally, and the ensemble skills maximize the overall state coverage. We adopt state-distribution constraints for the skill occupancy and the desired cluster for learning distinguishable skills. Theoretical analysis is provided for the state entropy and the resulting skill distributions. Based on extensive experiments on several challenging tasks, we find our method learns well-explored ensemble skills and achieves superior performance in various downstream tasks compared to previous methods.
Chenjia Bai, Rushuai Yang, Qiaosheng Zhang 0002, Ting Xiao 0002, Xuelong Li 0001
ICML6
2024 Relationship constraint deep metric learning
Yanbing Zhang, Ting Xiao 0002, Zhe Wang 0002, Wenyi Feng, Zhiling Fu, Hai Yang 0002
Appl. Intell.2
2024 Freezing partial source representations matters for image inpainting under limited data
Yanbing Zhang, Mengping Yang, Ting Xiao 0002, Zhe Wang 0002, Ziqiu Chi
Eng. Appl. Artif. Intell.3
2024 Adaptive weighted dictionary representation using anchor graph for subspace clustering
Wenyi Feng, Zhe Wang 0002, Ting Xiao 0002, Mengping Yang
Pattern Recognit.3
2024 Monotonic Quantile Network for Worst-Case Offline Reinforcement Learning
abstract
A key challenge in offline reinforcement learning (RL) is how to ensure the learned offline policy is safe, especially in safety-critical domains. In this article, we focus on learning a distributional value function in offline RL and optimizing a worst-case criterion of returns. However, optimizing a distributional value function in offline RL can be hard, since the crossing quantile issue is serious, and the distribution shift problem needs to be addressed. To this end, we propose monotonic quantile network (MQN) with conservative quantile regression (CQR) for risk-averse policy learning. First, we propose an MQN to learn the distribution over returns with non-crossing guarantees of the quantiles. Then, we perform CQR by penalizing the quantile estimation for out-of-distribution (OOD) actions to address the distribution shift in offline RL. Finally, we learn a worst-case policy by optimizing the conditional value-at-risk (CVaR) of the distributional value function. Furthermore, we provide theoretical analysis of the fixed-point convergence in our method. We conduct experiments in both risk-neutral and risk-sensitive offline settings, and the results show that our method obtains safe and conservative behaviors in robotic locomotion tasks.
Chenjia Bai, Ting Xiao 0002, Zhoufan Zhu, Lingxiao Wang 0003, Animesh Garg, Bin He 0003, Peng Liu 0008, Zhaoran Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Robust Structured Sparse Subspace Clustering with Neighborhood Preserving Projection
abstract
Sparse subspace clustering algorithm cluster the data points located on the union of low-dimensional subspaces through the ℒ1minimization program. However, the ℒ1-norm is not rotation invariant, and utilizing original data containing noise as the dictionary leads to poor performance. This paper proposes a method named robust structured sparse subspace clustering with neighborhood preserving projection (RSSSC). Firstly, RSSSC replaces the ℒ1minimization program with a structured re-weighting sparse regularization term, effectively recovering the sparse representation. Secondly, RSSSC uses the extracted features as the dictionary in the self-representation problem to replace the original dataset containing noise and outliers, thus making the model more robust. By fully considering the low-dimensional manifold structure of samples in the original high-dimensional space, RSSSC preserves the neighborhood structure while learning the optimal projection. We verify the effectiveness of the proposed method through experiments on real-world image datasets.
Wenyi Feng, Wei Guo 0023, Ting Xiao 0002, Zhe Wang 0002
ICME3
2023 Semantic-Aware Generator and Low-level Feature Augmentation for Few-shot Image Generation
abstract
Few-shot image generation aims to generate novel images for an unseen category with only a few samples. Prior studies fail to produce novel images with desirable diversity and fidelity. To ameliorate the generation quality, we in this paper propose a Semantic-Aware Generator (SAG) to provide explicit semantic guidance to the discriminator, and a Low-level Feature Augmentation (LFA) technique to provide fine-grained information, facilitating the diversity. Specifically, we observe that the generator feature layers contain different levels of semantic information. Such observation motivates us to employ intermediate feature maps of the generator as semantic labels to guide the discriminator, improving the semantic awareness of the generator. Moreover, spatially informative and diverse features obtained via LFA contribute to better generation quality. Together with the aforementioned module, we conduct extensive experiments on three representative benchmarks and the results demonstrate the effectiveness and advancement of our method.
Zhe Wang 0002, Jiaoyan Guan, Mengping Yang, Ting Xiao 0002, Ziqiu Chi
ACM Multimedia4
2023 Improving Few-shot Image Generation by Structural Discrimination and Textural Modulation
abstract
Few-shot image generation, which aims to produce plausible and diverse images for one category given a few images from this category, has drawn extensive attention. Existing approaches either globally interpolate different images or fuse local representations with pre-defined coefficients. However, such an intuitive combination of images/features only exploits the most relevant information for generation, leading to poor diversity and coarse-grained semantic fusion. To remedy this, this paper proposes a novel textural modulation (TexMod) mechanism to inject external semantic signals into internal local representations. Parameterized by the feedback from the discriminator, our TexMod enables more fined-grained semantic injection while maintaining the synthesis fidelity. Moreover, a global structural discriminator (StructD) is developed to explicitly guide the model to generate images with reasonable layout and outline. Furthermore, the frequency awareness of the model is reinforced by encouraging the model to distinguish frequency signals. Together with these techniques, we build a novel and effective model for few-shot image generation. The effectiveness of our model is identified by extensive experiments on three popular datasets and various settings. Besides achieving state-of-the-art synthesis performance on these datasets, our proposed techniques could be seamlessly integrated into existing models for a further performance boost. Our code and models are available at \hrefhttps://github.com/kobeshegu/SDTM-GAN-ACMMM-2023 here.
Mengping Yang, Zhe Wang 0002, Wenyi Feng, Qian Zhang 0068, Ting Xiao 0002
ACM Multimedia5
2020 Importance-weighted conditional adversarial network for unsupervised domain adaptation
Peng Liu 0008, Ting Xiao 0002, Cangning Fan, Wei Zhao 0008, Xianglong Tang, Hongwei Liu 0002
Expert Syst. Appl.2
2020 Domain adaptation based on domain-invariant and class-distinguishable feature learning using multiple adversarial networks
Cangning Fan, Peng Liu 0008, Ting Xiao 0002, Wei Zhao 0008, Xianglong Tang
Neurocomputing3
2019 Structure preservation and distribution alignment in discriminative transfer subspace learning
Ting Xiao 0002, Peng Liu 0008, Wei Zhao 0008, Hongwei Liu 0002, Xianglong Tang
Neurocomputing1