Xiyao Liu 0002

dblp:138/9719-2 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-9978-1286ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ObjecTok: Learning Holistic and Robust Object Tokens for MLLMs
abstract
Mainstream multimodal large language models (MLLMs) rely on patch-based tokenization methods, which compromise the integrity of objects and thereby limit the model's perception capabilities while triggering object-related hallucinations. To address this issue, we propose ObjecTok, an innovative object tokenization framework. ObjecTok generates a single, holistic object token for each object in an image. This token is produced by a specially trained object encoder that embeds the object's semantic, positional, and shape information into a single compact representation, thereby preserving the object's integrity. To mitigate the imperfections of upstream object proposer models, we introduce learnable confidence embeddings. These embeddings enable the MLLM to learn the reliability of each object's information, significantly enhancing the model's robustness. Additionally, ObjecTok employs a hybrid input strategy, combining object tokens with traditional image patch tokens, allowing the model to leverage both object-level information and global scene context. By integrating ObjecTok into the LLaVA architecture, we achieve notable performance improvements on multiple object-centric benchmarks, effectively reducing object hallucinations and enhancing perception capabilities. Experimental results robustly demonstrate that the object tokens generated by our ObjecTok framework hold great potential for building more powerful and reliable MLLMs.
Xiyao Liu 0002, Lianqing Liu, Zhi Han
AAAI2
2026 Zero-shot single-image 3D generation via multi-grained semantic guidance
Xiyao Liu 0002, Xiai Chen, Lianqing Liu, Zhi Han
Neurocomputing2
2026 Unsupervised Domain Adaptive Object Detection via Semantic Consistency and Compactness Learning
abstract
Unsupervised domain adaptive object detection methods enhance model robustness in the target domain without requiring target-domain annotations. Despite notable progress, existing methods face two major challenges: 1) insufficient and inefficient learning of holistic feature consistency due to cumbersome pixel-level style matching and semantic discrepancy elimination between domains as well as the overlooking of their collaborative effect; and 2) unreliable learning of category feature compactness caused by poor-quality target-domain samples, inaccurate pseudo-labels and noisy cross-domain contrast paradigms. To address these challenges, we propose a novel Semantic Consistency and Compactness Learning (SCCL) network. For consistency learning, we introduce a Visual Adaptation-guided Semantic Alignment (VSA) module that achieves style matching through simple feature adaptation and incorporates a novel adversarial-free self-supervised method for feature disentanglement. The collaboration between these two aspects enables sufficient and efficient consistency learning. For reliable compactness learning, we develop a plug-and-play Instance Center-Contrastive (ICC) head that, for the first time, comprehensively addresses all three potential causes of unreliable learning through three integrated innovations, concerning sample pseudo-label quality enhancement, reliable sample storage and updating, and a robust sample contrast paradigm. Besides, the mutual reinforcement effect of VSA and ICC simultaneously enhances feature transferability and discriminability. Extensive experiments across four UDA object detection benchmarks with two baselines show that SCCL achieves superior adaptability and robustness. Code will be available at https://github.com/TooZE23/SCCL.
Yiming Su, Chunhui Hao, Xiyao Liu 0002, Jiandong Tian
IEEE Trans. Image Process.5
2025 Concept agent network for zero-base generalized few-shot learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
Appl. Intell.3
2025 Label smoothing regularization-based no hyperparameter domain generalization
Xiyao Liu 0002, Fupeng Chu, Zhi Han
Knowl. Based Syst.3
2025 DUAL-GDFQ: A Dual-Generator, Dual-Phase Learning Approach for Data-Free Quantization
abstract
Data-free quantization (DFQ) seeks to maximize the performance of quantized networks without requiring original training data. Conventional methods, which use synthetic samples from generators for network fine-tuning, often yield inferior results compared to training conducted with real data. To mitigate this problem, we introduce a dual-generator, dual-phase learning generative data-free quantization (DUAL-GDFQ) method, which utilizes two generators: a knowledge-matching generator and a knowledge-promoting generator for replicating the original data distribution as well as keeping samples informative. Additionally, inspired by meta-learning, the proposed novel dual-phase learning scheme can effectively utilize the capabilities of both generators by aligning their gradient descent directions. Theoretical analysis and extensive experiments demonstrate that our method successfully minimizes performance degradation in quantized networks and can achieve performance levels comparable to training with real data.
Zhi Han, Xiyao Liu 0002
IEEE Signal Process. Lett.3
2025 Frequency-Spatial Complementation: Unified Channel-Specific Style Attack for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) addresses the challenges of recognizing targets with out-of-domain data when only a few instances are available. Many current CD-FSL approaches primarily focus on enhancing the generalization capabilities of models in spatial domain, which neglects the role of the frequency domain in domain generalization. To take advantage of frequency domain in processing global information, we propose a Frequency-Spatial Complementation (FSC) model, which combines frequency domain information with spatial domain information to learn domain-invariant information from attacked data style. Specifically, we design a Frequency and Spatial Fusion (FusionFS) module to enhance the ability of the model to capture style-related information. Besides, we propose two attack strategies, i.e., the Gradient-guided Unified Style Attack (GUSA) strategy and the Channel-specific Attack Intensity Calculation (CAIC) strategy, which conduct targeted attacks on different channels to provide more diversified style data during the training phase, especially in single-source domain scenarios where the source domain data style is homogeneous. Extensive experiments across eight target domains demonstrate that our method significantly improves the model's performance under various styles.
Zhong Ji, Zhilong Wang 0001, Xiyao Liu 0002, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.3
2024 Unbiased Faster R-CNN for Single-source Domain Generalized Object Detection
abstract
Single-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm per-formance significantly. However, existing methods attempt to extract domain-invariant features, neglecting that the bi-ased data leads the network to learn biased features that are non-causal and poorly generalizable. To this end, we pro-pose an Unbiased Faster R-CNN (UFR) for generalizable feature learning. Specifically, we formulate SDG in object detection from a causal perspective and construct a Struc-tural Causal Model (SCM) to analyze the data bias andfeature bias in the task, which are caused by scene confounders and object attribute confounders. Based on the SCM, we de-sign a Global-Local Transformation module for data aug-mentation, which effectively simulates domain diversity and mitigates the data bias. Additionally, we introduce a Causal Attention Learning module that incorporates a designed at-tention invariance loss to learn image-level features that are robust to scene confounders. Moreover, we develop a Causal Prototype Learning module with an explicit instance constraint and an implicit prototype constraint, which fur-ther alleviates the negative impact of object attribute con-founders. Experimental results on five scenes demonstrate the prominent generalization ability of our method, with an improvement of 3.9% mAP on the Night-Clear scene.
Shijun Zhou, Xiyao Liu 0002, Chunhui Hao, Baojie Fan, Jiandong Tian
CVPR3
2024 On the Approximation Risk of Few-Shot Class-Incremental Learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Jungong Han
ECCV (51)3
2024 Semantic-Aware Dynamic Generation Networks for Few-Shot Human-Object Interaction Recognition
abstract
Recognizing human-object interaction (HOI) aims at inferring various relationships between actions and objects. Although great progress in HOI has been made, the long-tail problem and combinatorial explosion problem are still practical challenges. To this end, we formulate HOI as a few-shot task to tackle both challenges and design a novel dynamic generation method to address this task. The proposed approach is called semantic-aware dynamic generation networks (SADG-Nets). Specifically, SADG-Net first assigns semantic-aware task representations for different batches of data, which further generates dynamic parameters. It obtains the features that highlight intercategory discriminability and intracategory commonality adaptively. In addition, we also design a dual semantic-aware encoder module (DSAE-Module), that is, verb-aware and noun-aware branches, to yield both action and object prototypes of HOI for each task space, which generalizes to novel combinations by transferring similarities among interactions. Extensive experimental results on two benchmark datasets, that is, humans interacting with common objects (HICO)-FS and trento universal HOI (TUHOI)-FS, illustrate that our SADG-Net achieves superior performance over state-of-the-art approaches, which proves its impressive effectiveness on few-shot HOI recognition.
Zhong Ji, Xiyao Liu 0002, Changxin Gao, Yanwei Pang, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Self-taught cross-domain few-shot learning with weakly supervised object localization and task-decomposition
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han
Knowl. Based Syst.1
2023 Dual Distillation Discriminator Networks for Domain Adaptive Few-Shot Learning
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han
Neural Networks1
2023 Memorizing Complementation Network for Few-Shot Class-Incremental Learning
abstract
Few-shot Class-Incremental Learning (FSCIL) aims at learning new concepts continually with only a few samples, which is prone to suffer the catastrophic forgetting and overfitting problems. The inaccessibility of old classes and the scarcity of the novel samples make it formidable to realize the trade-off between retaining old knowledge and learning novel concepts. Inspired by that different models memorize different knowledge when learning novel concepts, we propose a Memorizing Complementation Network (MCNet) to ensemble multiple models that complements the different memorized knowledge with each other in novel tasks. Additionally, to update the model with few novel samples, we develop a Prototype Smoothing Hard-mining Triplet (PSHT) loss to push the novel samples away from not only each other in current task but also the old distribution. Extensive experiments on three benchmark datasets, e.g., CIFAR100, miniImageNet and CUB200, have demonstrated the superiority of our proposed method.
Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.3
2022 Teachers cooperation: team-knowledge distillation for multiple cross-domain few-shot learning
Zhong Ji, Jingwei Ni, Xiyao Liu 0002, Yanwei Pang
Frontiers Comput. Sci.3
2022 Meta hyperbolic networks for zero-shot learning
Yan Xu 0016, Lifu Mu, Zhong Ji, Xiyao Liu 0002, Jungong Han
Neurocomputing4
2022 DGIG-Net: Dynamic Graph-in-Graph Networks for Few-Shot Human-Object Interaction
abstract
Few-shot learning (FSL) for human-object interaction (HOI) aims at recognizing various relationships between human actions and surrounding objects only from a few samples. It is a challenging vision task, in which the diversity and interactivity of human actions result in great difficulty to learn an adaptive classifier to catch ambiguous interclass information. Therefore, traditional FSL methods usually perform unsatisfactorily in complex HOI scenes. To this end, we propose dynamic graph-in-graph networks (DGIG-Net), a novel graph prototypes framework to learn a dynamic metric space by embedding a visual subgraph to a task-oriented cross-modal graph for few-shot HOI. Specifically, we first build a knowledge reconstruction graph to learn latent representations for HOI categories by reconstructing the relationship among visual features, which generates visual representations under the category distribution of every task. Then, a dynamic relation graph integrates both reconstructible visual nodes and dynamic task-oriented semantic information to explore a graph metric space for HOI class prototypes, which applies the discriminative information from the similarities among actions or objects. We validate DGIG-Net on multiple benchmark datasets, on which it largely outperforms existing FSL approaches and achieves state-of-the-art results.
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001
IEEE Trans. Cybern.1
2022 Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot Learning
abstract
Semantic information provides intra-class consistency and inter-class discriminability beyond visual concepts, which has been employed in Few-Shot Learning (FSL) to achieve further gains. However, semantic information is only available for labeled samples but absent for unlabeled samples, in which the embeddings are rectified unilaterally by guiding the few labeled samples with semantics. Therefore, it is inevitable to bring a cross-modal bias between semantic-guided samples and nonsemantic-guided samples, which results in an information asymmetry problem. To address this problem, we propose a Modal-Alternating Propagation Network (MAP-Net) to supplement the absent semantic information of unlabeled samples, which builds information symmetry among all samples in both visual and semantic modalities. Specifically, the MAP-Net transfers the neighbor information by the graph propagation to generate the pseudo-semantics for unlabeled samples guided by the completed visual relationships and rectify the feature embeddings. In addition, due to the large discrepancy between visual and semantic modalities, we design a Relation Guidance (RG) strategy to guide the visual relation vectors via semantics so that the propagated information is more beneficial. Extensive experimental results on three semantic-labeled datasets, i.e., Caltech-UCSD-Birds 200-2011, SUN Attribute Database and Oxford 102 Flower, have demonstrated that our proposed method achieves promising performance and outperforms the state-of-the-art approaches, which indicates the necessity of information symmetry.
Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.3
2022 Task-Oriented High-Order Context Graph Networks for Few-Shot Human-Object Interaction Recognition
abstract
Few-shot human-object interaction (FS-HOI) recognition aims at inferring new interactions between human actions and surrounding objects merely with a few available instances. It is beneficial to alleviate the long-tail and combinatorial explosion problems in human-object interaction (HOI). Nevertheless, the existing FS-HOI methods only focus on modeling the relationships between labeled samples and unlabeled samples in the Euclidean domain, which neglects the rich relational structures of the visual information among labeled samples and between human actions and objects. Accordingly, we tackle the few-shot HOI task in the non-Euclidean domain and present a graph-based model, namely, task-oriented high-order context graph network (THCG-Net). It contains a task attention module (TA-Module) and a high-order context graph module (HG-Module). In TA-Module, an attention mechanism is designed by utilizing task information to build a task-oriented space, in which the discriminative information for the current task (episode) is captured by embedding the visual features into the task-oriented space. The HG-Module is proposed to construct a task-level graph and takes the context information as high-order knowledge, which provides discriminative guidance for propagating visual information. It captures the discriminability among different categories while highlights the commonality of related categories adaptively, which effectively transfers knowledge to related categories. Extensive experimental results on two benchmark datasets, HICO-FS and TUHOI-FS, are provided. It demonstrates that our THCG-Net significantly outperforms the state-of-the-art approaches, which proves its impressive effectiveness in recognizing various human actions and surrounding objects in few-shot scenarios.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Ling Shao 0001, Zhongfei Zhang
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Few-Shot Human-Object Interaction Recognition With Semantic-Guided Attentive Prototypes Network
abstract
Extreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsical characteristic of HOI is diverse and interactive, we propose a Semantic-guided Attentive Prototypes Network (SAPNet) framework to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design two alternative prototypes calculation methods, i.e., Prototypes Shift (PS) approach and Hallucinatory Graph Prototypes (HGP) approach, which explore to learn a suitable category prototypes representations in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize 2 HOI benchmark datasets with 2 split strategies, i.e., HICO-NN, TUHOI-NN, HICO-NF, and TUHOI-NF. Extensive experimental results on these datasets have demonstrated the effectiveness of our proposed SAPNet approach.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Wanli Ouyang, Xuelong Li 0001
IEEE Trans. Image Process.2
2020 SGAP-Net: Semantic-Guided Attentive Prototypes Network for Few-Shot Human-Object Interaction Recognition
abstract
Extreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsic characteristic of HOI is diverse and interactive, we propose a Semantic-Guided Attentive Prototypes Network (SGAP-Net) to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design a novel decision method to alleviate the biases produced by different patterns of the same action in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize two HOI benchmark datasets, i.e., HICO-FS and TUHOI-FS, to realize the task of few-shot HOI. Extensive experimental results on both datasets have demonstrated the effectiveness of our proposed SGAP-Net approach.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
AAAI2