VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Liu 0015
dblp:42/7844-15
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-1168-6645ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ModSolAgent: Automated Finite Element Code Generation for Abaqus via LLM-Based AgentabstractFinite element simulation and solution process represents a critical component in engineering analysis. While large language models (LLMs) have demonstrated remarkable capabilities in general-purpose code generation from textual descriptions, their application to generating structured and specialized finite element simulation scripts presents unique challenges. A key challenge is AI-generated hallucination, as these tasks require precise intent analysis, complex planning and reasoning, and strict adherence to logical consistency during execution. This often results in outputs that seem convincing but are ultimately erroneous. To address this challenge, we propose ModSolAgent, an LLM-based agent that generates Python scripts for Abaqus software to perform finite element modeling and solving tasks. ModSolAgent ensures the accuracy and logical coherence of code generation by leveraging a structured reasoning instruction, dynamic retrieval guidance template, and iterative verification generation. Experimental results demonstrate that LLMs augmented by ModSolAgent achieve an 83.3% success rate on real-world finite element simulation tasks, effectively meeting most Abaqus scripting requirements while significantly outperforming baseline models. To further enhance accessibility, we construct the AbqInstruct dataset by distilling knowledge from ModSolAgent to fine-tune the lightweight open-source models. Experiments show that fine-tuning on AbqInstruct leads to substantial performance improvements, with the lightweight model achieving proficiency across most finite element modeling and solving tasks. This work establishes a paradigm for integrating LLMs with specialized engineering software and providing novel insights for other structured, domain-specific code generation scenarios in industrial applications. Zidi Li, Hong-Wei Ge, Guozhi Tang, Yuxuan Liu 0015 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person RetrievalabstractText-to-image person retrieval aims to identify target person images using natural language descriptions. Current state-of-the-art methods predominantly rely on single-round retrieval frameworks, where retrieval accuracy heavily depends on the quality of the initial textual descriptions. However, users sometimes struggle to provide detailed and distinctive descriptions in a single attempt, resulting in generic initial queries that lack discriminative details. This fundamental limitation of the single-round retrieval framework frequently leads to the misinterpretation of user intent and suboptimal retrieval performance. To address this limitation, we propose Dialogue-driven Interactive Dynamic Learning (DIDL) for text-to-image person retrieval. Specifically, we first introduce Collaborative Query Refinement (CQR), which progressively refines retrieval conditions through multi-round dialogues. Then, we design Dynamic Context Resampling (DCR) based on a bi-granular mask strategy that enhances the model's adaptation to dialogue-style contexts and effectively balances its attention between initial descriptions and supplementary information. Based on these components, we further propose cross-modal Probabilistic Context Matching Modeling (ProCMM) that establishes effective associations between static visual features and dynamic contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across all three benchmark datasets. Hong-Wei Ge, Yuxuan Liu 0015, Yaqing Hou |
ACM Multimedia | 3 |
| 2025 | Hypergraph-driven soft semantics flexible learning for visible-infrared person re-identification
Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu, Jiulin Fan |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Multiscale Recovery Diffusion Model With Unsupervised Learning for Video Anomaly Detection SystemabstractThe rapid development of intelligent industry and smart city increases the number of surveillance devices, greatly enhancing the need for unsupervised automatic anomaly detection in real-time video surveillance, which uses raw data without laborious manual annotations. Existing video anomaly detection (VAD) methods encounter limitations when utilizing pretext tasks, such as reconstruction or prediction to identify abnormal events, as these tasks are not completely consistent and complementary with the essential objective of anomaly detection. Motivated by recent advances in diffusion models, we propose a multiscale recovery diffusion model, which relies on the proposed novel and effective pretext task named recovery to introduce the notion of generation speed. It utilizes critical step-by-step generation of diffusion probabilistic models in unsupervised anomaly detection scenarios. By incorporating a proposed multiscale spatial-temporal subtraction module, our model captures more detailed appearance and motion information of foreground objects without relying on other high-level pretrained models. Furthermore, an innovative push–pull loss further extends the disparity between normal and abnormal events through pseudolabels. We validate our model on five established benchmarks: UCSD Ped1, UCSD Ped2, CUHK Avenue, ShanghaiTech, and UCF-Crime, achieving frame-level area under the curves of 86.01%, 99.23%, 92.35%, 82.49%, and 74.79%, respectively, surpassing other state-of-the-art unsupervised VAD methods. Hong-Wei Ge, Yuxuan Liu 0015, Guozhi Tang |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Find Hidden Modality Divergence: Adversarial Aware Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identifi-cation (Unsupervised VI-ReID) aims to learn discriminative identity features under the large modality gap without any labeled data. Currently, the state-of-the-art methods optimize cross-modality differences by using contrastive learning as the underlying paradigm. However, they neglect the problem of modality divergence during the cross-modality optimization process. This problem means that the interclass instances between the cross-modality intraclass gaps can make cross-modality intraclass instances difficult to get closer to each other in the feature space due to the effect of contrastive learning on these interclass instances. To alleviate the negative impact of the modality divergence problem, we propose an adversarial aware learning (ADAL) framework to explore the instances that generate modal divergence and adversarially optimize these explored instances. Specifically, on the one hand, we explore the optimization directions of each cluster during the cross-modality optimization process, and the cluster centroids generating positive optimization are facilitated, while the others generating negative optimization are penalized. On the other hand, we further consider the instance-level optimization process, which increases the affinities of the positive instance pairs with large cross-modality gaps to further improve the centroid-level optimization. Extensive experiments conducted on the visible-infrared person Re-ID datasets show that the proposed method is used as a universally applicable plug-in module to add the existing unsupervised VI-ReID methods, which outperforms the existing state-of-the-art approaches. Yuxuan Liu 0015, Hong-Wei Ge, Chunguo Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Ought to be salient and hidden: soft biological semantic-guided explicit-implicit learning for cloth-changing person re-identification
Wanda Zeng, Hong-Wei Ge, Yuxuan Liu 0015 |
Vis. Comput. | 3 |
| 2024 | Occluded Person Reidentification via a Universal Framework With Difference Consistency Guidance LearningabstractOccluded person reidentification (Re-ID) aims at learning discriminative identity features to match person images under the interference of occlusion situations in video surveillance of the visual Internet of Things (VIoT). Currently, occluded person Re-ID methods have made impressive improvements in nonperson occlusion situations. However, in real-world scenarios, the target person is commonly occluded by other nontarget persons, and the fine-grained differences between the persons make the model difficult to distinguish the discriminative identity features. To this end, we propose a difference consistency guidance (DCG) learning to enlarge the fine-grained differences by the guidance of the coarse-grained differences, which can distinguish the discriminative identity features in nontarget person occlusion situations. Then, DCG reduces the identity feature representation of the nonperson occlusion instances, which further improves the ability of the model in nonperson occlusion situations and improves the guidance ability of coarse-grained difference. Moreover, DCG can enhance the robustness of the model in the unsupervised occluded person Re-ID task and further improve the universal applicability of the model. Extensive experiment results under the supervised and unsupervised settings demonstrate the DCG outperforms the state-of-the-art methods in experiments conducted on the occluded person Re-ID benchmarks. Yuxuan Liu 0015, Hong-Wei Ge, Guozhi Tang |
IEEE Internet Things J. | 1 |
| 2024 | Localization and saturation of degradation space for weakly-supervised real-world super-resolution
Guozhi Tang, Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu |
Knowl. Based Syst. | 3 |
| 2024 | Representation Robustness and Feature Expansion for Exemplar-Free Class-Incremental LearningabstractDespite deep neural networks have made outstanding achievements in many static tasks, when faced with a continuous stream of data, they suffer from catastrophic forgetting since the previous data is usually inaccessible. Stored data or generative model is commonly used for maintaining the model performance but with memory utilization and privacy safety issues. Prototype-based methods address these issues by keeping only one prototype for each class but with limitations in its ability to trade-off the model stability and plasticity. In this paper, a novel exemplar-free class-incremental learning method is proposed which improves the stability of the representation learning and the decision boundary to a great degree. First, based on the results of our exploration into the impact of the batch normalization (BN) layer on representation learning, we propose to remove the BN layer (RBNL) in the incremental training phase to improve the stability of model representation learning. Then, to further maintain the feature space, we design the prototype mixing (PM), which expands the deep features by randomly and linearly combining prototypes of the old classes to generate hybrid prototypes with composite labels for fine-tuning the fully connected layer. Experimental results on three benchmark datasets, CIFAR-100, TinyImageNet, and ImageNet, show that our proposed method can effectively balance the stability and plasticity of the model, and outperforms the state-of-the-art works. Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Discriminative Identity-Feature Exploring and Differential Aware Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to learn discriminative representations for person retrieval from unlabeled data. Currently, state-of-the-art techniques accomplish this task by using instance contrastive learning, which contrasts the similarities of the instances in different views. However, existing contrastive methods only focus on the positive effects of inter-instance relationships, while neglecting the negative effects of intra-instance redundancy information. This redundancy information can generate invalid or spurious intra-class relationships during the instance contrasting process, which enlarges the intra-class gaps and increases the noisy pseudo-labels. To address this issue, we propose a discriminative identity-feature exploring and differential aware learning (DiDAL) framework to learn more discriminative intra-identity representations. Specifically, the DiDAL extracts intra-instance salient features by synthetic complementary attention, and further explores the discriminative identity features by modeling the relationship among these salient features based on graph neural networks. This strategy aims to reduce the intra-instance redundancy information. Moreover, DiDAL explores hard instances by leveraging the extracted intra-instance salient features, and matches an anchor with multiple hard positive instances to enhance the robustness of the model to noisy pseudo-labels. Extensive experiment results on two widely used person re-identification datasets and a vehicle re-identification dataset demonstrate the superiority of the proposed method compared with existing state-of-the-art methods. Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | Clothes-Changing Person Re-Identification via Universal Framework With Association and Forgetting LearningabstractClothes-changing person re-identification (Re-ID) aims at learning identity-relevant feature representations among clothing-changed persons. Currently, the state-of-the-art methods accomplish this task by using additional assistance (e.g., silhouettes, sketches, clothes labels, etc.) to explore identity-relevant information. However, humans do not require redundant assistance information to retrieve clothing-changed persons. It is commonly known that humans can recall targets they have seen before with a simple reminder. Inspired by human perception, we propose an association and forgetting learning (AFL) framework for clothes-changing person re-identification. Specifically, on the one hand, during the association learning process, the AFL framework constructs association factors for each identity to simulate the reminders found in human perception. Then, the original instances and the explored hardest positive instances are cross-correlated by the association factors to learn identity-relevant features. On the other hand, the model is forced to forget the identity-irrelevant features by the proposed forgetting learning module, which improves the intra-class compactness. Finally, we further propose a clustering relationship exploration (CRE) module to optimize the cluster distribution of clothes-changing instances, which enables AFL to also be effectively applied in unsupervised settings, improving the universal applicability of the model. Extensive experiment results obtained on clothes-changing person Re-ID datasets under supervised and unsupervised settings demonstrate the superiority of the proposed method over the existing state-of-the-art methods. Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002 |
IEEE Trans. Multim. | 1 |
| 2023 | Camera-aware progressive learning for unsupervised person re-identification
Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
Neural Comput. Appl. | 1 |
| 2023 | Attention-guided spatial-temporal graph relation network for video-based person re-identification
Hong-Wei Ge, Wenbin Pei, Yuxuan Liu 0015, Yaqing Hou, Liang Sun 0003 |
Neural Comput. Appl. | 4 |
| 2023 | AGPN: Action Granularity Pyramid Network for Video Action RecognitionabstractVideo action recognition is a fundamental task for video understanding. Action recognition in complex spatio-temporal contexts generally requires fusing of different multi-granularity action information. However, existing works do not consider spatio-temporal information modeling and fusion from the perspective of action granularity. To address this problem, this paper proposes an Action Granularity Pyramid Network (AGPN) for action recognition, which can be flexibly integrated into 2D backbone networks. The core module is the Action Granularity Pyramid Module (AGPM), a hierarchical pyramid structure with residual connections, which is established to fuse multi-granularity action spatio-temporal information. From top to bottom level in the designed pyramid structure, the receptive field decreases and action granularity becomes more refined. To enrich temporal information of the inputs, a Multiple Frame Rate Module (MFM) is proposed to mix different frame rates at a fine-grained pixel-wise level. Moreover, a Spatio-temporal Anchor Module (SAM) is employed to fix spatio-temporal feature anchors to promote the effectiveness of feature extraction. We conduct extensive experiments on three large-scale action recognition datasets, Something-Something V1 & V2 and Kinetics-400. The results demonstrate that our proposed AGPN outperforms the state-of-the-art methods for the tasks of video action recognition. Yatong Chen 0001, Hong-Wei Ge, Yuxuan Liu 0015, Xinye Cai, Liang Sun 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Complementary Attention-Driven Contrastive Learning With Hard-Sample Exploring for Unsupervised Domain Adaptive Person Re-IDabstractUnsupervised domain adaptive (UDA) methods for person re-identification (Re-ID) aim to transfer the knowledge of the labeled source domain to the unlabeled target domain without further annotations, which is challenging due to the drift of label distribution and the missing of target domain labels. Improving the clustering accuracy of pseudo-labels can help the model fit the target domain. However, the errors of pseudo-label noise will be accumulated during training, which is harmful to the model performance. Moreover, the hard samples can lead to a large gap between intra-class features and a small gap between inter-class features. To address these problems, this paper proposes a complementary attention-driven contrastive learning with hard-sample exploring (CACHE) algorithm. In CACHE, on one hand, the complementary attention module is used to improve the discriminability of the features. The obtained discriminative features can reduce noisy pseudo-labels and improve the clustering accuracy of pseudo labels; On the other hand, we explore the hard samples based on the instance relationship and cluster relationship for contrastive learning. This way can make the cluster more compact. Extensive experiments on three large-scale person re-identification benchmarks demonstrate the effectiveness of the proposed method, which significantly outperforms state-of-the-art methods in terms of mAP and CMC. Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |