Siyuan Chen 0005

dblp:84/5999-5 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-9272-4804ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language Models
abstract
Knowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines.
Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang
AAAI5
2026 Geometry-constrained open set recognition with frozen foundation model features for industrial inspection
Yingjun Xiao, Luyu Xie, Siyuan Chen 0005, Xiangjun Xiao, Lingxi Peng
Expert Syst. Appl.4
2026 Skew-normal distributions for modeling asymmetric moving tendencies in pedestrian trajectories
Siyuan Chen 0005, Yatie Xiao, Yangtao Wang, Yanzhao Xie, Tong Zhu 0003, Jinbiao Chen
Neurocomputing1
2026 LLM-augmented entity alignment: an unsupervised and training-free framework
abstract
Entity alignment (EA) is a fundamental task in knowledge graph (KG) integration, aiming to identify equivalent entities across different KGs for a unified and comprehensive representation. Recent advances have explored pre-trained language models (PLMs) to enhance the semantic understanding of entities, achieving notable improvements. However, existing methods face two major limitations. First, they rely heavily on human-annotated labels for training, leading to high computational costs and poor scalability. Second, some approaches use large language models (LLMs) to predict alignments in a multi-choice question format, but LLM outputs may deviate from expected formats, and predefined options may exclude correct matches, leading to suboptimal performance. To address these issues, we propose LEA, an LLM-augmented entity alignment framework that eliminates the need for labeled data and enhances robustness by mitigating information heterogeneity at both embedding and semantic levels. LEA first introduces an entity textualization module that transforms structural and textual information into a unified format, ensuring consistency and improving entity representations. It then leverages LLMs to enrich entity descriptions, enhancing semantic distinctiveness. Finally, these enriched descriptions are encoded into a shared embedding space, enabling efficient alignment through text retrieval techniques. To balance performance and computational cost, we further propose a selective augmentation strategy that prioritizes the most ambiguous entities for refinement. Experimental results on both homogeneous and heterogeneous KGs demonstrate that LEA outperforms existing models trained on 30 % labeled data, achieving a 30 % absolute improvement in Hit@1 score. As LLMs and text embedding models advance, LEA is expected to further enhance EA performance, providing a scalable and robust paradigm for practical applications. The code and dataset can be found at https://github.com/Longmeix/LEA.
Meixiu Long, Jiahai Wang, Junxiao Ma, Jianpeng Zhou, Siyuan Chen 0005
Neural Networks5
2026 Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002
Pattern Recognit.4
2026 DSA-GNN: Optimizing long-tail in graph structures via degree suppression with attention
Guanwei Huang, Qingxiao Guan, Can Tian, Siyuan Chen 0005, Wensheng Zhang 0002
Pattern Recognit.5
2026 MKGPL: graph prompt learning with multi-view knowledge for few-shot recognition
Yanzhao Xie, Man Qiu, Yangtao Wang, Siyuan Chen 0005, Meie Fang, Maobin Tang, Wensheng Zhang 0002
Pattern Recognit.4
2026 High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream Transformers
abstract
Recently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git .
Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2026 Harnessing Transferable Adversarial Examples via Multilayer Attention-Guided Spatial Transformations
abstract
Transfer-based adversarial attacks are key for evaluating the robustness of deep neural networks (DNNs) in black-box settings, yet their effectiveness is often constrained by limited cross-model transferability. Existing feature-level approaches typically rely on single-layer attention guidance or static perturbation patterns, which restrict adaptability across diverse architectures. In this work, we introduce a unified adversarial framework, named Multi-layer Attention-guided Spatial Transformations (MAT), to exploit class-discriminative cues from multiple feature layers to craft highly transferable adversarial examples. MAT integrates Multi-layer Attention Fusion (MAF) to capture complementary low-level and high-level semantics from multiple intermediate layers, Attention-guided Augmentation (AGA) to selectively perturb non-critical regions while preserving semantic integrity, and Spatial Random Transformation (SRT) to introduce stochastic spatial augmentations to diversify patterns during optimization. Unlike prior methods that use static or layer-specific attention, MAT dynamically adapts feature guidance to the architecture and task, which enhances generalization. We evaluate MAT against eleven state-of-the-art (SOTA) transfer-based attacks across nine CNN-based and Transformer-based architectures on ImageNet. Comprehensive experiments demonstrate that MAT consistently outperforms eleven state-of-the-art transfer-based attacks in both white-box and black-box settings, including against adversarially trained and input preprocessing-based defensive models, while maintaining higher semantic similarity to the original inputs. It highlights the superior adversarial robustness and excellent adaptability of MAT in adversarial machine learning. Our code is available athttps://github.com/dislab-gzhu/MAT.
Pengfei Dong, Yatie Xiao, Chi-Man Pun, Fei Peng 0001, Kongyang Chen, Qingxian Guan, Siyuan Chen 0005, Xiangyu Ye, Zhenbang Liu
IEEE Trans. Reliab.7
2025 Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text Retrieval
abstract
Image-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA.
Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016
ICME8
2025 Balanced residual distillation learning for 3D point cloud class-incremental semantic segmentation
Yuanzhi Su, Siyuan Chen 0005, Yuan-Gen Wang
Expert Syst. Appl.2
2025 Towards adversarial patch attacks on deep crowd-counting networks via density-aware normalized feature learning
Yatie Xiao, Siyuan Chen 0005, Kongyang Chen, Qingxiao Guan, Zhenbang Liu
Knowl. Based Syst.2
2025 A Hierarchical Framework With Spatio-Temporal Consistency Learning for Emergence Detection in Complex Adaptive Systems
abstract
Emergence, a global property of complex adaptive systems (CASs) constituted by interactive agents, is prevalent in real-world dynamic systems, e.g., network-level traffic congestions. Detecting its formation and evaporation helps to monitor the state of a system, allowing it to issue a warning signal for harmful emergent phenomena. Since there is no centralized controller of CAS, detecting emergence based on each agent's local observation is desirable but challenging. Existing works are unable to capture emergence-related spatial patterns, and fail to model the nonlinear relationships among agents. This article proposes a hierarchical framework with spatio-temporal consistency learning (HSTCL) to solve these two problems by learning the system representation and agent representations, respectively. Spatio-temporal encoders (STEs) composed of spatial and temporal transformers are designed to capture agents' nonlinear relationships and the system's complex evolution. Agents' and the system's representations are learned to preserve the spatio-temporal consistency by minimizing the spatial and temporal dissimilarities in a self-supervised manner in the latent space. Our method achieves more accurate detection than traditional methods and deep learning methods on three datasets with well-known yet hard-to-detect emergent behaviors. Notably, our hierarchical framework is generic in incorporating other deep learning methods for agent-level and system-level detection.
Siyuan Chen 0005, Jiahai Wang
IEEE Trans. Neural Networks Learn. Syst.1
2024 Locally-adaptive mapping for network alignment via meta-learning
Meixiu Long, Siyuan Chen 0005, Jiahai Wang
Inf. Process. Manag.2
2024 Heterogeneous Interaction Modeling With Reduced Accumulated Error for Multiagent Trajectory Prediction
abstract
Dynamical complex systems composed of interactive heterogeneous agents are prevalent in the world, including urban traffic systems and social networks. Modeling the interactions among agents is the key to understanding and predicting the dynamics of the complex system, e.g., predicting the trajectories of traffic participants in the city. Compared with interaction modeling in homogeneous systems such as pedestrians in a crowded scene, heterogeneous interaction modeling is less explored. Worse still, the error accumulation problem becomes more severe since the interactions are more complex. To tackle the two problems, this article proposes heterogeneous interaction modeling with reduced accumulated error (HIMRAE) for multiagent trajectory prediction. Based on the historical trajectories, our method infers the dynamic interaction graphs among agents, featured by directed interacting relations and interacting effects. A heterogeneous attention mechanism (HAM) is defined on the interaction graphs for aggregating the influence from heterogeneous neighbors to the target agent. To alleviate the error accumulation problem, this article analyzes the error sources from the spatial and temporal perspectives, and proposes to introduce the graph entropy and the mixup training strategy for reducing the two types of errors, respectively. Our method is examined on three real-world datasets containing heterogeneous agents, and the experimental results validate the superiority of our method.
Siyuan Chen 0005, Jiahai Wang
IEEE Trans. Neural Networks Learn. Syst.1
2023 Efficient Meta Neural Heuristic for Multi-Objective Combinatorial Optimization
abstract
Recently, neural heuristics based on deep reinforcement learning have exhibited promise in solving multi-objective combinatorial optimization problems (MOCOPs). However, they are still struggling to achieve high learning efficiency and solution quality. To tackle this issue, we propose an efficient meta neural heuristic (EMNH), in which a meta-model is first trained and then fine-tuned with a few steps to solve corresponding single-objective subproblems. Specifically, for the training process, a (partial) architecture-shared multi-task model is leveraged to achieve parallel learning for the meta-model, so as to speed up the training; meanwhile, a scaled symmetric sampling method with respect to the weight vectors is designed to stabilize the training. For the fine-tuning process, an efficient hierarchical method is proposed to systematically tackle all the subproblems. Experimental results on the multi-objective traveling salesman problem (MOTSP), multi-objective capacitated vehicle routing problem (MOCVRP), and multi-objective knapsack problem (MOKP) show that, EMNH is able to outperform the state-of-the-art neural heuristics in terms of solution quality and learning efficiency, and yield competitive solutions to the strong traditional heuristics while consuming much shorter time.
Jinbiao Chen, Jiahai Wang, Zizhen Zhang, Zhiguang Cao, Te Ye, Siyuan Chen 0005
NeurIPS6
2023 Multi-Agent Meta-Reinforcement Learning with Coordination and Reward Shaping for Traffic Signal Control
Jiahai Wang, Siyuan Chen 0005
PAKDD (2)3
2023 Enhanced edge convolution-based spatial-temporal network for network traffic prediction
Zehua Hu, Ke Ruan, Weihao Yu 0002, Siyuan Chen 0005
Appl. Intell.4
2022 Graph Neural Networks with Dynamic and Static Representations for Social Recommendation
Junfa Lin, Siyuan Chen 0005, Jiahai Wang
DASFAA (2)2
2022 Multiple userids identification with deep learning
Siyuan Chen 0005, Zhiyue Liu, Jiahai Wang
Expert Syst. Appl.2
2021 Neural Relational Inference with Efficient Message Passing Mechanisms
abstract
Many complex processes can be viewed as dynamical systems of interacting agents. In many cases, only the state sequences of individual agents are observed, while the interacting relations and the dynamical rules are unknown. The neural relational inference (NRI) model adopts graph neural networks that pass messages over a latent graph to jointly learn the relations and the dynamics based on the observed data. However, NRI infers the relations independently and suffers from error accumulation in multi-step prediction at dynamics learning procedure. Besides, relation reconstruction without prior knowledge becomes more difficult in more complex systems. This paper introduces efficient message passing mechanisms to the graph neural networks with structural prior knowledge to address these problems. A relation interaction mechanism is proposed to capture the coexistence of all relations, and a spatio-temporal message passing mechanism is proposed to use historical information to alleviate error accumulation. Additionally, the structural prior knowledge, symmetry as a special case, is introduced for better relation prediction in more complex systems. The experimental results on simulated physics systems show that the proposed method outperforms existing state-of-the-art methods.
Siyuan Chen 0005, Jiahai Wang
AAAI1
2021 A Semi-supervised Framework with Efficient Feature Extraction and Network Alignment for User Identity Linkage
Zehua Hu, Jiahai Wang, Siyuan Chen 0005
DASFAA (2)3
2021 Multi-agent Deep Reinforcement Learning with Spatio-Temporal Feature Fusion for Traffic Signal Control
Jiahai Wang, Siyuan Chen 0005, Zhiyue Liu
ECML/PKDD (4)3
2020 A Novel Framework with Information Fusion and Neighborhood Enhancement for User Identity Linkage
abstract
User identity linkage across social networks is an essential problem for cross-network data mining. Since network structure, profile and content information describe different aspects of users, it is critical to learn effective user representations that integrate heterogeneous information. This paper proposes a novel framework with INformation FUsion and Neighborhood Enhancement (INFUNE) for user identity linkage. The information fusion component adopts a group of encoders and decoders to fuse heterogeneous information and generate discriminative node embeddings for preliminary matching. Then, these embeddings are fed to the neighborhood enhancement component, a novel graph neural network, to produce adaptive neighborhood embeddings that reflect the overlapping degree of neighborhoods of varying candidate user pairs. The importance of node embeddings and neighborhood embeddings are weighted for final prediction. The proposed method is evaluated on real-world social network data. The experimental results show that INFUNE significantly outperforms existing state-of-the-art methods.
Siyuan Chen 0005, Jiahai Wang, Yanqing Hu
ECAI1