Xinyang Song

dblp:282/2113 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
abstract
The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation.
Xinyang Song, Weining Wang 0001, Shaozhen Liu, Jingdong Chen, Qi Li 0005, Zhenan Sun
AAAI1
2025 Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
abstract
Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently in a scenario. To tackle this, we propose a novel multi-character video generation framework in a tuning-free manner, which is based on the separated text and pose guidance. Specifically, we first extract character masks from the pose sequence to identify the spatial position for each character, and then single prompts for each character are obtained with LLMs for precise text guidance. Moreover, the spatial-aligned cross attention and multi-branch control module are proposed to generate fine-grained controllable multi-character video. The visualized results of generating video demonstrate the precise controllability of our method for multicharacter generation. We also verify the generality of our method by applying it to various personalized T2I models. Moreover, the quantitative results show that our approach achieves superior performance compared with previous works.
Beiyuan Zhang, Chunlei Fu, Xinyang Song, Zhenan Sun
ICASSP4
2024 Context Spatial Awareness Remote Sensing Image Change Detection Network Based on Graph and Convolution Interaction
abstract
Remote sensing images are characterized by high dimensionality, complex textures, and large scales. Traditional Convolutional Neural Network (CNN) methods may overlook spatial relationships and contextual information among pixels when dealing with remote sensing data. Therefore, Graph Convolutional Networks (GCN) have emerged as a promising solution. In this paper, we propose a Contextual Spatial Awareness Remote Sensing Image Change Detection Network Based on Graph and Convolution interaction (CSAGC). We aim to enhance the handling of contextual information by introducing multiple augmentation modules. In CSAGC, we propose a high-performance encoder called Congraph that integrates a CNN and a Graph Neural Network (GNN). By preserving the respective features of both branches, we effectively fuse local detailed features and global positional features, achieving superior feature extraction capabilities. Additionally, we design two modules to facilitate the integration of multiscale spatial information: Contextual Spatial Awareness Module (CSAM) and Spatial Integration Module (SIM). CSAM, a crucial module connecting the encoder and decoder, jointly explores contextual features using the current feature branch and high-low level feature branches, leveraging spatial positional information for better content acquisition. SIM, located in the decoder module, aims to integrate the multiscale information outputted by CSAM, complementing the contextual information and improving the overall network’s ability to capture spatial contextual information. We conducted extensive experiments on three datasets, namely LEVIR-CD, WHU-CD, and GZ-CD. The experimental results demonstrate that CSAGC exhibits excellent performance, achieving significant performance improvements compared to state-of-the-art (SOTA) methods.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Trust Management Strategy for Digital Twins in Vehicular Ad Hoc Networks
abstract
As an essential part of mobile networks, vehicular ad hoc networks (VANETs) are beneficial to the improvement of traffic efficiency and safety through real-time information sharing between vehicles. Digital Twins (DT) have been utilized to facilitate the design, testing, and deployment of VANETs. However, constructing Digital Twins still faces interference from malicious vehicles. Despite most vehicles following communication rules honestly, the reliability and authenticity of traffic messages cannot be guaranteed due to the network’s openness and vulnerability. Meanwhile, vehicles may suffer tracking attacks during the interaction without an effective privacy-preserving method, leading to the leakage of sensitive data. To address these issues, a decentralized trust management scheme embedded with blockchain that considers identity authentication is proposed to detect malicious DT-vehicles. In our method, each vehicle in the Digital Twin of VANETs (DT-VANETs) is equipped with a certificate recorded on the blockchain as a legal identity, which is also served as a pseudonym for security during message transmission. The trustworthiness of the vehicle is evaluated based on direct trust and recommendation trust. Direct interaction between vehicles consists of message authenticity verification and active detection, which are the basis of direct trust calculation. For other vehicles, these direct trust opinions are treated as second-hand information to obtain recommendation trust. Unreliable recommendations are filtered by our proposed RTF algorithm, further resisting cooperation attacks. Vehicles judged to be malicious will have their certificates revoked and removed from DT-VANETs, providing a guarantee for the establishment of trust in DT-VANETs. Experimental results show that the proposed scheme can effectively resist malicious attacks in DT-VANETs.
Bohan Li 0001, Xinyang Song, Tianlun Dai, Xiangping Bryce Zhai, Hao Wen 0009, Qinyong Lin, Huazhou Chen, Ken Cai
IEEE J. Sel. Areas Commun.2
2023 LHDACT: Lightweight Hybrid Dual Attention CNN and Transformer Network for Remote Sensing Image Change Detection
abstract
With the significant advancements of Deep Learning (DL) in the field of remote sensing imagery, a plethora of Change Detection (CD) methods based on CNNs, attention mechanisms, and transformers have emerged. Presently, a substantial amount of research has gradually relinquished control over parameter quantities in pursuit of enhanced outcomes, resulting in the inflation of networks with numerous stacked modules. This paper is dedicated to integrating lightweight approaches into the CD task.We introduce a Lightweight Hybrid Dual-Attention CNN and Transformer network (LHDACT) based on Depthwise Over-Parameterized Convolution (DO-Conv). In comparison to traditional convolution, DO-Conv combines both traditional and depthwise convolutions, achieving commendable performance enhancement with minimal additional cost. Furthermore, we leverage DO-Conv to enhance the Multi-Scale Average Pooling module (MSAP), ensuring global context with low computational overhead.To better discern regions of interest within complex images, we enhance the Dual Attention Module (DAM) by sharing weights across spatial and channel dimensions, thereby bolstering feature region identification. Lastly, we employ a compact transformer module to capture feature differences, enabling precise change detection CD. Our approach is evaluated on the LEVIR-CD, WHU-CD, and GZ-CD datasets, yielding F1 scores of 91.23%, 87.51%, and 85.32%, respectively. These results demonstrate high performance on a cost-effective scale.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Route Planning Based on Parallel Optimization in the Air-Ground Integrated Network
abstract
Recent advancement in propulsion technologies to reduce the need for travel or increase the share of sustainable unmanned devices has accelerated the shift toward sustainable transport. To achieve the optimization of route planning in the air-ground integrated network (AGIN), we design an optimization strategy of accompanying graph navigation for unmanned devices, which aims to reduce the power consumption and$CO_{2}$gas emissions. The optimization of accompanying graph navigation is composed of three strategies, namely, the navigation based on the complete maps, the navigation based on the partitioned maps, and the navigation without maps. We propose a Two-tiered Grid (TG) index and Distributed AGIN Navigation (DAN) to navigate on partitioned maps. The top layer of the TG-index is composed of the border vertices of the global road network, which reflects the overall traffic conditions of the global road network and provides coarse-grained navigation routes. The bottom layer is a grid index composed of subgraphs, which reflects traffic conditions in local areas and provides fine-grained navigation routes. The navigation optimization is implemented in several segments, which can be run by multi-processors and realize rapid response to a large number of concurrent queries.
Ken Cai, Tianlun Dai, Qinyong Lin, Xinyang Song, Qian Zhou 0005, Jinzhan Wei, Huazhou Chen, Bohan Li 0001
IEEE Trans. Intell. Transp. Syst.4
2023 T-PORP: A Trusted Parallel Route Planning Model on Dynamic Road Networks
abstract
Route planning over dynamic road networks is an increasingly fundamental problem of modern transportation systems for human society, especially in the field of Intelligent Supply Chain (ISC). Due to the high degree of urbanization and the high number of vehicles, longer response time caused by massive concurrent queries, as well as more attacks caused by malicious vehicles, results in low efficiency of the transportation system and huge waste of computation resources. Thus, it is necessary to provide an efficient and safe transportation service for intelligent transportation planning. To achieve it, we utilize and improve the trust model to prevent the waste of computation resources. Meanwhile, we introduce a Trusted Parallel Optimization on Route Planning (T-PORP) based on Dual-level Grid (DLG) index to continuously handle the process of route planning in parallel. Considering the evolving traffic condition, we employ an LSTM (Long Short-Term Memory) neural network to periodically predict the weights of roads. Experimental results indicate that T-PORP is effective to sorts of trust model attacks and reduces the response time by an average of about 46.7% and saves the processing time by an average of about 27.6% compared with CANDS (Continuous Optimal Navigation via Distributed Stream Processing) algorithm.
Bohan Li 0001, Tianlun Dai, Weitong Chen 0001, Xinyang Song, Yalei Zang, Zhelong Huang, Qinyong Lin, Ken Cai
IEEE Trans. Intell. Transp. Syst.4
2022 Pavise: Integrating Fault Tolerance Support for Persistent Memory Applications
abstract
Persistent memory (PM) allows programmers to bypass the file system and efficiently manage persistent data directly. As a consequence, the application is now responsible for a non-trivial task---maintaining data crash consistency. In addition, it is highly desirable for today's production-grade storage systems to have fault tolerance to restore from data corruptions. Systems may provide fault tolerance through data redundancy. However, direct PM accesses bypass the system and make the data vulnerable to corruption. Without system-level support, it is the application's responsibility to maintain both crash consistency and fault tolerance, creating a demand for software tools to alleviate the burden from the application programmer.
Han Jie Qiu, Sihang Liu 0001, Xinyang Song, Samira Manabi Khan, Gennady Pekhimenko
PACT3
2022 Remote Sensing Image Change Detection Transformer Network Based on Dual-Feature Mixed Attention
abstract
Change detection (CD) of high-resolution remote sensing (RS) images is a basic task in RS image processing tasks. In recent years, CD tasks have made many attempts in pure convolutional networks, attention mechanism, and transformer, and have achieved good results. Based on the power of attention and transformers, we hope to find a method that can handle the details of the image better and has better generalization ability. In this article, we propose a dual-feature mixed attention-based transformer network (DMATNet). First, we adopt a dual-feature extraction method, using a simple convolutional neural network (CNN) to extract coarse features, and a CNN based on progressive sampling to extract fine features. Then, we fuse the fine and coarse features with dual-feature mixed attention (DFMA) module. It can not only extract more specific regions of interest, but also overcome the misjudgment caused by oversampling, and synchronize feature extraction and target information integration. Finally, we use transformer to optimize these extracted information and feedback into the original features in the encoder to help remodel the pixel space. We merged the DMAT network into a deep feature difference-based CD framework and conducted extensive experiments on four datasets, LEVIR-CD, DSIFN-CD, WHU-CD, and CLCD, respectively, with tested F1 and interconnection over union (IoU) results of 90.75%/84.13%, 71.23%/55.32%, 85.70%/74.98%, and 66.56%/59.87%. Experimental results show that our DMAT-based model performs significantly better than the existing state-of-the-art attention and transformer-based methods.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 A Trust Management-Based Route Planning Scheme in LBS Network
Xinyang Song, Bohan Li 0001, Tianlun Dai, Jiaying Tian
ADMA1
2020 Blockchain-Based Privacy Preserving Trust Management Model in VANET
Ruochen Liang, Bohan Li 0001, Xinyang Song
ADMA3