Songchang Jin

dblp:143/6354 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-5959-0768ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Geometry-Aware Stereo Matching via Monocular Disparity Distribution Prior and Gradient Enhancement
abstract
Stereo matching recovers 3D scene information based on the correlation between corresponding pixels. Despite impressive progress, existing methods lack sufficient correlation priors in ill-posed regions such as occlusions, detailed and reflective regions. In this paper, we propose Geometry Aware Stereo Matching Network (GEAStereo) to enhance geometric structure perception and address this issue. We adaptively incorporate the Monocular Disparity Distribution Prior into the stereo cost volume, building Mono-Stereo Fusion Volume (MSFV), which effectively captures global geometric structures and rectifies the correlation information in ill-posed regions. Furthermore, we introduce rich detail information from gradient features and construct a Detail-Aware Volume (DAV) by aggregating the group-wise cost volume under the guidance of gradient spatial attention, thus enhancing the correlation modeling in detailed structures. Jointly, MSFV and DAV provide rich correlation priors for disparity iterative optimization. Experimental results show that our method achieves competitive results on the ETH3D and KITTI2015 benchmarks. Compared with the state-of-the-art methods, our method demonstrates stronger performance in zero-shot generalization.
Junze Zhang, Luoxi Jing, Yuanyuan Wang 0002, Guoli Yang, Songchang Jin, Chunping Qiu
AAAI6
2026 Hierarchical area graph for object navigation with adaptive entropy-driven exploration
Jing Xie 0021, Dian-xi Shi, Junze Zhang, Yuetian Wang, Songchang Jin
Adv. Eng. Informatics6
2026 D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu
Neural Networks6
2025 Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios
abstract
Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods.
Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang
AAAI3
2025 Multi-Agent Hierarchical Graph Attention Actor-Critic Reinforcement Learning
abstract
Multi-agent systems often face challenges such as elevated communication demands and intricate interactions. We propose an innovative hierarchical graph attention actor-critic reinforcement learning method to address the issues, which uses the hierarchical graph attention to capture the relationships of cooperation or competition among agents, and the agent enables a better understand of the dynamic environment. Specifically, we model the interaction among agents as a graph and encode the observations of the agents as a feature embedding vector with constant dimensionality to improve scalability. Through the "inter-agent" and "inter-group" attention layers, the embedding vector of each agent is updated into an information-condensed and contextualized state representation, which can adaptively extract the state-dependent relationship between agents, model the interaction at both the individual and group level, and thus learn more "advanced" strategies. Finally, we experiment on multiple multi-agent tasks to validate our proposed method’s effectiveness, stability, and scalability.
Tongyue Li, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Huanhuan Yang
ICASSP3
2025 Enhancing Visual Localization with Cross-Domain Image Generation
abstract
Visual localization aims to predict the absolute camera pose for a single query image. However, predominant methods focus on single-camera images and scenes with limited appearance variations, limiting their applicability to cross-domain scenes commonly encountered in real-world applications. Furthermore, the long-tail distribution of cross-domain datasets poses additional challenges for visual localization. In this work, we propose a novel cross-domain data generation method to enhance visual localization methods. To achieve this, we first construct a cross-domain 3DGS to accurately model photometric variations and mitigate the interference of dynamic objects in large-scale scenes. We introduce a text-guided image editing model to enhance data diversity for addressing the long-tail distribution problem and design an effective fine-tuning strategy for it. Then, we develop an anchor-based method to generate high-quality datasets for visual localization. Finally, we introduce positional attention to address data ambiguities in cross-camera images. Extensive experiments show that our method achieves state-of-the-art accuracy, outperforming existing cross-domain visual localization methods by an average of 59% across all domains. Project page: https://yzwang-sjtu.github.io/CDG-Loc.
Yuanze Wang, Yichao Yan, Shiming Song 0003, Songchang Jin, Yilan Huang, Xingdong Sheng, Dian-xi Shi
ICML4
2025 UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block
abstract
Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event and image data provides significant advantages, yet effective integration remains challenging. Existing CNN-based fusion methods struggle with occlusions and depth disparities due to limited receptive fields, while Transformer-based fusion methods often lack deep modality interaction. To address these issues, we propose UniCT Depth, an event-image fusion method that unifies CNNs and Transformers to model local and global features. We propose the Convolution-compensated ViT Dual SA (CcViT-DA) Block, designed for the encoder, which integrates Context Modeling Self-Attention (CMSA) to capture spatial dependencies and Modal Fusion Self-Attention (MFSA) for effective cross-modal fusion. Furthermore, we design the tailored Detail Compensation Convolution (DCC) Block to improve texture details and enhances edge representations. Extensive experiments show that UniCT Depth outperforms existing image, event, and fusion-based monocular depth estimation methods across key metrics.
Luoxi Jing, Dian-xi Shi, Zhe Liu 0029, Songchang Jin, Chunping Qiu, Ziteng Qiao, Jianqiang Xia
IJCAI4
2025 Remote sensing image encryption algorithm based on DNA convolution
Jingxi Tian, Songchang Jin, Dian-xi Shi, Shaowu Yang
J. Supercomput.4
2024 ICF-Loc: An Infrared-Based Coarse-to-Fine Approach for UAV Visual Geolocation under GPS-Denied Environments
abstract
Visual geolocation plays a crucial role when GPS is unavailable in the Unmanned Aerial Vehicles (UAVs). Many methods rely on visible light cameras, which may not perform well in low-light or foggy conditions. To address this issue, we propose an advanced UAV visual geolocation method called ICF-Loc, which utilizes infrared images in a coarse-to-fine approach. ICF-Loc consists of two stages: a retrieval-based coarse localization stage and a matching-based fine localization stage. The goal of the coarse localization stage is to identify the satellite image that is most similar to the UAV’s infrared image. To bridge the distribution gap between the visible and infrared domains, we propose a feature transfer module. In the fine localization stage, the UAV’s infrared image is matched with the satellite image obtained during coarse localization to estimate the UAV’s position and orientation accurately. We have designed a cross-modal image matching method based on the Fourier transform for precise estimation. The experimental results demonstrate the effectiveness of our proposed approach on both synthetic and real-world datasets.
Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li
ICME4
2024 Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006
IJCAI6
2024 SVR-AVT: Scale Variation Robust Active Visual Tracking
abstract
Active Visual Tracking (AVT) is a significant research area with extensive applications in fields such as drones and autonomous driving. AVT involves controlling camera motion based on visual observations to track target object(s). In dynamic environments, especially with the presence of distractors, AVT faces the challenge of scale variation. Existing methods struggle to effectively handle these scale changes. To address this problem, this paper proposes a novel Scale Variation Robust Active Visual Tracking method (SVR-AVT). We first introduce a multi-scale multi-stage curriculum learning approach. By progressively increasing the complexity of tracking tasks, the tracker adapts to target of various scales. Secondly, we design a scale attention network, which adaptively extracts important scale features through multiple convolutional branches with different receptive fields and a scale attention mechanism. Moreover, we employ maximum position entropy learning to encourage the target to explore the environment more extensively. Experimental results in 3D environments demonstrate that SVR-AVT significantly outperforms existing methods in handling distraction and scale variation, and exhibits strong generalization capability in unseen environments.
Zhang Biao, Songchang Jin, Qianying Ouyang, Huanhuan Yang, Yuxi Zheng, Chunlian Fu, Dian-xi Shi
IJCNN2
2024 Improved Communication and Collision-Avoidance in Dynamic Multi-Agent Path Finding
abstract
Multi-Agent Path Finding (MAPF) is a classic problem with a wide range of applications. To cope with more complex situations in reality, Dynamic MAPF (DMAPF) has received much attention. The existing DMAPF definition lacks completeness or considers too simple situations. In this paper, we comprehensively model DMAPF based on realistic scenarios. Consequently, dynamic scenarios bring many problems. The dynamics of agent tasks bring the problem of more difficult coordination and cooperation of the multi-agent system, and the dynamics of obstacles bring the problem of increased collisions. To address these problems, this paper proposes a fully decentralised multi-agent reinforcement learning method CO3, which uses COmmon knowledge in selective COmmunication and proposes obstacle COllision avoidance mechanism. Firstly, common knowledge for communication improves cooperation between agents, which improves system performance and reduces collisions between agents. Secondly, the obstacle collision avoidance mechanism consists of a collision avoidance helper module and a critical region. The collision avoidance helper module improves the agents’ alertness to nearby obstacles, and the critical region gives an early warning to the agents to beware of distant obstacles. The obstacle collision avoidance mechanism can effectively reduce collisions between agents and obstacles. Finally, experiments show that CO3 can solve the DMAPF problem quite well, and the number of collisions is significantly lower than other learning-based methods in a dynamic environment.
Jing Xie 0021, Yongjun Zhang 0006, Qianying Ouyang, Huanhuan Yang, Dian-xi Shi, Songchang Jin
IJCNN7
2024 JFDI: Joint Feature Differentiation and Interaction for domain adaptive object detection
Ziteng Qiao, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Chunping Qiu
Neural Networks3
2024 Sequence Matching for Image-Based UAV-to-Satellite Geolocalization
abstract
UAV-to-satellite geolocalization offers accurate drift-free navigation in the absence of external positioning signals. Increased deep-learning-based approaches have demonstrated their potential for high accuracy by framing the problem as a one-to-all retrieval task. However, in real-world scenario, the problem is not just a one-to-all retrieval task, which leads to a gap between research and applications. Based on this observation, We attempt to look closer to the problem instead of designing sophisticated network architectures or objective functions. In this study, we proposed a flexible and simple coarse-to-fine sequence-matching solution with targeted joint use of deep learning and classical machine learning approaches. Our goal is to improve geolocalization accuracy by matching UAV images with a few relevant reference image patches instead of all images. To this end, we first coarsely constructed a sequence of reference satellite image patches corresponding to the UAV trajectory, for which we proposed a deep feature- and manifold learning-based image-sorting method. Once the reference satellite patches are sorted and aligned with the UAV trajectory, the reference sequence is determined. Given a query UAV frame, the search area can be decreased from two to one dimensions. In particular, both deep-learning-based and classical image-matching algorithms can provide competitive accuracy when integrating sequence constraints. We demonstrate that classical manifold learning-based and image matching methods perform exceptionally well for UAV-to-satellite geolocalization when utilized jointly with suitable deep learning techniques. We validated the approach’s unique outperformance on two challenging and realistic UAV-to-satellite geolocalization datasets. Dataset, code and models are available for research purposes at https://seqmatch.geovisuallocalization.com/.
Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li, Zhe Liu 0029, Ziteng Qiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot
abstract
Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. Therefore, this paper presents a collision-free CPP approach (CFC) for the obstacle-constrained environment, which enhances time efficiency by constructing the variable-speed Dubins paths and ensures robot safety by building a risk potential surface for representing the possibility of collision. Furthermore, CFC models the CPP problem as an asymmetric traveling salesman problem (ATSP) and utilizes a graph pruning strategy to reduce the computational cost. Comparison tests with other Dubins coverage methods demonstrate that CFC provides shorter coverage times and better runtimes than the other Dubins coverage methods while preventing collision risk between the robot and obstacles. Physical experiments in a laboratory setting demonstrate the applicability of CFC to the physical robot.
Lin Li 0075, Dian-xi Shi, Songchang Jin, Yixuan Sun, Xing Zhou 0004, Shaowu Yang, Hengzhu Liu
ICRA3
2023 NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and Navigation
abstract
Visual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acquiring a large number of posed images and dense 3D labels in the real world is challenging and costly. In this paper, we present a novel visual localization method that achieves accurate localization while using only a few posed images compared to other localization methods. To achieve this, we first use a few posed images with coarse pseudo-3D labels provided by NeRF to train a coordinate regression network. Then a coarse pose is estimated from the regression network with PNP. Finally, we use the image-based visual servo (IBVS) with the scene prior provided by NeRF for pose optimization. Furthermore, our method can provide effective navigation prior, which enable navigation based on IBVS without using custom markers and depth sensor. Extensive experiments on 7-Scenes and 12-Scenes datasets demonstrate that our method outperforms state-of-the-art methods under the same setting, with only 5\% to 25\% training data. Furthermore, our framework can be naturally extended to the visual navigation task based on IBVS, and its effectiveness is verified in simulation experiments.
Yuanze Wang, Yichao Yan, Dian-xi Shi, Wenhan Zhu, Jianqiang Xia, Jeff Tan, Songchang Jin, Ke Gao 0012, Xiaokang Yang 0001
NeurIPS7
2021 Evaluating the Parallel Execution Schemes of Smart Contract Transactions in Different Blockchains: An Empirical Study
Chengzhi Li, Heng Wu 0001, Heran Gao, Songchang Jin, Tao Huang 0001, Wenbo Zhang 0006
ICA3PP (3)5
2020 Dual-template Siamese Network with Cross-Correlation Fusion for Tracking
abstract
Siamese network based trackers treat tracking as a maximum matching between the detection region and the template, in which the template is either fixed or updated. The template-fixed trackers degrade accuracy since the appearance variations of the object in the subsequent frames are lacked; while the template-updated trackers depress robustness once the tracking drift occurs. In this paper, we apply a dual-template tracking strategy in the Siamese region proposal network to compensate for the separate merits of fixed template and mutative template. As for the dual-template, the one template is fixed and the other template is constantly updated. Meanwhile, we integrate a learnable convolutional neural network to fuse two Cross-Correlation responses generated from two template tracking branches. Based on Siamese network, we try to utilize the complementary fusion of different template responses to promote tracking performances. Extensive experiments on five tracking benchmarks of OTB100, VOT2016, VOT2018, GOT-10k and UAV123 demonstrate that our approach achieves the state-of-the-art tracking performances.
Dian-xi Shi, Ying Kang, Zunlin Fan, Songchang Jin, Yongjun Zhang 0006
ICTAI5
2020 Balanced Multi-Region Coverage Path Planning for Unmanned Aerial Vehicles
abstract
Nowadays, Unmanned Aerial Vehicles (UAVs) are playing increasingly important roles in agriculture, rescuing and surveillance due to their small size, low cost and strong adaptability. Coverage Path Planning (CPP) is a fundamental problem for UAV applications, which means to find a path covering all the targets or regions of interest. Researches on CPP in a single region have been studied for decades, but rare to be devoted to covering multiple scattered regions of multiple UAVs. This paper proposes an attempt to solve this problem in a short time and take the task balance of multiple UAVs into account meanwhile. This work constructs a model for multiple UAVs to cover scattered regions firstly. In view of the high computational complexity of the precise solution, we keep innovating a heuristic measure based on the model to make the solving process feasible. To settle the problem of imbalanced time consumption caused by super regions, we improve the previously built model furthermore. A series of experiments validated that the presented approaches exhibit the effectiveness and balance of time consumption in diverse scenarios.
Xiaoxiao Yu, Songchang Jin, Dian-xi Shi, Lin Li 0075, Ying Kang, Junbo Zou
SMC2
2016 Collaborative Partitioning for Multiple Social Networks with Anchor Nodes
Fenglan Li, Anming Ji, Songchang Jin, Shuqiang Yang, Qiang Liu 0004
WAIM (1)3
2014 Synergistic partitioning in multiple large scale social networks
abstract
Social networks have been part of people's daily life and plenty of users have registered accounts in multiple social networks. Interconnections among multiple social networks add a multiplier effect to social applications when fully used. With the sharp expansion of network size, traditional standalone algorithms can no longer support computing on large scale networks while alternatively, distributed and parallel computing become a solution to utilize the data-intensive information hidden in multiple social networks. As such, synergistic partitioning, which takes the relationships among different networks into consideration and focuses on partitioning the same nodes of different networks into same partitions. With that, the partitions containing the same nodes can be assigned to the same server to improve the data locality and reduce communication overhead among servers, which are very important for distributed applications. To date, there have been limited studies on multiple large scale network partitioning due to three major challenges: 1) the need to consider relationships across multiple networks given the existence of intricate interactions, 2) the difficulty for standalone programs to utilize traditional partitioning methods, 3) the fact that to generate balanced partitions is NP-complete. In this paper, we propose a novel framework to partition multiple social networks synergistically. In particular, we apply a distributed multilevel k-way partitioning method to divide the first network into k partitions. Based on the given anchor nodes which exist in all the social networks and the partition results of the first network, using MapReduce, we then develop a modified distributed multilevel partitioning method to divide other networks. Extensive experiments on two real data sets demonstrate that our method can significantly outperform baseline independent-partitioning method in accuracy and scalability.
Songchang Jin, Jiawei Zhang 0001, Philip S. Yu, Shuqiang Yang, Aiping Li
IEEE BigData1