EDBT 2026 Demo / reviewers in the wild / expert
Yuxuan Zheng
dblp:211/3406
· DBLP profile ↗
20ranked-venue papers
8as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cooperative Policy Agreement: Learning Diverse Policy for Offline MARLabstractOffline Multi-Agent Reinforcement Learning (MARL) aims to learn optimal joint policies from pre-collected datasets without further interaction with the environment. Despite the encouraging results achieved so far, we identify the policy mismatch problem that arises from employing diverse offline MARL datasets, a highly important ingredient for cooperative generalization yet largely overlooked by existing literature. Specifically, in the case that offline datasets exhibit various optimal joint policies, policy mismatch often occurs when individual actions from different optimal joint actions are combined in a way that results in a suboptimal joint action. In this paper, we introduce a novel Cooperative Policy Agreement (CPA) method, that not only mitigates the policy mismatch problem but also learns to generate diverse joint policies. CPA firstly introduces an autoregressive decision-making mechanism among agents during offline training. This mechanism enables agents to access the actions previously taken by other agents, thereby facilitating effective joint policy matching. Moreover, diverse joint policies can be directly obtained through sequential action sampling from the autoregressive model. Then we further incorporate a policy agreement mechanism to convert these autoregressive joint policies into decentralized policies with a non-autoregressive form, while still ensuring the diversity of the generated policies. This mechanism guarantees that the proposed CPA adheres to the Centralized Training with Decentralized Execution (CTDE) constraint. Experiments conducted on various benchmarks demonstrate that CPA yields superior performance to state-of-the-art competitors. Yihe Zhou, Yuxuan Zheng, Kai-Xuan Chen 0001, Tongya Zheng, Jie Song 0011, Mingli Song, Shunyu Liu 0001 |
AAAI | 2 |
| 2025 | Bi-Level Mean Field: Dynamic Grouping for Large-Scale MARLabstractLarge-scale Multi-Agent Reinforcement Learning (MARL) often suffers from the curse of dimensionality, as the exponential growth in agent interactions significantly increases computational complexity and impedes learning efficiency. To mitigate this, existing efforts that rely on Mean Field (MF) simplify the interaction landscape by approximating neighboring agents as a single mean agent, thus reducing overall complexity to pairwise interactions. However, these MF methods inevitably fail to account for individual differences, leading to aggregation noise caused by inaccurate iterative updates during MF learning. In this paper, we propose a Bi-level Mean Field (BMF) method to capture agent diversity with dynamic grouping in large-scale MARL, which can alleviate aggregation noise via bi-level interaction. Specifically, BMF introduces a dynamic group assignment module, which employs a Variational AutoEncoder (VAE) to learn the representations of agents, facilitating their dynamic grouping over time. Furthermore, we propose a bi-level interaction module to model both inter- and intra-group interactions for effective neighboring aggregation. Experiments across various tasks demonstrate that the proposed BMF yields results superior to the state-of-the-art methods. Our code is available at https://github.com/Chreer/BMF. Yuxuan Zheng, Yihe Zhou, Feiyang Xu, Mingli Song, Shunyu Liu 0001 |
ECAI | 1 |
| 2025 | Odyssey : Empowering Minecraft Agents with Open-World SkillsabstractRecent studies have delved into constructing generalist agents for open-world environments like Minecraft. Despite the encouraging results, existing efforts mainly focus on solving basic programmatic tasks, e.g., material collection and tool-crafting following the Minecraft tech-tree, treating the ObtainDiamond task as the ultimate goal. This limitation stems from the narrowly defined set of actions available to agents, requiring them to learn effective long-horizon strategies from scratch. Consequently, discovering diverse gameplay opportunities in the open world becomes challenging. In this work, we introduce Odyssey, a new framework that empowers Large Language Model (LLM)-based agents with open-world skills to explore the vast Minecraft world. Odyssey comprises three key parts: (1) An interactive agent with an open-world skill library that consists of 40 primitive skills and 183 compositional skills. (2) A fine-tuned LLaMA-3 model trained on a large question-answering dataset with 390k+ instruction entries derived from the Minecraft Wiki. (3) A new agent capability benchmark includes the long-term planning task, the dynamic-immediate planning task, and the autonomous exploration task. Extensive experiments demonstrate that the proposed Odyssey framework can effectively evaluate different capabilities of LLM-based agents. All datasets, model weights, and code are publicly available to motivate future research on more advanced autonomous agent solutions. Shunyu Liu 0001, Yaoru Li, Kongcheng Zhang, Zhenyu Cui, Wenkai Fang, Yuxuan Zheng, Tongya Zheng, Mingli Song |
IJCAI | 6 |
| 2025 | HiTuner: Hierarchical Semantic Fusion Model Fine-Tuning on Text-Attributed GraphsabstractText-Attributed Graphs (TAGs) are vital for modeling entity relationships across various domains. Graph Neural Networks have become cornerstone for processing graph structures, while the integration of text attributes remains a prominent research. The development of Large Language Models (LLMs) provides new opportunities for advancing textual encoding in TAGs. However, LLMs face challenges in specialized domains due to their limited task-specific knowledge, and fine-tuning them for specific tasks demands significant resources. To cope with the above challenges, we propose HiTuner, a novel framework that leverages fine-tuned Pre-trained Language Models (PLMs) with domain expertise as tuner to enhance the hierarchical LLM contextualized representations for modeling TAGs. Specifically, we first strategically select hierarchical hidden states of LLM to form a set of diverse and complementary descriptions as input for the sparse projection operator. Concurrently, a hybrid representation learning is developed to amalgamate the broad linguistic comprehension of LLMs with task-specific insights of the fine-tuned PLMs. Finally, HiTuner employs a confidence network to adaptively fuse the semantically-augmented representations. Empirical results across benchmark datasets spanning various domains validate the effectiveness of the proposed framework. Our codes are available at: https://github.com/ZihanFang11/HiTuner Zihan Fang 0002, Zhiling Cai, Yuxuan Zheng, Shide Du, Yanchao Tan, Shiping Wang |
IJCAI | 3 |
| 2025 | Worst-case robust optimization based on an adaptive incremental Kriging metamodel
Jie Han 0004, Yuxuan Zheng, Kai Wang 0024, Chunhua Yang 0001, Xin Yuan 0008 |
Expert Syst. Appl. | 2 |
| 2025 | ChangeDA: Depth-Augmented Multitask Network for Remote Sensing Change Detection via Differential AnalysisabstractIn the field of remote sensing change detection (RSCD), accurately identifying significant changes between bi-temporal images is essential for environmental monitoring, urban planning, and disaster assessment. In recent years, advancements in deep learning for computer vision (CV) have transformed RSCD, significantly enhancing its effectiveness. However, existing methods often overlook the importance of depth information, focusing primarily on 2-D information. This limits their ability to capture subtle changes and structural details in 3-D space. To address these limitations, we introduce ChangeDA—a depth-augmented multitask network designed to enhance the effectiveness of RSCD. ChangeDA introduces a depth encoder module to extract implicit depth information from optical images, enabling the utilization of 3-D structural information without reliance on external data sources. Through the depth infusion module (DIM), depth information is integrated into the dual-temporal feature maps, significantly enhancing the network’s ability to perceive changes in 3-D spatial structures. In addition, ChangeDA includes a differential feature extractor (DFE) tailored to pinpoint differential features between sequential images, and an adaptive all-feature fusion (AAFF) strategy that significantly improves recognition accuracy and generalization capability through cross-level feature integration. Performance evaluations on four prominent single-modal datasets—LEVIR-CD, S2Looking, WHU-CD, and SYSU-CD—yielded state-of-the-art (SOTA)${F}1$-scores of 92.27%, 66.42%, 94.12%, and 82.74%, respectively. Furthermore, ChangeDA also achieved outstanding results on the multimodal 3DCD dataset, with an${F}1$score of 63.52% in 2-D CD and an RMSE of 1.20 in the 3-D CD task. These results demonstrate ChangeDA’s robust adaptability across diverse targets and real-world scenarios. Jiangtao Meng, Xinying Xu, Pengyue Li, Gang Xie 0001, Jinchang Ren, Yuxuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | A Spatial-Spectral Difference Guided DDPM for Hyperspectral Image Change Detection
Yuxuan Zheng, Jiongce Zhao, Jiahui Qu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Non-Overlapping Sampling with Extreme Data Utilization in Hyperspectral ImageryabstractData-driven hyperspectral classification has gained significant attention across various applications, leading to the development of numerous novel neural networks for effective classification. However, existing spatial-spectral works often overlook the critical aspect of how training and testing samples are split, resulting in data leaks between these sets and suboptimal performance in real-world scenarios. To address this issue, we propose a non-overlapping sampling approach based on the extremal animals theory to achieve extreme data utilization, i.e., maximizing the number of testing samples. Specifically, it begins by determining the number of training samples n in each connectivity area according to the label map and pre-defined sampling ratio. Subsequently, it searches for all possible polyominos with the minimum perimeter p(n) and at least n pixels intersecting with the connectivity area. Finally, we remove unnecessary pixels according to the priority of their class in the connectivity area and degree in polyominos. Experimental results on Indian Pines validate the superior effectiveness of our proposed non-overlap sampling strategy. Yanzi Shi, Yuxuan Zheng, Yaping Yin |
IGARSS | 2 |
| 2024 | A Lightweight and Effective Multi-View Knowledge Distillation Framework for Text-Image RetrievalabstractLarge-scale dual-stream Vision-Language Pre-training (VLP) models provide an efficient solution for text-image retrieval tasks. Despite this, their performance often falls short of the most current single-stream models, primarily due to limited fine-grained text-image interactions. Recent trends indicate a union of these two types of networks. Some methods adopt a retrieve and rerank strategy, their performance improvements largely hinge on the single-stream encoder during inference. Other approaches utilize knowledge distillation to strengthen either the single-stream encoder or the dual-stream encoder, surpassing their previous capabilities. However, existing distillation techniques typically focus on a single knowledge type, neglecting the richer insights available in the teacher model. To bridge this gap, we introduce a Lightweight and Effective Multi-View Knowledge Distillation approach, named LEMKD, for text-image retrieval. This method effectively utilizes response-based, feature-based and relation-based knowledge, transferring the knowledge from the single-stream encoder to the dual-stream encoder. Our approach is executed on the widely used MS-COCO and Flickr30K datasets. Results demonstrate that LEMKD not only matches the exceptional performance of the most advanced single-stream models but also excels in dual-stream encoder performance amidst the recent integration of single-stream and dual-stream models. Yuxiang Song, Yuxuan Zheng, Shangqing Zhao, Xinlin Zhuang, Zhaoguang Long, Changzhi Sun, Aimin Zhou, Man Lan |
IJCNN | 2 |
| 2024 | Rank and Pairs of Rank and Dimension of Kernel of ZpZp²-Linear CodesabstractA code$C$is called$Z_{p}Z_{p^{2}}$-linear if it is the Gray image of a$Z_{p}Z_{p^{2}}$-additive code. For any prime number$p$larger than 3, the bounds of the rank of$Z_{p}Z_{p^{2}}$-linear codes are given. For each value of the rank and the pairs of rank and the dimension of the kernel of$Z_{p}Z_{p^{2}}$-linear codes, we give detailed construction of the corresponding codes. As an example, the rank and the dimension of the kernel of$Z_{5}Z_{25}$-linear codes are studied. Xiaoxiao Li 0002, Minjia Shi, Shukai Wang, Yuxuan Zheng |
IEEE Trans. Inf. Theory | 5 |
| 2022 | Hyperspectral Pansharpening With Adaptive Feature Modulation-Based Detail Injection NetworkabstractRecently, deep learning-based methodologies have attained unprecedented performance in hyperspectral (HS) pansharpening, which aims to improve the spatial quality of HS images (HSIs) by making use of details extracted from the high-resolution panchromatic (HR-PAN) image. However, it remains challenging to incorporate the details into the pansharpened image effectively, while alleviating the spectral distortion simultaneously. To tackle this problem, in this article, we propose an adaptive feature modulation-based detail injection network (AFM-DIN) for HS pansharpening, which mainly consists of four phases: high-frequency details generation of the HR-PAN image, multiscale feature extraction of the upsampled HSI, AFM-based detail injection and reconstruction of the HR-HSI. First, a novel octave convolution unit is employed to decompose the HR-PAN image into high and low frequencies, and then merge the high-frequency features together to generate the comprehensive PAN-details. Second, the spatial and spectral separable 3D convolution units with multiple kernel sizes are designed to extract multiscale features of the upsampled HSI in a computationally efficient manner. Subsequently, by taking the critical PAN-details as prior, the proposed AFM module is able to not only incorporate the detail information effectively, but also adjust the injected details adaptively to ensure the spectral fidelity. Finally, the anticipated HR-HSI is obtained through adding the upsampled HSI to the predicted HSI-details reconstructed from informative modulated features. Extensive comparison experiments with several state-of-the-arts conducted on simulated and real HS data sets demonstrate that our proposed AFM-DIN can achieve superior pansharpening accuracy in both spatial and spectral aspects. Yunsong Li 0001, Yuxuan Zheng, Jiaojiao Li 0001, Rui Song 0003, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Edge-Conditioned Feature Transform Network for Hyperspectral and Multispectral Image FusionabstractDespite recent advances achieved by deep learning techniques in the fusion of low-spatial-resolution hyperspectral image (LR-HSI) and high-spatial-resolution multispectral image (HR-MSI), it remains a challenge to reconstruct the high-spatial-resolution HSI (HR-HSI) with more accurate spatial details and less spectral distortions, since the low-level structure information such as sharp edges tends to be weakened or lost as the network depth grows. To tackle this issue, we creatively propose an edge-conditioned feature transform network (EC-FTN) in this article, which is mainly composed of three parts, namely, feature extraction network (FEN), feature fusion and transformation network (FFTN), and image reconstruction network (IRN). First, two computationally efficient FENs with 3-D convolutions and reshaping layers are employed to extract the joint spectral-spatial features of input images. Then, the FFTN conditioned on the edge map prior can fuse and transform the features adaptively, in which a fusion node and several cascaded feature modulation modules (FMMs) equipped with feature-wise modulation layers are constructed. Specifically, the edge map is generated via transfer learning, i.e., by applying the Sobel operator to feature maps of the red-green-blue (RGB) version of HR-MSI resulting from the pretrained VGG16 model without extra training. Finally, the desired HR-HSI is recovered from the transformed features through IRN. Furthermore, we elaborately design a weighted combinatorial loss function consisting of mean absolute error, image gradient difference, and spectral angle terms to guide the training. Experiments on both ground-based and remotely sensed datasets demonstrate that our EC-FTN outperforms state-of-the-art methods in visual and quantitive evaluations, as well as in fine details reconstruction. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Jie Guo 0009, Xianyun Wu, Yanzi Shi, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Pansharpening of Hyperspectral Images with Detail Guided Feature ModulationabstractPansharpening of hyperspectral image (HSI), which makes use of the detail information contained in the high-resolution panchromatic (HR-PAN) image to sharpen the low-resolution HSI (LR-HSI), is an essential technology to enhance the spatial resolution of HSI. In this paper, we propose a detail guided feature modulation residual network (DGFM-Net) to address the HS pansharpening problem, which is able to effectively integrate details extracted from the PAN image into the pansharpened result. Specifically, we elaborately design a novel feature modulation (FM) module with the guidance of PAN detail information to modulate HSI features flexibly and incorporate PAN details adaptively. The modulated features are then fed to the residual reconstruction (RR) block to recover the difference between the upsampled HSI and the HR-HSI by efficient residual learning. Finally, the upsampled HSI is combined with the estimated residual HSI to produce the desired HR-HSI. Experiments on the Pavia Center data set confirm that the proposed DGFM-Net outperforms several state-of-the-art HS pansharpening methods. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Kailang Cao |
IGARSS | 1 |
| 2021 | Hyperspectral Target Detection With RoI Feature Transformation and Multiscale Spectral AttentionabstractTarget detection plays a core issue in hyperspectral remote sensing, but faces serious challenges of how to deal with the spatial and spectral redundancies and spectral variations. In this article, a novel network block is developed, called RFT-MSA block (abbreviated as RM), which includes the region-of-interest (RoI) feature transformation (RFT) and the multiscale-spectral-attention (MSA) module as to reduce the spatial and spectral redundancies simultaneously and provide strong discrimination. Furthermore, a deep spatial-spectral network (DSSN) is presented by stacking several RM and deconvolutional (DC) blocks for hyperspectral target detection in an unsupervised manner, and a feature loss term is investigated to simultaneously restrict the target to be sparse and minimize the energy of the background. The proposed algorithm mainly consists of three steps. First, an RoI map is detected using a classical detector (no statistic assumption is needed) with an edge-preserving filter. Then, the hyperspectral image (HSI) and the corresponding RoI map are considered as inputs to the DSSN for extracting the spatial and spectral feature of interest (SSFI). Finally, we apply the nearest neighbors (NNs) to the SSFI for detection-map refinement. The experimental results on one synthetic and three real HSIs demonstrate that the proposed algorithm outperforms other benchmark approaches in detection performance and robustness. In addition, further analysis also demonstrates the effectiveness of the proposed RM block. Yanzi Shi, Jiaojiao Li 0001, Yuxuan Zheng, Bobo Xi, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Deep Residual Spatial Attention Network for Hyperspectral PansharpeningabstractIn this paper, we propose a deep residual spatial attention network (DRSAN) for hyperspectral (HS) pansharpening. Different from the existing methods, our newly proposed method not only considers the spatial information of both the panchromatic (PAN) and the HS image simultaneously, but also adaptively learns more informative features of spatial locations for details enhancement, which mainly includes four steps. Firstly, the spatial details of the enhanced PAN image are obtained through the structure tensor. Then we extract the spatial information of the upsampled HSI by using the guided filter. The integrated spatial information of both PAN and HS images is subsequently fed into the DRSAN to map the residual HSI between the upsampled HSI and the reference HSI, where several residual spatial attention blocks (RSABs) are cascaded to exploit more useful details information. Finally, the fused HSI is generated by the summation of the upsampled HSI and the reconstructed residual HSI. Extensive visual and quantitative assessments validate the superiority of our proposed DRSAN over the state-of-the-art HS pansharpening methods. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Yanzi Shi, Jiahui Qu |
IGARSS | 1 |
| 2020 | Deep Residual Learning for Boosting the Accuracy of Hyperspectral PansharpeningabstractRecently, deep learning (DL) has gained impressive achievements in the field of remote sensing image fusion. However, most of the previous DL-based fusion methods are originally designed for multispectral pansharpening, which cannot be readily employed to hyperspectral pansharpening due to the much wider spectral range and lower spatial resolution of a hyperspectral image (HSI). In this letter, a novel framework based on deep residual learning is proposed for hyperspectral pansharpening. The proposed framework consists mainly of two parts. First, the initialized HSI with the enhanced spatial resolution is generated through contrast limited adaptive histogram equalization (CLAHE) and guided filter. Then, a deep residual convolutional neural network (DRCNN) is introduced to map the residuals between the initialized HSI and the reference HSI for further boosting the fusion accuracy. Experimental results demonstrate that the proposed framework can achieve superior performance compared with the existing state-of-the-art pansharpening methods, especially in terms of edge details enhancement. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Kailang Cao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Hyperspectral Pansharpening Using Deep Prior and Dual Attention Residual NetworkabstractConvolutional neural networks (CNNs) have recently achieved impressive improvements on hyperspectral (HS) pansharpening. However, most of the CNN-based HS pansharpening approaches would have to first upsample the low-resolution hyperspectral image (LR-HSI) using bicubic interpolation or data-driven training strategy, which inevitably lose some details or greatly rely on the learning process. In addition, most previous methods regard the pansharpening as a black-box problem and treat diverse features equally, thus hindering the discriminative ability of CNNs. To conquer these issues, a novel HS pansharpening method using deep hyperspectral prior (DHP) and dual-attention residual network (DARN) is proposed in this article. Specifically, we first upsample the LR-HSI to the scale of the panchromatic (PAN) image through the DHP algorithm, which can better preserve spatial and spectral information without learning from large data sets. The upsampled result is then concatenated with the PAN image to form the input of the DARN, where several channel-spatial attention residual blocks (CSA ResBlocks) are stacked to map the residual HSI between the reference HSI and the upsampled HSI. In each CSA ResBlock, two complementary attention modules, i.e., channel attention and spatial attention modules, are designed to adaptively learn more informative features of spectral channels and spatial locations simultaneously, which can effectively boost the fusion accuracy. Finally, the fused HSI is obtained by the summation of the upsampled HSI and the reconstructed residual HSI. The experimental results of both simulated and real HS data sets demonstrate that the performance of our DHP-DARN method is superior over the state-of-the-art HS pansharpening approaches. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Jie Guo 0009, Xianyun Wu, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Hyperspectral Pansharpening Based on Guided Filter and Deep Residual LearningabstractRecently, deep learning technology has gained impressive effectiveness in the field of hyperspectral pansharpening. However, the existing methods with relatively shallow architectures ignores the deep features of hyperspectral image (HSI) and panchromatic (PAN) image, which leads to a limitation of the fusion performance. To address this issue, a novel hyperspectral pansharpening framework based on guided filter and deep residual learning is proposed in this paper. The proposed framework mainly consists of two parts: generating the initial HSI through enhancing spatial information while preserving the original spectral information, and mapping the residuals between the initialized HSI and the reference HSI for further improvement of the fusion accuracy. Experimental results demonstrate that the proposed framework can achieve superior fusion accuracy compared with other state-of-the-art hyperspectral pansharpening methods while providing better edge information. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001 |
IGARSS | 1 |
| 2018 | Hyperspectral pansharpening via improved PCA approach and optimal weighted fusion strategy
Yunsong Li 0001, Jiahui Qu, Wenqian Dong, Yuxuan Zheng |
Neurocomputing | 4 |
| 2018 | Efficient coarse-to-fine spectral rectification for hyperspectral image
Weiying Xie, Yunsong Li 0001, Weiping Zhou, Yuxuan Zheng |
Neurocomputing | 4 |