EDBT 2026 Demo / reviewers in the wild / expert
Yichen Zhao
dblp:157/7815
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An overview of domain-specific foundation model: key technologies, applications and challenges
Haolong Chen, Hanzhi Chen, Zijian Zhao 0002, Kaifeng Han, Guangxu Zhu, Yichen Zhao, Wei Xu 0001, Qingjiang Shi |
Sci. China Inf. Sci. | 6 |
| 2025 | Progressive language-aware encoding and decoding for referring expression comprehension
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001 |
Sci. China Inf. Sci. | 1 |
| 2025 | GCCO-ECP: A New Energy Saving Cluster IoT Protocol for Agricultural WSN Information Monitoring SystemabstractIn modern agriculture, efficient agricultural wireless sensor networks (AWSNs) are crucial for monitoring and management. Traditional algorithms struggle with the NP-hard cluster head selection problem in AWSNs. To address this, GCCO-ECP employs the Gaussian Chaos Opposition-based Learning Cheetah Optimization (GCCO) algorithm. This protocol takes into account various factors, including node energy, node degree, average distance, and delay, to enhance the clustering process. Additionally, it incorporates innovative strategies based on Gaussian mutation, chaos theory, opposition-based learning, and cheetah optimization principles to optimize the clustering scheme. Experimental results demonstrate that GCCO-ECP surpasses the LEACH-C, DMaOWOA, and MC-CRITIC-KM protocols, achieving remarkable improvements in network energy consumption, network lifespan, and transmission delay. Specifically, it optimizes network lifespan by 102.58%, throughput by 92.18%, and transmission delay by 11.73%, thereby showcasing its superiority in agricultural WSN information monitoring systems. The protocol’s design aims to enhance the diversity and convergence of population evolution, ensuring robust and efficient performance in real-world agricultural applications. Chuchu Rao, Dikun Wen, Qike Cao, Yichen Zhao, Yeshen Lan |
IEEE Internet Things J. | 4 |
| 2025 | Bilinear Parallel Fourier Transformer for Multimodal Remote Sensing ClassificationabstractVision Transformers (ViTs) have shown promise in multimodal fusion image classification, yet face performance challenges in complex remote sensing scenarios. Single fusion frameworks often fail to fully utilize multimodal diversity, and the uneven distribution of image categories complicates the accurate construction of spatial structures by Transformers. Additionally, traditional cross-entropy tends to favor majority classes, neglecting minority classes, resulting in suboptimal predictions and reduced overall accuracy (OA). To solve these challenges, we propose a novel deep neural network, a bilinear parallel Fourier Transformer (BPFT). We propose a novel dual-fusion feature interaction (DFFI) module that utilizes two distinct types of fused features for learning, namely the spatial-spectral fusion feature and the global fusion feature. Besides, we introduce a dual-feature interaction (DFI) module to improve the utilization of fused feature information. To enable the Transformer to better establish spatial structural relationships, we employ the Fourier transform in place of the self-attention mechanism. To address the focus on minority class labels, we propose an exponential label smoothing cross-entropy loss function. This loss function comprises two components: exponential cross-entropy and label smoothing. The exponential cross-entropy component applies a strong penalty to misclassified samples, thereby increasing attention on minority class labels. To validate the efficacy of our approach, extensive experiments are conducted across two multimodal remote sensing datasets: Augsburg and Berlin, encompassing hyperspectral imaging (HSI) data and synthetic aperture radar (SAR) data. The results of these experiments affirm the superior performance of our proposed BPFT model compared to existing state-of-the-art models in multimodal remote sensing image classification tasks. Yaxiong Chen, Qicong Wang, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | VGRSS: Datasets and Models for Visual Grounding in Remote Sensing Ship ImagesabstractThis paper introduces a task named Visual Grounding of Remote Sensing Ship Images (VGRSS). The goal of VGRSS is to locate ship objects in remote sensing images guided by natural language. Extensive research has been conducted on multimodal processing of remote sensing images and text to retrieve rich information from remote sensing images using natural language. However, due to the unique characteristics of remote sensing ship images, ship localization using natural language remains a challenge. Therefore, in this work, we construct datasets for the VGRSS task and explore deep learning models. Specifically, our contributions can be summarized as follows: First, we construct two remote sensing ship datasets for visual grounding. One is based on the optical remote sensing dataset, named RSSVG, while the other is based on the synthetic aperture radar (SAR) dataset, named SARVG. Second, we propose a Language-Guided Visual Feature Enhancement (LVFE) module. This module enhances visual features through language guidance before Visual-Linguistic Fusion. Third, we propose a Visual-Linguistic Fusion (VLF) module based on multimodal feature stacking. This module inputs the stacked language and visual features, and then performs feature fusion using a Transformer, enabling effective cross-modal interaction and integration. Fourth, we introduce a novel loss calculation method by incorporating Enhanced Intersection over Union (EIoU) into the loss function. Finally, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed RSSVG and SARVG datasets, then provide insightful analysis based on the results. This work offers valuable insights for developing better VGRSS models. Yaxiong Chen, Liwen Zhan, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Visual Grounding with Multi-modal Conditional AdaptationabstractVisual grounding is the task of locating objects specified by natural language expressions. Existing methods extend generic object detection frameworks to tackle this task. They typically extract visual and textual features separately using independent visual and textual encoders, then fuse these features in a multi-modal decoder for final prediction. However, visual grounding presents unique challenges. It often involves locating objects with different text descriptions within the same image. Existing methods struggle with this task because the independent visual encoder produces identical visual features for the same image, limiting detection performance. Some recently approaches propose various language-guided visual encoders to address this issue, but they mostly rely solely on textual information and require sophisticated designs. In this paper, we introduce Multi-modal Conditional Adaptation (MMCA), which enables the visual encoder to adaptively update weights, directing its focus towards text-relevant regions. Specifically, we first integrate information from different modalities to obtain multi-modal embeddings. Then we utilize a set of weighting coefficients, which generated from the multimodal embeddings, to reorganize the weight update matrices and apply them to the visual encoder of the visual grounding model. Extensive experiments on four widely used datasets demonstrate that MMCA achieves significant improvements and state-of-the-art results. Ablation experiments further demonstrate the lightweight and efficiency of our method. Our source code is available at: https://github.com/Mr-Bigworth/MMCA. Ruilin Yao, Shengwu Xiong 0001, Yichen Zhao |
ACM Multimedia | 3 |
| 2024 | A Conditional Diffusion Model Based WiFi Sensing Enhancement MethodabstractDriven by the rapid development of deep learning approaches, many novel WiFi sensing based applications have emerged, such as human activity recognition, pose estimation and indoor localization. However, due to the limited richness of collected WiFi data, the performance of WiFi sensing based models still lags behind conventional vision based models in terms of recognition accuracy and generalization. To break through the bottleneck of insufficient WiFi data, we propose a diffusion model based data augmentation scheme for human activity recognition task, in which the training dataset is composed of both real data and synthetic data. In particular, to reduce training overheads of the diffusion model, it is trained by taking activity classes as input conditions. Therefore, a single model is able to generate multiple types of WiFi data corresponding to activities, thereby avoiding the need to train separate models for each individual activity. Simulation results show that the generated WiFi data samples are visually indistinguishable from real ones, even when the model is trained on a small-scale dataset. Moreover, it also shows that adding an appropriate amount of synthetic data into training dataset can indeed improve the performance of WiFi sensing in most cases. Mingfeng Xu, Kaifeng Han, Yichen Zhao, Jiamo Jiang |
PIMRC | 4 |
| 2024 | Global-Group Attention Network With Focal Attention Loss for Aerial Scene ClassificationabstractAerial scene classification, aiming at assigning a specific semantic class to each aerial image, is a fundamental task in the remote sensing community. Aerial scene images have more diverse and complex geological features. While some statistics of images can be well fit using convolution, it limits such models to capturing the global context hidden in aerial scenes. Furthermore, to optimize the feature space, many methods add class information to the feature embedding space. However, they seldom combine model structure with class information to obtain more separable feature representations. In this article, we propose to address these limitations in a unified framework (i.e., CGFNet) from two aspects: focusing on the key information of input images and optimizing the feature space. Specifically, we propose a global-group attention module (GGAM) to adaptively learn and selectively focus on important information from input images. GGAM consists of two parallel branches: the adaptive global attention branch (AGAB) and the region-aware attention branch (RAAB). AGAB utilizes an adaptive pooling operation to better model the global context in aerial scenes. As a supplement to AGAB, RAAB combines grouping features with spatial attention to spatially enhance the semantic distribution of features (i.e., selectively focus on effective regions of features and ignore irrelevant semantic regions). In parallel, a focal attention loss (FA-Loss) is exploited to introduce class information into attention vector space, which can improve intraclass consistency and interclass separability. Experimental results on four publicly available and challenging datasets demonstrate the effectiveness of our method. The source code will be released at:https://github.com/zoecheno/CGFNet. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Co-Enhanced Global-Part Integration for Remote-Sensing Scene ClassificationabstractRemote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Evaluation on User Equipment Chip for Deep Learning based Channel Estimation in 5G Advanced SystemabstractNowadays, the integration of conventional 5G communication systems and artificial intelligence (AI) has become one of the important trends for the evolution of future advanced systems. With the employment of AI, many works have verified that significant performance gains can be achieved for varieties of tasks. However, these works mainly focus on the performance improvement and ignore the practical feasibility. In this paper, we investigate the performance of a convolutional neural network (CNN) based channel estimation scheme to run on two powerful mobile terminal chips with different quantization types. In particular, the accuracy of channel estimation and the computation capability of terminal chips for supporting model inference are evaluated. Simulation results show that the accuracy losses caused by both quantization types are small in the mobile scenario of speed at 120 km/h. However, in the case of speed at 350 km/h, the INT8 quantization type leads to a performance degradation while the $\Gamma$P16 quantization type can still maintain a satisfying performance. In addition, the results also show that further optimization for AI based computing is essential for guaranteeing a promising pratical deployment more reliably. Finally, the achievable demodulation performance gain is much smaller than the channel estimation gain, which should be explored fully in the future works. Yichen Zhao, Mingfeng Xu, Jiasong Mu, Jiamo Jiang, Hengjiang Wang |
IWCMC | 1 |
| 2023 | Aerial image recognition in discriminative bi-transformer
Yichen Zhao, Yaxiong Chen, Xiongbo Lu, Lei Zhou 0008, Shengwu Xiong 0001 |
Signal Process. | 1 |