EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Zou
dblp:133/0594
· DBLP profile ↗
26ranked-venue papers
8as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Computer networks · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Joint Port Selection and Cramér-Rao Bound Optimization for ISAC in Fluid Antenna SystemabstractIntegrated sensing and communication (ISAC) has been envisioned as a key enabler for next-generation wireless systems. Meanwhile, fluid antenna system (FAS) has emerged as a promising flexible antenna technology, offering enhanced multiple access capabilities and significant spatial diversity gains. This paper studies the problem of joint port selection and Cramér-Rao Bound (CRB) optimization in an ISAC-enabled FAS. Specifically, the objective is to minimize the CRB while ensuring compliance with communication quality of service (QoS) requirements, transmit power constraints, and port selection limitations. To address this problem, we propose an iterative optimization algorithm leveraging the majorization-minimization (MM) and block coordinate descent (BCD) methods. Numerical results validate the effectiveness of the proposed algorithm, demonstrating its ability to achieve near-optimal solutions with efficient search. Lvxin Xu, Jiaqi Zou, Songlin Sun, Jintao Wang 0001 |
ICC | 2 |
| 2025 | Hierarchical Bi-directional LiDAR-Camera Fusion Framework
Yefei Yang, Yanyun Tao, Jiaqi Zou |
ICIC (14) | 3 |
| 2025 | Hierarchical Distribution-Aware Network for Point Cloud CompletionabstractIn the field of 3D vision, 3D point cloud completion is a critical task in many practical applications. This study proposes a point cloud completion network designed to mitigate the impact of uneven point cloud distribution on completion tasks. By combining explicit neighborhood aggregation methods with implicit association techniques, the network achieves balanced feature representation across regions with varying distributions. Our approach comprises a Distribution-Geometry Feature Extractor (DGFE), a Seed Generator (SG), and a Point Generator (PG). DGFE leverages the proposed Dynamic Differential Distribution-Aware Module (D3AM) to process and enhance point cloud features layer by layer. SG generates seed point clouds and their corresponding features, while PG utilizes Cross-Resolution Upsampling Blocks (CRUB) to progressively generate denser point clouds by integrating point clouds and features across different resolutions. Experimental results demonstrate that our method achieves state-of-the-art performance on the PCN, ShapeNet-55/34, and KITTI benchmark datasets. Jiaqi Zou, Yanyun Tao, Yefei Yang |
IJCNN | 1 |
| 2025 | SaLIC: Saliency-Enhanced Learned Image Compression for Balanced QualityabstractPerception-optimized Learned Image Compression (LIC) methods have recently made significant progress. They surpass both traditional image compression algorithms and non-perceptually optimized LIC methods in image sharpness, detail representation, and subjective perception, even at similar or lower bitrates. However, LIC methods optimized for perception often generate false details and textures in reconstructions, leading to underperformance in objective metrics like PSNR and MS-SSIM, which limits their applicability. In this paper, we introduce a comprehensive loss metric based on saliency detection that aids in achieving exceptional perceptual quality while minimizing distortions. By applying this metric in training, we develop the SaLIC model, i.e., Saliency-Enhanced Learned Image Compression. User study results indicate that, at similar or lower bitrates, SaLIC exhibits better human perceptual quality compared to HiFiC and VVC; quantitative results show that the PSNR of SaLIC significantly outperforms HiFiC (by 1-2dB), and the MS-SSIM of SaLIC even surpasses VVC, achieving a balance between perception and distortion. Mingwei He, Jiaqi Zou, Songlin Sun, Jintao Wang 0001 |
ISCAS | 2 |
| 2025 | Low Latency Immersive Visual Communication with Scalable Gaussian Splatting CodingabstractImmersive visual communication has many important applications and Gaussian Splatting (GS) is a recent breakthrough that uses learnable geometry and color representation to capture 3D world with a very efficient parallelizable rendering pipeline. However, the compression of GS data still lacks efficiency and is quite complex in computation which prevents its adoption and deployment as a streaming solution in the real world. In this work, we develop a lightweight scalable GS coding scheme that exploits the correlation between adjacent quality layers and come up with a lightweight novel prediction and residual coding scheme that creates layered representation and is friendly to the MPEG DASH-like receiver-driven scalable streaming solutions. Simulation results demonstrate the efficiency of the proposed compression solution, as well as low latency/complexity in the decoding and rendering process. To the best of our knowledge, this is the first high-efficiency and low-complexity scalable GS coding solution that can be deployed with the existing MPEG DASH framework. Lingyu Shi, Jiaqi Zou, Songlin Sun, Geert Van der Auwera, Zhu Li 0001 |
MMSP | 2 |
| 2025 | Empirical Analysis of LLMDPP: Advancing Log Parsing in the LLM EraabstractIn the field of software engineering, the automated analysis of log data is crucial for operations and maintenance teams. This study introduces LLMDPP, a novel log parser that leverages Large Language Models (LLMs) and Determinantal Point Process (DPP) sampling techniques to enhance the efficiency and accuracy of online log parsing. LLMDPP transforms raw log messages into structured log templates through a few-shot learning, thus simplifying the processing and analysis of log data. The study explores the accuracy of LLMDPP in log-parsing tasks and compares the effectiveness of different encoding functions (TF-IDF and Flan-T5-small embedding layer) in DPP sampling. Experimental results show that LLMDPP outperforms traditional methods in both Global Accuracy and Parsing Accuracy, with the semantic information encoding function performing better when the number of samples is low. Furthermore, we simulate an online scenario to evaluate the parsing effectiveness of different sampling methods on unseen log datasets. The results indicate that the DPP sampling method has an advantage in maintaining sample diversity and fairness, which can improve parsing accuracy. Siqin Zhang, Haijing Nan, Xueyu Hou, Jiaqi Zou, Zicong Miao |
MobiSys | 5 |
| 2025 | Energy Efficiency Optimization for Rate-Splitting Multiple Access in ISAC SystemsabstractWe consider a rate-splitting multiple access (RSMA) assisted dual-functional integrated sensing and communications (ISAC) system, where the ISAC base station (BS) has the dual capability to simultaneously communicate with downlink users and to probe detection signals to a target. For this system, we focus on the problem of energy efficiency (EE) maximization and propose a new algorithmic framework that aims to optimize the beamforming matrices of RSMA such as to maximize the EE. Our framework is applicable to the optimization of both common and private streams’ beamforming matrices, and it accounts for a variety of constraints which includes power consumption constraint, communication rate constraints and a sensing quality constraint which is expressed with the aid of Cramér-Rao bound (CRB). Finally, the performance of our framework is compared to that of SDMA and NOMA based ISAC, and the superiority of RSMA-ISAC to SDMA-ISAC and NOMA-ISAC is revealed. George A. Ropokis, Jiaqi Zou, Constantinos B. Papadias, Songlin Sun |
PIMRC | 3 |
| 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic SegmentationabstractThe divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to solve such problem. Recent works show that self-training is a powerful approach to UDA. However, existing methods have difficulty in balancing the scalability and performance. In this paper, we propose a hard-aware instance adaptive self-training framework for UDA on the task of semantic segmentation. To effectively improve the quality and diversity of pseudo-labels, we develop a novel pseudo-label generation strategy with an instance adaptive selector. We further enrich the hard class pseudo-labels with inter-image information through a skillfully designed hard-aware pseudo-label augmentation. Besides, we propose the region-adaptive regularization to smooth the pseudo-label region and sharpen the non-pseudo-label region. For the non-pseudo-label region, consistency constraint is also constructed to introduce stronger supervision signals during model optimization. Our method is so concise and efficient that it is easy to be generalized to other UDA methods. Experiments on GTA5 $\rightarrow$→ Cityscapes, SYNTHIA $\rightarrow$→ Cityscapes, and Cityscapes $\rightarrow$→ Oxford RobotCar demonstrate the superior performance of our approach compared with the state-of-the-art methods. Chuang Zhu, Kebin Liu 0002, Wenqi Tang, Ke Mei, Jiaqi Zou, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Semantic Communication for VR Music Live Streaming With Rate SplittingabstractVirtual reality (VR) live streaming has established a remarkable transformation of music performances that facilitates a unique interaction between artists and their audiences within a virtual environment, offering an experience that significantly surpasses the conventional constraints of live music events. This article proposes a novel framework for enhancing VR music live streaming through the integration of semantic communication and rate splitting. The framework aims to improve user experience by efficiently transmitting music and speech components. It utilizes a semantic encoder to separately extract semantic information for music and speech, to capture the unique characteristics of music and speech. After having the extracted feature, we propose a rate-splitting-based algorithm in the transmission of music and speech to enhance user utility by designating music as a common message for all users and speech as a private message targeted to specific users based on their preferences. Simulation results demonstrate significant performance gain compared to the baseline methods. Jiaqi Zou, Lvxin Xu, Songlin Sun |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Joint Design for Communication Beamforming and Radar Waveform in MIMO Radar and MU-MIMO Communication Co-Existing SystemabstractThis paper considers a joint design for communication beamforming matrices and radar waveform in a scenario where a multiple-input multiple-output (MIMO) radar and a MIMO communication system coexist. We propose a joint optimizing method to maximize the sum rate by designing the communication beamforming matrices and the radar waveform. We give consideration to power constraints for both radar and communication system. In order to guarantee the radar performance, we also consider the similarity constraint of the radar waveform. This paper formulates two convex optimization problems and proposes an alternating iterative algorithm. Finally, the simulation results verify that the proposed algorithm can effectively raise the sum rate. Jiaqi Zou, Songlin Sun |
WCNC | 2 |
| 2024 | Integrated Sensing and Communications: Recent Advances and Ten Open ChallengesabstractIt is anticipated that integrated sensing and communications (ISAC) would be one of the key enablers of next-generation wireless networks (such as beyond 5G (B5G) and 6G) for supporting a variety of emerging applications. In this paper, we provide a comprehensive review of the recent advances in ISAC systems, with a particular focus on their foundations, physical-layer system design, networking aspects and ISAC applications. Furthermore, we discuss the corresponding open questions of the above that emerged in each issue. Hence, we commence with the information theory of sensing and communications (S&C), followed by the information-theoretic limits of ISAC systems by shedding light on the fundamental performance metrics. Next, we discuss their clock synchronization and phase offset problems, the associated Pareto-optimal signaling strategies, as well as the associated super-resolution physical-layer ISAC system design. Moreover, we envision that ISAC ushers in a paradigm shift for the future cellular networks relying on network sensing, transforming the classic cellular architecture, cross-layer resource management methods, and transmission protocols. In ISAC applications, we further highlight the security and privacy issues of wireless sensing. Finally, we close by studying the recent advances in a representative ISAC use case, namely the multi-object multi-task (MOMT) recognition problem using wireless signals. Shihang Lu, Fan Liu 0005, Yunxin Li, Kecheng Zhang, Hongjia Huang, Jiaqi Zou, Xinyu Li 0007, Yuxiang Dong, Fuwang Dong, Jia Zhu 0001, Yifeng Xiong, Weijie Yuan 0001, Yuanhao Cui, Lajos Hanzo |
IEEE Internet Things J. | 6 |
| 2024 | Energy-Efficient Beamforming Design for Integrated Sensing and Communications SystemsabstractIn this paper, we investigate the design of energy-efficient beamforming for an ISAC system, where the transmitted waveform is optimized for joint multi-user communication and target estimation simultaneously. We aim to maximize the system energy efficiency (EE), taking into account the constraints of a maximum transmit power budget, a minimum required signal-to-interference-plus-noise ratio (SINR) for communication, and a maximum tolerable Cramér-Rao bound (CRB) for target estimation. We first consider communication-centric EE maximization. To handle the non-convex fractional objective function, we propose an iterative quadratic-transform-Dinkelbach method, where Schur complement and semi-definite relaxation (SDR) techniques are leveraged to solve the subproblem in each iteration. For the scenarios where sensing is critical, we propose a novel performance metric for characterizing the sensing-centric EE and optimize the metric adopted in the scenario of sensing a point-like target and an extended target. To handle the nonconvexity, we employ the successive convex approximation (SCA) technique to develop an efficient algorithm for approximating the nonconvex problem as a sequence of convex ones. Furthermore, we adopt a Pareto optimization mechanism to articulate the tradeoff between the communication-centric EE and sensing-centric EE. We formulate the search of the Pareto boundary as a constrained optimization problem and propose a computationally efficient algorithm to handle it. Numerical results validate the effectiveness of our proposed algorithms compared with the baseline schemes and the obtained approximate Pareto boundary shows that there is a non-trivial tradeoff between communication-centric EE and sensing-centric EE, where the number of communication users and EE requirements have serious effects on the achievable tradeoff. Jiaqi Zou, Songlin Sun, Christos Masouros, Yuanhao Cui, Ya-Feng Liu, Derrick Wing Kwan Ng |
IEEE Trans. Commun. | 1 |
| 2024 | Pretrain a Remote Sensing Foundation Model by Promoting Intra-Instance SimilarityabstractSelf-supervised learning (SSL) has gained significant traction within the remote sensing community, with pretraining a foundation model on large-scale unlabeled datasets for the interpretation of remote sensing images (RSIs) emerging as a trending direction. This approach aims to supplant the conventional practice of loading ImageNet pretrained weights, offering a more versatile and potentially more effective solution for handling RSIs. Among SSL techniques, contrastive learning excels in extracting general representations in the field of remote sensing. However, its excessive focus on inter-instance discrimination hinders the effectiveness of pretraining due to the diverse and complex geographical information present in RSIs. Moreover, the typical two-variations-as-one-pair pattern may be suboptimal, particularly given the temporal information specific to RSIs. In this article, we propose a novel method called promoting intra-instance similarity (PIS) for short, which leverages the temporal information specific to RSIs and increases the intra-instance variations to expand the positive representation space. Additionally, by PIS within this space, our foundation models develop the ability to extract more general and instance-invariant features that prove beneficial for various downstream tasks. Experiments show that our PIS method achieves state-of-the-art (SOTA) performance on ten datasets across four downstream remote sensing tasks, demonstrating the generalizability and efficacy of the proposed method. Through our preliminary investigation into intra-instance characteristics, we believe there exists substantial potential in this aspect, holding considerable promise for further exploration. The codes are available on the website:https://github.com/ShawnAn-WHU/PIS.git. Xiao An, Wei He 0003, Jiaqi Zou, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | PSFormer: Pyramid Superpixel Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a core processing procedure in the remote sensing community, which has been recently well studied using vision transformers (ViTs). However, due to the high computational and memory complexities, existing transformer-based classification methods tend to restrict the spatial extent of the transformer to small cropped HSI patches instead of the whole HSI data, thus sacrificing the essential strength of transformers in long-range interaction modeling and overlooking the beneficial multiscale features in HSI data. Inspiringly, here we propose PSFormer, a novel pyramid superpixel transformer (PSFormer) method specifically for HSI classification, in order to make full use of the transformer to excavate multiscale local-global features in HSI data. Specifically, a progressive superpixel merging strategy is introduced to flexibly control the scale of feature maps. Furthermore, a unique transformer backbone design based on a spectral attention layer and a classification head with a gate mechanism are developed, to adaptively exploit valuable local-global information at different scales with low computational cost. Extensive experimental results on five widely used datasets demonstrate the superiority of PSFormer over other state-of-the-art networks. For the sake of reproducibility, the related code of the PSFormer method will be open-sourced at:https://github.com/immortal13. Jiaqi Zou, Wei He 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Sensing-Centric Energy-Efficient Waveform Design for Integrated Sensing and CommunicationsabstractIn this paper, we consider the energy-efficient waveform design for integrated sensing and communications systems, simultaneously performing multi-user communications and point-like/extended target sensing. We propose a performance metric to measure sensing-centric energy efficiency (EE) for the first time, namely sensing-centric EE. We formulate a problem to optimize sensing-centric EE with power budget, signal-to-interference-and-noise ratio (SINR) constraints for communication and a Cramér-Rao bound (CRB) constraint for sensing. For the point-like target case, we give the first-order approximations for the non-convex formulations and develop an effective iterative algorithm to handle the nonconvexity. For the extended target case, we show that the considered problem can be relaxed into semidefinite programming and the optimum can be reconstructed. Simulation results demonstrate significant performance gains on sensing-centric EE over the benchmarks. Jiaqi Zou, Songlin Sun, Christos Masouros, Yuanhao Cui |
GLOBECOM | 1 |
| 2023 | A Geographically Weighted Regression-Based Soil Moisture Product Using Cygnss GNSS-R DataabstractThe use of the Cyclone Global Navigation Satellite System (CYGNSS) for soil moisture (SM) estimation is of interest. However, the advantage of the variable resolution of CYGNSS was not fully utilized, leading to the loss of detailed information. Geographically Weighted Regression (GWR) permits the co-existence of diverse spatial relationships across different geographic regions, with the regression coefficient varying spatially rather than being globally constant, thus enabling coefficient adjustments within specific spatial boundaries. Advanced GWR-based SM estimation offers a significant improvement over other competing estimation models. This study demonstrated that the CYGNSS with high temporal and spatial resolution has the potential for high-resolution independent SM retrieval. Yan Jia 0004, Jiaqi Zou, Zhiyu Xiao, Qingyun Yan, Yinqing Zhen, Shuanggen Jin |
IGARSS | 2 |
| 2023 | A New Semantic Segmentation Diagram for Intelligent Transportation Based on Heterogeneous Knowledge BaseabstractSemantic segmentation is regarded as an important technology for future communication and sensing networks due to its promising ability to extract features of transmit data. It integrates the functionality of computer vision and can realize high-fidelity transmission in the channel with lower bandwidth. Knowledge Base (KB) is a key component for the semantic segmentation framework. In this paper, we consider the scenario of intelligent transportation and propose a heterogeneous KB to extract the to-be-transmitted features and information at multiple levels, i.e., the raw level, symbol level, feature level and image level. The compressed features are obtained by a deep-learning-based framework. The proposed KB is shared with the transmitter and the receiver, by a service-oriented multi-stream transmission algorithm to meet various requirements of services. Experiments results demonstrate significant performance gain in terms of quality of service and encoding efficiency. Jingyuan Tang, Jiaqi Zou, Songlin Sun |
WCNC | 2 |
| 2022 | An Empirical Study on Multi-Source Cross-Project Defect Prediction ModelsabstractMulti-source cross-project defect prediction (MSCPDP) refers to transferring defect knowledge from multiple source projects to the target project. MSCPDP has drawn increasing attention of academic and industry communities owing to its advantages compared with single-source cross-project defect prediction (SSCPDP) and some MSCPDP models have been proposed. However, to the best of our knowledge, there are no empirical studies to investigate the effect of different MSCPCP models on the performance of MSCPDP. To comprehensively investigate the performance of different MSCPDP models, we first conduct the literature research about MSCPDP studies, and then identify and compare 7 state-of-the-art MSCPDP models in terms of multiple performance measures including PD, PF, area under ROC curve (AUC), F1, precision, Matthews correlation coefficient (MCC), and Popt20% on 20 publicly available defect datasets. Furthermore, a robust multiple comparison method, i.e., the Scott-Knott effect-size difference (ESD) test, is used for statistical test. The experiment results show that 1) Burak’s Filter always performs best in terms of precision, AUC, MCC, Popt20% except for F1;2) MSCPDP models outperform the mean performance of SSCPDP models on most datasets; 3) the performance of MSCPDP models still needs to be further improved. We suggest software engineers use MSCPDP models but not SSCPDP models for CPDP and pay more attention to both the distribution difference of different datasets and the problems of sample similarity and weight when building MSCPDP models. Xuanying Liu, Zonghao Li, Jiaqi Zou, Haonan Tong |
APSEC | 3 |
| 2022 | Assessment of Signal Degradation Performance on Vegetations for GNSS-R SM RetrievalabstractGlobal Navigation Satellite System-Reflectometry (GNSS-R) is a remote sensing technique and can be regarded as a bistatic radar system. GNSS-R uses GNSS signals as signal sources and obtains the Earth's surface environmental parameters, such as soil moisture (SM), by receiving the L-band microwave signal reflected from the Earth's surface. However, the surface vegetation could be one of the main factors influencing the accuracy of GNSS-R land applications since the plants, including branches and leaves, attenuate the GNSS signal. Also, the evaluation of signal attenuations caused by plant canopy is quite difficult. In this paper, we present a sensitivity study of received GPS signals (L1 and L2 bands) to the vegetation leaf area index (LAI) over different types of plants. The relationship of GPS signal Signal-to-noise ratio (SNR) attenuations (above-canopy and below-canopy) versus LAIs is established through field experiments. The results show that the SNR received at the L2 band is with a larger standard deviation (SD) than at the L1 band for each satellite. The sensitivity of L1 and L2 bands signal to LAI is revealed, which shows a larger sensitivity and a relatively good Person correlation coefficient (R) for lower vegetation biomass. In addition, the sensitivity of the L2 band signal to LAI is lower than the L1 band signal, and with a lower R. This study is significant for improving the quantitative representation of error estimations in GNSS-R SM retrieval. Yan Jia 0004, Shuanggen Jin, Qingyun Yan, Jiaqi Zou |
IGARSS | 4 |
| 2022 | Multi-Stage Pseudo-Label Iteration Framework for Semi-Supervised Land-Cover MappingabstractLand-cover mapping is a pivotal pathway for Earth observation. Nevertheless, the lack of labeled data and the domain gap of different mapping regions are still challenges inhibiting the large-scale implementation of common mapping methods. In this article, a multi-stage pseudo-label iteration framework is presented for the semi-supervised land-cover mapping track of the 2022 Data Fusion Contest (DFC-SLM). The proposed framework combines the multi-stage training process with the pseudo-label technique to tackle the issues of large-scale land-cover mapping task when limited labeled samples are available. Firstly, the multi-stage training process promotes to sufficiently explore the discriminative and robust features of land covers with limited labeled data, where the proportion of samples in confusing classes gradually increases. Secondly, the pseudo-label technique generates extra supervision information from the unlabeled imagery. Overall, experimental results obtained from several cities in France achieved a mIoU of 52.96%, and won 2nd place on the final leaderboard of the 2022 DFC-SLM. Zhuohong Li, Jiaqi Zou, Fangxiao Lu, Hongyan Zhang 0001 |
IGARSS | 2 |
| 2022 | Energy Efficiency Optimization for Integrated Sensing and Communications SystemsabstractIn this paper, we consider an energy efficient waveform design in integrated sensing and communication (ISAC) systems. The transmitted waveform simultaneously serves multiple communication users and estimates the parameters of a moving target. In order to improve its energy efficiency (EE) while guaranteeing target estimation performance, we maximize the EE of the emitted dual-use waveform, under a Cramér-Rao bound (CRB) constraint. However, the considered optimization problem is a fractional function that is highly non-convex. Thus, we firstly adopt fractional programming based on Dinkelbach’ method and then, solve the sub-problem by leveraging semi-definite relaxation (SDR). Numerical results demonstrate superior performance than the benchmark and show the trade-off between EE and CRB. Jiaqi Zou, Yuanhao Cui, Songlin Sun |
WCNC | 1 |
| 2022 | EMS-GCN: An End-to-End Mixhop Superpixel-Based Graph Convolutional Network for Hyperspectral Image ClassificationabstractThe lack of labels is one of the major challenges in hyperspectral image (HSI) classification. Widely used Deep Learning (DL) models such as convolutional neural networks (CNNs) experience serious performance degradation when training samples are limited. In contrast, graph convolutional networks (GCNs) can simultaneously exploit the insufficient labeled data and massive unlabeled data of HSI in a semisupervised learning fashion. However, in order to reduce computational cost and mitigate noise, existing GCN-based classification methods usually perform superpixel segmentation as a preprocessing step and implement feature extraction as well as node classification on the predefined superpixel graph, where one superpixel might incorporate pixels with different labels. Moreover, the local spectral–spatial information within superpixels is generally ignored. To alleviate these two issues, we propose an end-to-end mixhop superpixel-based GCN (EMS-GCN) framework for HSI classification. Specifically, we first introduce the differentiable superpixel segmentation algorithm to map the pixel representations into a superpixel feature space, which allows refining the superpixel boundary with the training of the network. After that, a superpixel graph is constructed and fed into a novel mixhop superpixel-based GCN, where both the local information within superpixels and long-range information among superpixels are extracted, while the structure of the superpixel graph is updated at the same time. Finally, the enhanced superpixel representations are mapped back into a pixel feature space to conduct pixel-wise classification. Extensive experiments demonstrate the effectiveness of the proposed EMS-GCN method compared with other state-of-the-art methods. Hongyan Zhang 0001, Jiaqi Zou, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | LESSFormer: Local-Enhanced Spectral-Spatial Transformer for Hyperspectral Image ClassificationabstractCurrently, the convolutional neural networks (CNNs) have become the mainstream methods for hyperspectral image (HSI) classification, due to their powerful ability to extract local features. However, CNNs fail to effectively and efficiently capture the long-range contextual information and diagnostic spectral information of HSI. In contrast, the leading-edge vision transformers are capable of capturing long-range dependencies and processing sequential data such as spectral signatures. Nevertheless, pre-existing transformer-based classification methods generally generate inaccurate token embeddings from a single spectral or spatial dimension of raw HSIs and encounter difficulty modeling locality with insufficient training data. To mitigate these limitations, we propose a novel local-enhanced spectral-spatial transformer method (i.e., LESSFormer) specifically devised for HSI classification. Two effective and efficient modules are designed in LESSFormer, i.e., the HSI2Token module and the local-enhanced transformer encoder. The former is devised to transform HSI into the adaptive spectral-spatial tokens, and the latter is built to further enhance the representation ability of tokens by reinforcing the local information explicitly with a simple attention mask as well as retaining the long-range information in the meantime. Extensive experimental results on the new Xiong’an dataset and the widely used Pavia University and Houston University datasets have shown the superiority of LESSFormer over other state-of-the-art networks. Jiaqi Zou, Wei He 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Instance Adaptive Self-training for Unsupervised Domain Adaptation
Ke Mei, Chuang Zhu, Jiaqi Zou, Shanghang Zhang |
ECCV (26) | 3 |
| 2020 | Multi-Scale Video Inverse Tone Mapping with Deformable AlignmentabstractInverse tone mapping(iTM) is an operation to transform low-dynamic-range (LDR) content to high-dynamic-range (HDR) content, which is an effective technique to improve the visual experience. ITM has developed rapidly with deep learning algorithms in recent years. However, the great majority of deep-learning-based iTM methods are aimed at images and ignore the temporal correlations of consecutive frames in videos. In this paper, we propose a multi-scale video iTM network with deformable alignment, which increases time consistency in videos. We first align the input consecutive LDR frames at the feature level by deformable convolutions and then simultaneously use multi-frame information to generate the HDR frame. Additionally, we adopt a multi-scale iTM architecture with a pyramid pooling module, which enables our network to reconstruct details as well as global features. The proposed network achieves better performance compared to other iTM methods on quantitative metrics and gain a significant visual improvement. Jiaqi Zou, Ke Mei, Songlin Sun |
VCIP | 1 |
| 2020 | A fault diagnosis model of marine diesel engine cylinder based on modified genetic algorithm and multilayer perceptron
Liangsheng Hou, Jiaqi Zou, Changjiang Du, Jundong Zhang |
Soft Comput. | 2 |