VLDB 2026 Research / reviewers in the wild / expert
Jinxin Guo
dblp:238/6905
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASC-SRN: Adaptive sparse sampling convolution-based infrared super-resolution network
Jinxin Guo, Weida Zhan, Depeng Zhu, Yu Chen 0079 |
Expert Syst. Appl. | 3 |
| 2026 | Towards SAR-to-optical image translation via perception correlative learning and global-local feature collaboration
Weida Zhan, Depeng Zhu, Jinxin Guo, Ziqiang Hao |
Pattern Recognit. | 6 |
| 2026 | Hierarchical Semantics Interaction for Compressed Video Action RecognitionabstractDirectly recognizing action based on compressed video shows significant advantages, such as low storage demands, efficient decoding, and fast inference speeds. Existing methods on compressed video achieve promising performance by separately modeling spatial and motion cues and directly fusing recognition results of I-frames and P-frames. However, these approaches overlook the following inherent attributes of compressed videos: 1) Temporal misalignment between I-frames and P-frames impairs the accuracy of action recognition. 2) Spatiotemporal sparsity of compressed video frames severely hinders the semantic modeling of complex actions. 3) Semantic discrepancy between I-frames and P-frames, which capture appearance and motion information respectively, leads to suboptimal performance when they are fused directly. To address these challenges, we propose a Hierarchical Semantics Interaction Network (HSINet) that ensures refined semantic modeling of compressed video through alignment, interaction, and calibration within a hierarchical fusion framework. Specifically, to resolve temporal misalignment, we propose an efficient cross-modal temporal alignment module that fully combines the spatiotemporal information of P-frames and the spatial information of I-frames, and includes both pre-alignment and fine alignment stages. To mitigate semantic degradation caused by sparsity, we propose a cross-modal semantics interaction module to provide the multi-scale semantics interaction between spatial and temporal representation learning and enhance representations’ spatiotemporal awareness. To calibrate the semantic imbalance between I-frames and P-frames, we propose a modal imbalance calibration module that optimizing directional differences via cosine similarity in hyperspherical space. Experiments on HMDB-51, UCF-101, and Kinetics-400 benchmarks, demonstrate the effectiveness of hierarchical semantic interaction for compressed video action recognition. Jinxin Guo, Yang Yang 0121, Huaiwen Zhang, Shengsheng Qian, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Toward Thermal Infrared Image Colorization via Large Kernel Convolution and Patch-Wise Graph Contrastive LearningabstractIn the context of intelligent transportation systems (ITS), accurate thermal infrared image colorization plays a vital role in improving visual interpretation of traffic scenes, facilitating downstream analysis of vehicles and pedestrians under complex environmental conditions. Existing methods predominantly focus on point-wise correspondences, overlooking the spatial topology among neighboring pixels, which often leads to unrealistic and structurally inconsistent results. Moreover, most current methods employ small receptive fields that struggle to capture long-range dependencies and high-level semantics in complex infrared scenes. To address the above issue, we propose a Large Kernel convolution and Patch-wise Graph Contrastive Learning based generative adversarial network, called LKPG-GAN. Firstly, we introduce graph contrastive learning strategy to ensure topological consistency of color features and avoid unnecessary content deformation during the colorization process. Secondly, we propose Feature Selection Extraction Module extracts image features at the larger scale, enhancing the network’s ability to extract advanced features from thermal infrared images. Additionally, we design Feature Refinement Block within the network, which enables more effective extraction of color and structural features by the color infrared image generator network. Finally, we incorporate Attention Gate Module within the decoder, allowing operations on feature maps with a large receptive field to prevent information loss during the sampling process. Experimental evaluations on the KAIST dataset and the FLIR dataset prove that our proposed LKPG-GAN exceeds existing state-of-the-art methods in colorization performance. Furthermore, by focusing on typical urban traffic scenes contained in these datasets, our method establishes a solid foundation for further application and research in ITS thermal image processing, providing an effective tool for enhanced scene visualization and interpretation. Source code will be available athttps://github.com/cyanymore/LKPG-GAN Yu Chen 0079, Weida Zhan, Depeng Zhu, Jinxin Guo, Ziqiang Hao, Deng Han |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Reference-based infrared image colorization via feature enhancement and context refinement
Weida Zhan, Yu Chen 0079, Depeng Zhu, Jinxin Guo, Ziqiang Hao, Deng Han, Jin Li 0063 |
Knowl. Based Syst. | 6 |
| 2025 | MSSFC-Net: Enhancing Building Interpretation With Multiscale Spatial-Spectral Feature CollaborationabstractBuilding interpretation from remote sensing imagery primarily involves two fundamental tasks: building extraction and change detection. However, most existing methods address these tasks independently, overlooking their inherent correlation and failing to exploit shared feature representations for mutual enhancement. Furthermore, the diverse spectral, spatial, and scale characteristics of buildings pose additional challenges in jointly modeling spatial-spectral multi-scale features and effectively balancing precision and recall. The limited synergy between spatial and spectral representations often results in reduced detection accuracy and incomplete change localization. To address these challenges, we propose a Multi-Scale Spatial-Spectral Feature Cooperative Dual-Task Network (MSSFC-Net) for joint building extraction and change detection in remote sensing images. The framework integrates both tasks within a unified architecture, leveraging their complementary nature to simultaneously extract building and change features. Specifically, a Dual-branch Multi-scale Feature Extraction module (DMFE) with Spatial-Spectral Feature Collaboration (SSFC) is designed to enhance multi-scale representation learning, effectively capturing shallow texture details and deep semantic information, thus improving building extraction performance. For temporal feature aggregation, we introduce a Multi-scale Differential Fusion Module (MDFM) that explicitly models the interaction between differential and dual-temporal features. This module refines the network’s capability to detect large-area changes and subtle structural variations in buildings. Extensive experiments conducted on three benchmark datasets demonstrate that MSSFC-Net achieves superior performance in both building extraction and change detection tasks, effectively improving detection accuracy while maintaining completeness. Dehua Huo, Weida Zhan, Jinxin Guo, Depeng Zhu, Yu Chen 0079, Yueyi Han, Deng Han, Jin Li 0063 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | BFCD-Net: Enhanced Dual-Branch Network Framework for Infrared Small Target DetectionabstractInfrared small target detection has recently benefftted greatly from deep neural models. However, existing frameworks generally ignore effective global information modeling, leading to poor detection when the target has high similarity to the background. We propose a dual-branch detection framework BFCD-Net, to address the above challenges through a clever combination of neural low-rank background modeling and multi-dimensional feature characterization. In neural low-rank background modeling. We adopt low-rank prior neural functions and neural regularization to effectively capture the global and local properties of the background. Specifically, the BFCD-Net framework contains a feature enhancement module, a distinguishing convolution kernel module, and a dynamic detector. The feature enhancement module improves network perception of targets by enlarging image resolution and enhancing the features of target. The distinguishing convolution kernel module enhances contrast between targets and backgrounds through adaptive feature segmentation and differential convolution operations. The dynamic detector optimizes feature representation and decision processes through adaptive weight learning. Additionally, we introduce an anti-wavelet transform fusion module to enhance multi-scale feature integration capabilities, and employs a three-dimensional total variation constraint to further improve the robustness of background modeling. Meanwhile, we propose a LGdicefunction. It improves the accuracy of the model in infrared small target detection tasks by explicitly focusing on the discrepancy between the predicted results and the real target. Experimental results demonstrate that BFCD-Net delivers superior detection performance across challenging datasets, especially on the NUDT-SIRST dataset, the IoU and Precision reach 84.55% and 94.99%. BFCD-Net provides an innovative solution for efficient and robust infrared small target detection. Weida Zhan, Jinxin Guo, Depeng Zhu, Yu Chen 0079, Deng Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | YOLO-BS: A Better Object Detection Model for Real-Time Driver Behavior Detection
Yang Xi, Jinxin Guo, Ming Ma 0006 |
ICIC (12) | 2 |
| 2024 | MFHOD: Multi-modal image fusion method based on the higher-order degradation model
Jinxin Guo, Weida Zhan, Yu Chen 0079, Jin Li 0063 |
Expert Syst. Appl. | 1 |
| 2024 | A feature refinement and adaptive generative adversarial network for thermal infrared image colorization
Yu Chen 0079, Weida Zhan, Depeng Zhu, Ziqiang Hao, Jin Li 0063, Jinxin Guo |
Neural Networks | 8 |
| 2023 | MTFD: Multi-Teacher Fusion Distillation for Compressed Video Action RecognitionabstractAs an important work in computer vision, some recent representative works such as Two-stream networks, 3D ConvNets, and Transformer-based networks have achieved outstanding performance. However, due to the high computational cost, the explosion of computation time and parameters, they cannot meet the needs of real-time applications. The current work utilizes the keyframes, residual and motion information retained by compressed video for computation, which greatly reduces the computational effort but still cannot satisfy real-time applications. Therefore, we propose a multi-teacher fusion distillation framework for compressed video action recognition (MTFD). Unlike the traditional method of transferring the knowledge of single or multiple teachers directly into the student model, we also perform knowledge transfer between teachers. MTFD achieves better knowledge distillation through mutual guidance and information fusion between teachers. Furthermore, we improve the network’s ability to extract motion information, which ultimately reduces the computational effort while maintaining high accuracy. Jinxin Guo, Shaojie Li 0001, Ming Ma 0006 |
ICASSP | 1 |
| 2023 | META: Motion Excitation With Temporal Attention for Compressed Video Action RecognitionabstractCompressed video action recognition has gained significant attention recently due to its ability to replace the raw video with I-frames and compressed motion clues, such as motion vectors and residuals. This results in substantial reductions in storage and computation costs. However, this task suffers from coarse and a lack of structures that can capture long-range spatiotemporal dependencies. To address these issues, this paper proposes a novel module called Motion Excitation with Temporal Attention (META) and utilizes network structures that can capture long-range dependencies. The META module stimulates motion information between I-frames and enhances the motion representation of the motion vectors. It first assigns different weights to feature-level frames, and then calculates the feature-level temporal differences from spatiotemporal features. Finally, it utilizes these differences to excite the motion-sensitive channels of the features. In addition, for compressed video action recognition, we have introduced a new network structure that combines CNN and Transformer. It can seamlessly integrate the merits of convolution and self-attention in a concise transformer format. The proposed method is evaluated on the challenging HMDB-51 and UCF-101 datasets. The extensive comparison results and ablation studies demonstrate the effectiveness and strength of the proposed method. Shaojie Li 0001, Jinxin Guo, Ming Ma 0006 |
ICPADS | 3 |
| 2023 | LAE-Net: Light and Efficient Network for Compressed Video Action Recognition
Jinxin Guo, Ming Ma 0006 |
MMM (2) | 1 |
| 2023 | IFF-Net: I-Frame Fusion Network for Compressed Video Action RecognitionabstractCompressed video action recognition has received significant attention due to its potential for reducing storage and computational costs. However, the current methods typically only capture a few RGBs and compressed motion cues (e.g., motion vectors and residuals), which are insufficient for modeling actions at their full temporal extent. To address this issue, we propose a Time Domain Fusion (TDF) Module that can extract both low-frequency and high-frequency components from the video and integrate them seamlessly, resulting in the effective integration of abundant motion information into a single frame. More importantly, by using the TDF module, we introduced a new network called I-Frame Fusion Network (IFF-Net). The IFF -Net interacts with the original network (I-frame, motion vector, and residual) in two ways: explicit and implicit. Explicit interaction involves extracting the new representation and the original compressed representation information separately and then performing a later fusion. In contrast, implicit interaction uses the distillation approach, with the IFF-Net acting as the teacher to guide the I-frame network to learn full temporal expressions. Our approach performs better than state-of-the-art methods on the UCF-101 and HMDB-51 datasets for compressed video action recognition. Shaojie Li 0001, Jinxin Guo, Ming Ma 0006 |
SMC | 2 |
| 2022 | Multi-Knowledge Attention Transfer Framework for Action Recognition
Jinxin Guo, Ming Ma 0006 |
ICANN (1) | 2 |
| 2020 | Aerosol Optical Depth Estimate Using Ground-Measured Spectral Skylight Ratio MethodabstractAerosol Optical Depth (AOD) is an important physical quantity of atmospheric turbidity and a key factor for the atmospheric correction of optical remote sensing image. Several approaches have been developed to retrieve AOD from remote sensing data, and at ground level the sunphotometer is usually used to measure AOD. However, the sunphotometer is expensive and not portable for field campaign, which make it difficult to synchronously obtain AOD data along with other relevant measurement. In order to overcome this problem, this study proposed a new way to estimate AOD using Ground-measured Spectral Skylight Ratio method (GSSR) on the basis that skylight ratio is highly determined by such parameter. In the GSSR method, a look-up table (LUT) containing spectral skylight ratio in 400-1000nm was first established from the MODTRAN code under various AOD levels, solar zenith angles (SZA), and atmospheric types. Based on the LUT and actual skylight ratio measured from a ground spectrometer, the AOD was then determined using a shortest distance between the measurement and the spectral values in the LUT. Finally, the AOD derived from the GSSR method was validated using the AOD data from the CE318 sunphotometer data, and it was found that AOD error was about 0.0185, which demonstrated that the GSSR method can be used to estimate AOD accurately, and can be regarded as an alternative method to estimate AOD in the field work if no sunphotometer is available. Jing Nie 0003, Huazhong Ren, Hui Zeng 0004, Jiaji Dong, Jinxin Guo, Yitong Zheng |
IGARSS | 5 |
| 2019 | A New Index for Sandy Land Detection Based On Thermal Infrared Emissivity DataabstractSpatial distribution and disappearance of sandy land is important for ecosystem management of desert regions and provides highly valuable information on desertification and climate change studies in arid environments. Based on the field measurement in the Gurbantonggut Desert, Xinjiang, China and the analysis of the spectral features of sandy land, a new sand differential emissivity index (SDEI) was proposed first for sandy land detection. Compared with the previous vegetation index, which can only distinguish green plants from bare land, SDEI can make a distinction well between sandy land and dry vegetation. For large regional mapping of sandy land, SDEI was applied on the ASTER Global Emissivity Dataset based on the Google Earth Engine platform. And then, four emissivity simulation schemes of different mixed pixels were conducted to determine the best threshold of sandy land mapping. The results show that when the threshold value is larger than 0.041, the sand distribution can be well extracted. Finally, the sandy land area of China extracted by SDEI is 160.67×104km2for year 2008, which is close to the data released by the China’s State Forestry Administration. These experimental results indicated that SDEI is applicable to identification of sandy land, and therefore satellite remotely-sensed thermal infrared observations have good potential in sandy land detection. Huazhong Ren, Yunzhu Tao, Yitong Zheng, Yuanheng Sun, Jing Nie 0003, Jinxin Guo, Rongyuan Liu, Wenjie Fan 0001 |
IGARSS | 7 |
| 2019 | Identify Urban Area From Remote Sensing Image Using Deep Learning MethodabstractUrban area is the main and important space of human activities with a large number of population. Compared with rural and other natural areas, the dense buildings and high-intensity land use are the most different features of urban areas. Therefore, the urban area has obvious texture in remote sensing images. Effective and accurate identification of urban area can play an important role in urban study, urban planning and other urban-related fields. In this paper, a new method based on urban and non-urban scene classification using Convolutional Neural Network (CNN) technique is developed to identify the boundary of urban areas and is applied in Beijing as an example. An acceptable result of the urban area identification was obtained, indicating a great potential of deep learning method in urban related studies. Jinxin Guo, Huazhong Ren, Yitong Zheng, Jing Nie 0003, Yuanheng Sun, Qiming Qin |
IGARSS | 1 |