VLDB 2026 Research / reviewers in the wild / expert
Shujie Chen 0001
dblp:05/9884-1 · also Shu-Jie Chen 0001
· DBLP profile ↗
14ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-9502-5846ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationabstractWe propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a homography estimation network and a modality transfer network. To realize stable training, we introduce an effective split optimization strategy to train each network separately within its respective sub-problem. We also formulate an extra homography feature space supervision to enhance feature consistency, further boosting the estimation accuracy. Moreover, we employ a simple yet effective distillation training technique to reduce model parameters and improve cross-domain generalization ability while maintaining comparable performance. The training stability of SSHNet enables its cooperation with various homography estimation architectures. Experiments reveal that the SSHNet using IHN as homography estimation network, namely SSHNet-IHN, outperforms previous unsupervised approaches by a significant margin. Even compared to supervised approaches MHN and LocalTrans, SSHNet-IHN achieves 47.4% and 85.8% mean average corner errors (MACEs) reduction on the challenging OPT-SAR dataset. Source code is available at https://github.com/Junchen-Yu/SSHNet. Junchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang 0002, Zhu Yu 0001, Shujie Chen 0001, Bailin Yang |
CVPR | 6 |
| 2025 | EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image RegistrationabstractPrevious deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due to their inherent limitations. To address this, we propose an Exponential-Decay Free-Form Deformation Network (EDFFDNet), which employs free-form deformation with an exponential-decay basis function. This design achieves higher efficiency and performs well in scenes with depth disparities, benefiting from its inherent locality. We also introduce an Adaptive Sparse Motion Aggregator (ASMA), which replaces the MLP motion aggregator used in previous methods. By transforming dense interactions into sparse ones, ASMA reduces parameters and improves accuracy. Additionally, we propose a progressive correlation refinement strategy that leverages global-local correlation patterns for coarse-to-fine motion estimation, further enhancing efficiency and accuracy. Experiments demonstrate that EDFFDNet reduces parameters, memory, and total runtime by 70.5%, 32.6%, and 33.7%, respectively, while achieving a 0.5 dB PSNR gain over the state-of-the-art method. With an additional local refinement stage,EDFFDNet-2 further improves PSNR by 1.06 dB while maintaining lower computational costs. Our method also demonstrates strong generalization ability across datasets, outperforming previous deep learning methods. Haokai Zhu, Bo Qu, Si-Yuan Cao, Runmin Zhang, Shujie Chen 0001, Bailin Yang |
ICCV | 5 |
| 2025 | Hybrid Reinforcement Learning in parameterized action space via fluctuates constraint
Chengcheng Yan, Shujie Chen 0001, Zheng Peng 0004 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Representation alignment contrastive regularisation for multi-object trackingabstractAbstract Achieving high‐performance in multi‐object tracking algorithms heavily relies on modelling spatial‐temporal relationships during the data association stage. Mainstream approaches encompass rule‐based and deep learning‐based methods for spatial‐temporal relationship modelling. While the former relies on physical motion laws, offering wider applicability but yielding suboptimal results for complex object movements, the latter, though achieving high‐performance, lacks interpretability and involves complex module designs. This work aims to simplify deep learning‐based spatial‐temporal relationship models and introduce interpretability into features for data association. Specifically, a lightweight single‐layer transformer encoder is utilised to model spatial‐temporal relationships. To make features more interpretative, two contrastive regularisation losses based on representation alignment are proposed, derived from spatial‐temporal consistency rules. By applying weighted summation to affinity matrices, the aligned features can seamlessly integrate into the data association stage of the original tracking workflow. Experimental results showcase that our model enhances the majority of existing tracking networks' performance without excessive complexity, with minimal increase in training overhead and nearly negligible computational and storage costs. Shujie Chen 0001, Zhonglin Liu, Jianfeng Dong, Xun Wang 0007 |
IET Comput. Vis. | 1 |
| 2025 | TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and DetailsabstractThermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches. Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 6 |
| 2024 | Non-autoregressive transformer with fine-grained optimization for user-specified indoor layout
Chao Song 0001, Shujie Chen 0001, Zhaoyi Jiang, Bailin Yang |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningabstractThis paper targets unsupervised skeleton-based action representation learning and proposes a new Hierarchical Contrast (HiCo) framework. Different from the existing contrastive-based solutions that typically represent an input skeleton sequence into instance-level features and perform contrast holistically, our proposed HiCo represents the input into multiple-level features and performs contrast in a hierarchical manner. Specifically, given a human skeleton sequence, we represent it into multiple feature vectors of different granularities from both temporal and spatial domains via sequence-to-sequence (S2S) encoders and unified downsampling modules. Besides, the hierarchical contrast is conducted in terms of four levels: instance level, domain level, clip level, and part level. Moreover, HiCo is orthogonal to the S2S encoder, which allows us to flexibly embrace state-of-the-art S2S encoders. Extensive experiments on four datasets, i.e., NTU-60, NTU-120, PKU-I and PKU-II, show that HiCo achieves a new state-of-the-art for unsupervised skeleton-based action representation learning in two downstream tasks including action recognition and retrieval, and its learned action representation is of good transferability. Besides, we also show that our framework is effective for semi-supervised skeleton-based action recognition. Our code is available at https://github.com/HuiGuanLab/HiCo. Jianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen 0001, Xun Wang 0007 |
AAAI | 4 |
| 2022 | Partially Relevant Video RetrievalabstractCurrent methods for text-to-video retrieval (T2VR) are trained and tested on video-captioning oriented datasets such as MSVD, MSR-VTT and VATEX. A key property of these datasets is that videos are assumed to be temporally pre-trimmed with short duration, whilst the provided captions well describe the gist of the video content. Consequently, for a given paired video and caption, the video is supposed to be fully relevant to the caption. In reality, however, as queries are not known a priori, pre-trimmed video clips may not contain sufficient content to fully meet the query. This suggests a gap between the literature and the real world. To fill the gap, we propose in this paper a novel T2VR subtask termed Partially Relevant Video Retrieval (PRVR). An untrimmed video is considered to be partially relevant w.r.t. a given textual query if it contains a moment relevant to the query. PRVR aims to retrieve such partially relevant videos from a large collection of untrimmed videos. PRVR differs from single video moment retrieval and video corpus moment retrieval, as the latter two are to retrieve moments rather than untrimmed videos. We formulate PRVR as a multiple instance learning (MIL) problem, where a video is simultaneously viewed as a bag of video clips and a bag of video frames. Clips and frames represent video content at different time scales. We propose a Multi-Scale Similarity Learning (MS-SL) network that jointly learns clip-scale and frame-scale similarities for PRVR. Extensive experiments on three datasets (TVR, ActivityNet Captions, and Charades-STA) demonstrate the viability of the proposed method. We also show that our method can be used for improving video corpus moment retrieval. Jianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 0001, Shujie Chen 0001, Xirong Li 0001, Xun Wang 0007 |
ACM Multimedia | 5 |
| 2021 | Estimating Generalized Gaussian Blur Kernels for Out-of-Focus Image DeblurringabstractOut-of-focus blur is a common image degradation phenomenon that occurs in case of lens defocusing. The out-of-focus blur kernel is usually modeled as a Gaussian function or a uniform disk in previous work. In this paper, we propose that it can be more accurately depicted using the generalized Gaussian (GG) function. This is motivated by the theoretical analysis of the out-of-focus blur and the practical observation of real blur kernels. We show that as the out-of-focus blur kernels are of specific shapes, the GG function can be further simplified to a single-parameter model. We estimate the parameter of the GG blur kernel from image patches containing step edges, and obtain the clear image by non-blind image deblurring. Experimental results validate that the proposed GG blur kernel estimation algorithm outperforms the state-of-the-art ones deploying either parametric (disk and Gaussian) or nonparametric kernels, and consequently benefits the image deblurring process. Yuqi Liu 0005, Xin Du 0005, Shujie Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Boosting Structure Consistency for Multispectral and Multimodal Image RegistrationabstractMultispectral imaging plays a vital role in the area of computer vision and computational photography. As spectral band images can be misaligned due to imaging device movement or alternation, image registration is necessary to avoid spectral information distortion. The current registration measures specialized for multispectral data are typically robust yet complex, requiring excessive computation. The common measures such as sum of squared differences (SSD) and sum of absolute differences (SAD) are computationally efficient whereas they perform poorly on multispectral data. To cope with this challenge, we propose a structure consistency boosting (SCB) transform that aims at boosting the structural similarity of multispectral images. With SCB, the common measures can be employed for multispectral image registration. The SCB transform exploits the fact that inherent edge structures maintain relative saliency locally despite the nonlinear variation between band images. A statistical prior of the natural image, which is based on the gradient-intensity correlation, is explored to build a parametric form of SCB. Experimental results validate that the SCB transform outperforms current similarity enhancement algorithms, and performs better than the state-of-the-art multispectral registration measures. Thanks to the generality of the statistical prior, the SCB transform is also applicable to various multimodal data such as flash/no-flash images and medical images. Si-Yuan Cao, Shujie Chen 0001, Chunguang Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Normalized Total Gradient: A New Measure for Multispectral Image RegistrationabstractImage registration is a fundamental issue in multispectral image processing, and is challenged by two main characteristics of multispectral images. First, the regional intensities can be essentially different between band images. Second, the local contrasts of two difference band images are inconsistent or even reversed. Conventional measures can align images with different regional intensity levels, but may fail in the circumstance of severe local intensity variation. In this paper, a new measure called normalized total gradient is proposed for multispectral image registration. The measure is based on the key assumption (observation) that the gradient of the difference between two aligned band images is sparser than that between two misaligned ones. A registration framework, which incorporates image pyramid and global/local optimization, is further introduced for affine transform. Experimental results validate that the proposed method is not only effective for multispectral image registration, but also applicable to general unimodal/multimodal image registration tasks. It performs better than or comparable to the existing methods, both quantitatively and qualitatively. Shujie Chen 0001, Chunguang Li 0001, John H. Xin |
IEEE Trans. Image Process. | 1 |
| 2016 | Fast Multispectral Imaging by Spatial Pixel-Binning and Spectral UnmixingabstractMultispectral imaging system is of wide application in relevant fields for its capability in acquiring spectral information of scenes. Its limitation is that, due to the large number of spectral channels, the imaging process can be quite time-consuming when capturing high-resolution (HR) multispectral images. To resolve this limitation, this paper proposes a fast multispectral imaging framework based on the image sensor pixel-binning and spectral unmixing techniques. The framework comprises a fast imaging stage and a computational reconstruction stage. In the imaging stage, only a few spectral images are acquired in HR, while most spectral images are acquired in low resolution (LR). The LR images are captured by applying pixel binning on the image sensor, such that the exposure time can be greatly reduced. In the reconstruction stage, an optimal number of basis spectra are computed and the signal-dependent noise statistics are estimated. Then the unknown HR images are efficiently reconstructed by solving a closed-form cost function that models the spatial and spectral degradations. The effectiveness of the proposed framework is evaluated using real-scene multispectral images. Experimental results validate that, in general, the method outperforms the state of the arts in terms of reconstruction accuracy, with additional 20× or more improvement in computational efficiency. Zhi-Wei Pan, Chunguang Li 0001, Shujie Chen 0001, John H. Xin |
IEEE Trans. Image Process. | 4 |
| 2015 | Multispectral Image Out-of-Focus Deblurring Using Interchannel CorrelationabstractOut-of-focus blur occurs frequently in multispectral imaging systems when the camera is well focused at a specific (reference) imaging channel. As the effective focal lengths of the lens are wavelength dependent, the blurriness levels of the images at individual channels are different. This paper proposes a multispectral image deblurring framework to restore out-of-focus spectral images based on the characteristic of interchannel correlation (ICC). The ICC is investigated based on the fact that a high-dimensional color spectrum can be linearly approximated using rather a few number of intrinsic spectra. In the method, the spectral images are classified into an out-of-focus set and a well-focused set via blurriness computation. For each out-of-focus image, a guiding image is derived from the well-focused spectral images and is used as the image prior in the deblurring framework. The out-of-focus blur is modeled as a Gaussian point spread function, which is further employed as the blur kernel prior. The regularization parameters in the image deblurring framework are determined using generalized cross validation, and thus the proposed method does not need any parameter tuning. The experimental results validate that the method performs well on multispectral image deblurring and outperforms the state of the arts. Shujie Chen 0001 |
IEEE Trans. Image Process. | 1 |