EDBT 2026 Demo / reviewers in the wild / expert
Jialu Sui
dblp:347/7093
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-7450-766XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal DatasetabstractIn recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose and personality traits, yet previous research has largely ignored pose data in computational models. To address this gap, we develop a novel multimodal dataset that incorporates full-body pose data. The dataset includes video recordings of 287 participants completing a virtual interview with 36 questions, along with self-reported Big Five personality scores as labels. To effectively utilize this multimodal data, we introduce the Psychology-Inspired Network (PINet), which consists of three key modules: Multimodal Feature Awareness (MFA), Multimodal Feature Interaction (MFI), and Psychology-Informed Modality Correlation Loss (PIMC Loss). The MFA module leverages the Vision Mamba Block to capture comprehensive visual features related to personality, while the MFI module efficiently fuses the multimodal features. The PIMC Loss, grounded in psychological theory, guides the model to emphasize different modalities for different personality dimensions. Experimental results show that the PINet outperforms several state-of-the-art baseline models. Furthermore, the three modules of PINet contribute almost equally to the model’s overall performance. Incorporating pose data significantly enhances the model’s performance, with the pose modality ranking mid-level in importance among the five modalities. These findings address the existing gap in personality-related datasets that lack full-body pose data and provide a new approach for improving the accuracy of personality prediction models, highlighting the importance of integrating psychological insights into AI frameworks. Bin Tang 0010, Keqi Pan, Miao Zheng, Jialu Sui, Dandan Zhu 0001, Cheng-Long Deng, Shu-Guang Kuai |
AAAI | 5 |
| 2024 | FreestyleRet: Retrieving Images from Style-Diversified Queries
Hao Li 0073, Yanhao Jia, Peng Jin 0001, Zesen Cheng, Kehan Li 0002, Jialu Sui, Chang Liu 0047, Li Yuan 0007 |
ECCV (23) | 6 |
| 2024 | Semantic Distortion-Aware Network with Cloud Classification for Remote Sensing Cloud RemovalabstractCloud Removal (CR) utilizing Deep Learning (DL) has been widely employed to enhance the downstream applications of Remote Sensing (RS) satellite imagery affected by cloud coverage. Segmenting thin and thick cloud images into separate training sets and utilizing the CR model with a targeted learning strategy will result in improved performance. In this work, we propose a one-stop automatic cloud processing scheme for CR, including cloud classification and effective CR. To address the varying visibility of thin and thick cloudy images, we train a cloud classification network to distinguish between these two types of cloudy images, subsequently feeding them into two distinct CR networks. Furthermore, we propose a Semantic Distortion-Aware Network (SDAN) designed for similar cloud images after classification. Within SDAN, the Distortion Swin Transformer Block (DSTB) enhances the capability to extract contextual semantic information by incorporating global feature extraction and expression. This enhancement allows for targeted learning for thin or thick cloud images. Experiments conducted on the CR dataset named RICE demonstrate the enhanced performance of our model compared to various existing CR methods. Jialu Sui, Shanjun Xie, Jianuo Jiang, Man-On Pun |
IGARSS | 1 |
| 2024 | A Sam-Empowered Dual-Stream Framework for Scene-Level Local Climate Zone Classification Using Google Earth and Sentinel ImagesabstractRecent advancements in remote sensing (RS)-based methods have shown remarkable effectiveness in large-scale local climate zone classification. However, conventional convolutional neural network (CNN)-based methods encounter limitations in effectively incorporating ground object priors. Additionally, commonly used medium-scale data sources, such as Sentinel-2, face challenges in capturing detailed ground object information. In light of these obstacles, we propose a data fusion method that integrates ground object priors extracted from high-resolution Google imagery with Sentinel-2 multi-spectral imagery. The proposed method introduces a novel dual-stream fusion framework (DF4LCZ-Net), which fully leverages instance-based location features from Google imagery and combines them with the scene-level spatial-spectral features extracted from Sentinel-2 images. For effective feature extraction of Google imagery, a SAM-based Graph Convolutional Network (GCN) branch is designed in the framework. We conducted experiments on a newly created multi-source remote sensing image dataset, and the final classification results show the superiority of our method. Qianqian Wu 0004, Xianping Ma, Jialu Sui, Man-On Pun |
IGARSS | 3 |
| 2024 | FLDCF: A Collaborative Framework for Forgery Localization and Detection in Satellite ImageryabstractSatellite images are highly susceptible to forgery due to various editing techniques. Traditional forgery detection methods, designed for natural images, often fail when applied to satellite images because of differences in sensing technology and processing protocols. The rise of generative models, such as diffusion models, has further complicated the detection of forgeries in satellite images. This study tackles these challenges from both methodological and data perspectives. We introduce a multitask forgery localization and detection collaborative framework (FLDCF), comprising a multiview forgery localization network (M-FLnet) and a forgery detection network. The M-FLnet, leveraging a content-based prior, generates forgery masks that serve as auxiliary information to improve the detection network’s accuracy. Conversely, the detection network refines these masks, reducing noise for authentic images. Furthermore, two novel forgery datasets, namely, Fake-Vaihingen and Fake-LoveDA, are derived from the Vaihingen and LoveDA satellite image sets, respectively, by exploiting the latest generative models. These datasets represent the first open-source datasets for forgery localization and detection in remote sensing. Extensive experimental results on Fake-Vaihingen and Fake-LoveDA demonstrate that the proposed FLDCF can effectively detect sophisticated forgeries in satellite imagery. The source code and datasets in this work are available athttps://github.com/littlebeen/Forgery-localization-for-remote-sensing. Jialu Sui, C.-C. Jay Kuo, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Diffusion Enhancement for Cloud Removal in Ultra-Resolution Remote Sensing ImageryabstractThe presence of cloud layers severely compromises the quality and effectiveness of optical remote sensing (RS) images. However, existing deep-learning (DL)-based cloud removal (CR) techniques, which usually take the fidelity-driven losses as constraints, e.g.,$L_{1}$or$L_{2}$losses, tend to generate smooth results, often failing to reconstruct visually pleasing results and cause semantic loss. To tackle this challenge, this work proposes to encompass enhancements at the data and methodology fronts. On the data side, an ultra-resolution benchmark named CUHK cloud removal (CUHK-CR) of 0.5 m spatial resolution is established. This benchmark incorporates rich detailed textures and diverse cloud coverage, serving as a robust foundation for designing and assessing CR models. From the methodology perspective, a novel diffusion-based framework for CR named diffusion enhancement (DE) is introduced. This framework aims to gradually recover texture details, leveraging a reference visual prior providing foundational structure of the images to enhance inference accuracy. Additionally, a weight allocation (WA) network is developed to dynamically adjust the weights for feature fusion, thereby further improving performance, particularly in the context of ultra-resolution image generation. Furthermore, a coarse-to-fine training strategy is applied to effectively expedite training convergence while reducing the computational complexity required to handle ultra-resolution images. Extensive experiments on the newly established CUHK-CR and existing datasets such as RICE confirm that the proposed DE framework outperforms existing DL-based methods in terms of both perceptual quality and signal fidelity. Jialu Sui, Yiyang Ma, Wenhan Yang, Man-On Pun, Jiaying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | DF4LCZ: A SAM-Empowered Data Fusion Framework for Scene-Level Local Climate Zone ClassificationabstractRecent advances in remote sensing technologies have highlighted their capability for accurate classification of local climate zones (LCZs). However, traditional methods using convolutional neural networks (CNNs) often fall short of effectively incorporating prior knowledge of ground objects. In addition, data sources such as Sentinel-2 struggle with capturing detailed information on ground objects. To address these issues, we introduce a novel data fusion approach that combines high-resolution Google imagery, which provides ground object priors, with Sentinel-2 multispectral imagery. Our method, the Dual-stream Fusion framework for LCZ classification (DF4LCZ), merges instance-based location features from Google imagery and spatial-spectral features from Sentinel-2. This framework is enhanced by a graph convolutional network (GCN) module, powered by the segment anything model (SAM), to improve feature extraction from Google imagery. Concurrently, a 3D-CNN architecture is utilized to process the spectral-spatial features of Sentinel-2 imagery. The effectiveness of DF4LCZ is demonstrated through experiments conducted on a specialized multisource remote sensing image dataset for LCZ classification. The related code and dataset are available athttps://github.com/ctrlovefly/DF4LCZ. Qianqian Wu 0004, Xianping Ma, Jialu Sui, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | DTRN: Dual Transformer Residual Network for Remote Sensing Super-ResolutionabstractThe synergy of the transformer and the convolutional neural network (CNN) has been well regarded as a promising technique for single image super-resolution (SISR) based on low-quality satellite remote sensing images. In this work, a Dual Transformer Residual Network (DTRN) consisting of one transformer branch and one CNN-based residual branch is proposed. More specifically, the transformer branch is designed to capture the global relationships of feature maps by exploiting three pairs of token embedding blocks and convolutional transformer blocks (CTB). Furthermore, the residual branch employs several residual blocks (resblocks) to effectively learn hierarchical features through global feature fusion. Extensive experiments on a large-scale remote sensing dataset called OLI2MSI confirm the superior performance of the proposed DTRN as compared to the existing SISR methods. Jialu Sui, Xianping Ma, Man-On Pun |
IGARSS | 1 |