VLDB 2026 Research / reviewers in the wild / expert
Xin Jin 0005
dblp:68/3340-5
· DBLP profile ↗
82ranked-venue papers
15as first author
74since 2021 · last 2026
0000-0003-2211-2006ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 8 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 25 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space ModelabstractRemote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an important issue is to obtain a color image with high spatial resolution when there is only a PAN image at the input. The existing methods improve spatial resolution using super-resolution (SR) technology and spectral recovery using colorization technology. However, the SR technique cannot improve the spectral resolution, and the colorization technique cannot improve the spatial resolution. Moreover, the pansharpening method needs two registered inputs and can not achieve SR. As a result, an integrated approach is expected. We designed a novel multi-function model (MFmamba) to realize the tasks of SR, spectral recovery, joint SR and spectral recovery through three different inputs. Firstly, MFmamba utilizes UNet++ as the backbone, and a Mamba Upsample Block (MUB) is combined with UNet++. Secondly, a Dual Pool Attention (DPA) is designed to replace the skip connection in UNet++. Finally, a Multi-scale Hybrid Cross Block (MHCB) is proposed for initial feature extraction. Many experiments show that MFmamba is competitive in evaluation metrics and visual results and performs well in the three tasks when only the input PAN image is used. Qianqian Wang 0013, Xin Jin 0005, Michal Wozniak 0001, Shaowen Yao 0001, Wei Zhou 0011 |
AAAI | 3 |
| 2026 | Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningabstractPan-sharpening aims to generate high-resolution multispectral images by integrating the spectral richness of low-resolution multispectral images with the spatial details of high-resolution panchromatic images. Although frequency-domain modeling shows great potential in this field, most existing methods are still limited to spatial-domain processing or fail to effectively capture the contextual interactions between frequency and spatial features. To address these issues, we propose a novel multi-scale frequency-spatial collaborative fusion approach. A Frequency-Spatial U-Net is introduced as the backbone network, in which frequency-spatial modeling blocks are embedded to progressively enhance the frequency-guided spatial contextual modeling capability across layers. To this end, we design a Dual Branch Frequency Attention module that adaptively enhances high- and low-frequency information. In addition, we introduce fine-mid-coarse resolution branches and devise a main-auxiliary multi-scale reconstruction loss to facilitate collaborative optimization. The effectiveness of the proposed model is validated through extensive experiments, demonstrating superior performance in both qualitative and quantitative evaluations. Moreover, our model achieves the fastest inference time among all compared methods, striking an excellent balance between accuracy and efficiency. Huangqimei Zheng, Chengyi Pan, Wei Zhou 0011, Xin Jin 0005 |
AAAI | 5 |
| 2026 | Attention-guided network for infrared unmanned aerial vehicle target detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wangming Lan, Wei Zhou 0011 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Adaptive distributed multi-objective collaborative traffic signal control framework based on multi-agent reinforcement learning
Peisong Huang, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao |
Future Gener. Comput. Syst. | 4 |
| 2026 | CS-DRL: A soft policy update approach for wireless bandwidth allocation using deep reinforcement learning
Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao |
Future Gener. Comput. Syst. | 4 |
| 2026 | STAD-GWR: Continual learning for depression video diagnosis via spatial-temporal attention distillation and gradient-weight-aware regularization
Dongting Cai, Xin Jin 0005 |
Neurocomputing | 2 |
| 2026 | YCSC-UNet: A Y-shaped composite spatial channel network based on U-Net for breast lesion ultrasound image segmentation
Wenjun Ma, Qianqian Wang 0013, Dongjian Yu, Yilong Huang, Xin Jin 0005 |
Neurocomputing | 7 |
| 2026 | ECMSFNet: Real-Time Infrared Small Target Detection Network With Efficient Convolution and Multiscale Feature FusionabstractWith the continuous development of fields such as national defense and military applications, the importance of infrared small target detection (IRSTD) technology based on thermal imaging has become increasingly prominent. However, in practical application scenarios, it remains difficult to effectively extract the features of weak and small targets under low signal-to-noise ratio conditions, while simultaneously suppressing background clutter, preserving target details, and balancing detection accuracy and speed. To address the issues mentioned above, this work proposes a real-time IRSTD network (ECMSFNet) based on efficient convolution and attention-guided multi-scale feature weighting and fusion. During the feature extraction stage, a dual-branch hybrid convolution module (DBHConv) is designed to extract infrared small target features more efficiently. In the feature fusion stage, a three-branch attention-guided module (TBAG) is designed to enhance the input features from both spatial and channel dimensions. By using a three-branch parallel structure to process input features in a differentiated manner, noise is effectively filtered while target detail information is preserved. In addition, to further address the issue of missed detections, a multi-scale feature weighting and fusion module (MSFWF) is designed at the added detection head to adaptively weight the features and optimize the feature propagation path, thereby improving the model’s detection accuracy. Extensive experimental results on multiple datasets demonstrate that the method proposed in this paper outperforms other advanced approaches and achieves a real-time detection speed of 74.63 frames per second. https://github.com/liubiaohua/ECMSFNet. Xin Jin 0005, Biaohua Liu, Shaowen Yao 0001, Puming Wang |
IEEE Internet Things J. | 1 |
| 2026 | GIL-DDI: multi-view graph invariant learning for unknown drug-drug interaction prediction
Yuanxian Li, Yuan Du, Zhenli He, Xin Jin 0005, Cheng Xie 0001 |
Knowl. Inf. Syst. | 5 |
| 2026 | TDFG-GAN: Top-down-feature guided GAN for thermal infrared image colorization
Hongyue Huang, Wei Zhou 0011, Xin Jin 0005 |
Pattern Anal. Appl. | 7 |
| 2026 | YOLO-LIRTU: a lightweight infrared real-time UAV detection framework
Wangming Lan, Hongyue Huang, Xin Jin 0005 |
Pattern Anal. Appl. | 6 |
| 2026 | GILMRec: Graph Invariant Learning for Multimodal RecommendationabstractMultimodal recommendation is a crucial technology on social media platforms. It is widely applied in scenarios such as product recommendation and advertising delivery. However, existing multimodal recommendation approaches often overlook invariant semantic features that persist across modalities, leading to decreased robustness and generalization. To address this limitation, we proposeGILMRec, a novelgraphinvariantlearning-basedmultimodal social mediarecommendation framework. The GILMRec introduces an invariant feature learning strategy to extract invariant features separately from visual and textual modalities and employs an attention-based fusion mechanism to integrate them into a unified embedding. Specifically, we construct modality-specific similarity graphs and apply top-$t$neighbor aggregation, enhancing the consistency of invariant features while effectively suppressing modality-specific noise. Extensive experiments on three Amazon benchmark datasets and a large-scale dataset [baby, sports, clothing, and compact discs (CDs)] demonstrate that GILMRec consistently outperforms twelve state-of-the-art baselines. The results confirm the efficiency of invariant features in capturing robust multimodal representations and improving recommendation performance, particularly in sparse data scenarios. Changlong Fu, Cheng Xie 0001, Zhenli He, Xin Jin 0005, Yun Yang 0003 |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2026 | Designated Masking Propagation Learning for Self-Supervised Heterogeneous Graph RepresentationabstractSelf-supervised heterogeneous graph representation learning (SSHGRL) is a key technique for embedding heterogeneous graphs, enabling effective analysis and modeling of social networks and other graph-structured data, which are central to knowledge discovery and the study of social systems. However, existing SSHGRL methods are hardly applied to large-scale heterogeneous graph environments due to the normally used metapath decomposing mechanism being graph-size-sensitive. Moreover, the existing self-supervised signals are normally created from Shared Mutual Information (SMI) of different graph views that ignore the Non-SMI (NMI) contained in the same view. This results in the model tending to learn insufficient graph representation. To this end, this article proposes a designated masking propagation (DMP) mechanism to process heterogeneous graphs without using metapath. Moreover, based on the DMP graph view, a novel sufficient representation is proposed to learn the effective graph representation by combining both NMI and SMI. Extensive experiments on eight large- and medium-scale heterogeneous graph datasets demonstrate the superiority of our method, setting new state-of-the-art performance in various big data contexts. Haoran Duan 0002, Beibei Yu, Cheng Xie 0001, LinYu Li 0001, Zhenli He, Xin Jin 0005 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2025 | DiffRetouch: Using Diffusion to Retouch on the Shoulder of ExpertsabstractImage retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched results and tends to learn an average style during training, but also lacks sample diversity during inference. In this paper, we propose a diffusion-based method, named DiffRetouch. Thanks to the excellent distribution modeling ability of diffusion, our method can capture the complex fine-retouched distribution covering various visual-pleasing styles in the training data. Moreover, four image attributes are made adjustable to provide a user-friendly editing mechanism. By adjusting these attributes in specified ranges, users are allowed to customize preferred styles within the learned fine-retouched distribution. Additionally, the affine bilateral grid and contrastive learning scheme are introduced to handle the problem of texture distortion and control insensitivity respectively. Extensive experiments have demonstrated the superior performance of our method on visually appealing and sample diversity. Zheng-Peng Duan, Jiawei Zhang 0002, Zheng Lin 0005, Xin Jin 0005, Xundong Wang, Dongqing Zou, Chunle Guo, Chongyi Li |
AAAI | 4 |
| 2025 | Classic Video Denoising in a Machine Learning World: Robust, Fast, and ControllableabstractDenoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures due to discrepancies between training data distributions and the wide variety of noise patterns found in real-world videos. These methods also tend to be slow and lack user control. In contrast, traditional denoising methods perform reliably on in-the-wild videos and run relatively quickly on modern hardware. However, they require manually tuning parameters for each input video, which is not only tedious but also requires skill. We bridge the gap between these two paradigms by proposing a differentiable denoising pipeline based on traditional methods. A neural network is then trained to predict the optimal denoising parameters for each specific input, resulting in a robust and efficient approach that also supports user control. Xin Jin 0005, Simon Niklaus, Zhoutong Zhang, Zhihao Xia, Chunle Guo, Jiawen Chen 0001, Chongyi Li |
CVPR | 1 |
| 2025 | Towards RAW Object Detection in Diverse ConditionsabstractExisting object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might lose crucial information for detection, especially under complex light and weather conditions. We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions. Based on AODRaw that supports RAW and sRGB object detection, we provide a comprehensive benchmark for evaluating current detection methods. We find that sRGB pre-training constrains the potential of RAW object detection due to the domain gap between sRGB and RAW, prompting us to directly pre-train on the RAW domain. However, it is harder for RAW pre-training to learn rich representations than sRGB pre-training. To assist RAW pre-training, we distill the knowledge from an off-the-shelf model pre-trained on the sRGB domain. As a result, we achieve substantial improvements under diverse and adverse conditions without relying on extra pre-processing modules. The code and dataset are available at https://github.com/lzyhha/AODRaw. Zhongyu Li 0006, Xin Jin 0005, Bo-Yuan Sun, Chunle Guo, Ming-Ming Cheng |
CVPR | 2 |
| 2025 | DiT4SR: Taming Diffusion Transformer for Real-World Image Super-ResolutionabstractLarge-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UNet-based architecture in image generation, which also raises the question: Can we adopt the advanced DiT-based diffusion model for Real-ISR? To this end, we propose our DiT4SR, one of the pioneering works to tame the large-scale DiT model for Real-ISR. Instead of directly injecting embeddings extracted from low-resolution (LR) images like ControlNet, we integrate the LR embeddings into the original attention mechanism of DiT, allowing for the bidirectional flow of information between the LR latent and the generated latent. The sufficient interaction of these two streams allows the LR stream to evolve with the diffusion process, producing progressively refined guidance that better aligns with the generated latent at each diffusion step. Additionally, the LR guidance is injected into the generated latent via a cross-stream convolution layer, compensating for DiT's limited ability to capture local information. These simple but effective designs endow the DiT model with superior performance in Real-ISR, which is demonstrated by extensive experiments. Project Page: https://adam-duan.github.io/projects/dit4sr/. Zheng-Peng Duan, Jiawei Zhang 0002, Xin Jin 0005, Zheng Xiong, Dongqing Zou, Jimmy S. J. Ren, Chunle Guo, Chongyi Li |
ICCV | 3 |
| 2025 | CoT-NER: A Reasoning Method via Chain of Thought for Chinese Named Entity Recognition
ZhaoKe Long, Shengfa Miao, Puming Wang, Xin Jin 0005, Jing Niu, ShuangFeng Cai |
ICIC (24) | 5 |
| 2025 | GCTAM: Global and Contextual Truncated Affinity Combined Maximization Model For Unsupervised Graph Anomaly DetectionabstractAnomalies often occur in real-world information networks/graphs, such as malevolent users, malicious comments, banned users, and fake news in social graphs. The latest graph anomaly detection methods use a novel mechanism called truncated affinity maximization (TAM) to detect anomaly nodes without using any label information and achieve impressive results. TAM maximizes the affinities among the normal nodes while truncating the affinities of the anomalous nodes to identify the anomalies. However, existing TAM-based methods truncate suspicious nodes according to a rigid threshold that ignores the specificity and high-order affinities of different nodes. This inevitably causes inefficient truncations from both normal and anomalous nodes, limiting the effectiveness of anomaly detection. To this end, this paper proposes a novel truncation model combining contextual and global affinity to truncate the anomalous nodes. The core idea of the work is to use contextual truncation to decrease the affinity of anomalous nodes, while global truncation increases the affinity of normal nodes. Extensive experiments on massive real-world datasets show that our method surpasses peer methods in most graph anomaly detection tasks. In highlights, compared with previous state-of-the-art methods, the proposed method has +15% ~ +20% improvements in two famous real-world datasets, Amazon and YelpChi. Notably, our method works well in large datasets, Amazin-all and YelpChi-all, and achieves the best results, while most previous models cannot complete the tasks. Zhenli He, Cheng Xie 0001, Xin Jin 0005 |
IJCAI | 5 |
| 2025 | Spatial-Aware Multi-Modal Information Fusion for Food Nutrition EstimationabstractFood nutrition assessment plays a crucial role in maintaining health, preventing diseases, and promoting scientific dietary habits. However, existing nutrition assessment methods often fail to fully consider the relationships between tasks, leading to limited overall performance. Specifically, these methods suffer from three major challenges: (1) task conflicts, where different tasks compete during joint optimization, leading to suboptimal overall performance; (2) varying training difficulties among tasks, leading to imbalanced learning and subpar model generalization; and (3) the small-scale and complex distribution of datasets, which limits the robustness of learned representations. To address these issues, we propose a novel method that reduces interference between tasks, dynamically focuses on more challenging tasks, and incorporates 3D spatial awareness to enhance multi-modal feature representation. First, we decouple the prediction network from the backbone and introduce a CAMTH (Cross-Attention-Based Multi-Task Head Module), effectively mitigating task interference and fully leveraging each task's learning potential. Second, we improve the loss function to adaptively focus on more challenging tasks, improving overall model performance. Third, we design a 3D-FEM (3D Feature Extraction Module) and MMFF (Multi-Modal Feature Fusion Module), enabling the model to fully exploit the spatial information of food and enhance the food's multi-modal feature representation. We validate our method through extensive experiments on the Nutrition5K dataset, comparing it with state-of-the-art (SOTA) models. The results show that our method achieves superior performance in nutrition estimation, demonstrating the effectiveness of our method. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Multimedia | 3 |
| 2025 | UltraLED: Learning to See Everything in Ultra-High Dynamic Range ScenesabstractUltra-high dynamic range (UHDR) scenes exhibit pronounced exposure disparities between bright and dark regions. Such conditions are Ultra-high dynamic range (UHDR) scenes exhibit significant exposure disparities between bright and dark regions. Such conditions are commonly encountered in nighttime scenes with light sources. Even with standard exposure settings, a bimodal intensity distribution with boundary peaks often emerges, making it difficult to preserve both highlight and shadow details simultaneously. RGB-based bracketing methods can capture details at both ends using short-long exposure pairs, but are susceptible to misalignment and ghosting artifacts. We found that a short-exposure image already retains sufficient highlight detail. The main challenge of UHDR reconstruction lies in denoising and recovering information in dark regions. In comparison to the RGB images, RAW images, thanks to their higher bit depth and more predictable noise characteristics, offer greater potential for addressing this challenge. This raises a key question: can we learn to see everything in UHDR scenes using only a single short-exposure RAW image? In this study, we rely solely on a single short-exposure frame, which inherently avoids ghosting and motion blur, making it particularly robust in dynamic scenes. To achieve that, we introduce UltraLED, a two-stage framework that performs exposure correction via a ratio map to balance dynamic range, followed by a brightness-aware RAW denoiser to enhance detail recovery in dark regions. To support this setting, we design a 9-stop bracketing pipeline to synthesize realistic UHDR images and contribute a corresponding dataset based on diverse scenes, using only the shortest exposure as input for reconstruction. Extensive experiments show that UltraLED significantly outperforms existing single-frame approaches. Our code and dataset are made publicly available at https://srameo.github.io/projects/ultraled. Yuang Meng, Xin Jin 0005, Lina Lei, Chunle Guo, Chongyi Li |
NeurIPS | 2 |
| 2025 | Transferable adversarial attacks for multi-model systems coupling image fusion with classification modelsabstractAbstract Image preprocessing models typically serve as the initial step in advanced visual tasks, aiming to enhance the performance of subsequent tasks. For example, multi-focus image fusion technology significantly improves the performance of downstream semantic classification tasks. However, with the advancement of adversarial attack techniques, these models are facing significant challenges. Previous research has only explored the impact of adversarial attacks on the performance of individual models, lacking an in-depth investigation into the robustness of tasks involving the combination of multiple models. This study aims to delve into the robustness issues of tasks that combine multi-focus image fusion and image classification. To address this challenge, we have designed a new adversarial attack generator specifically for scenarios that combine multi-focus image fusion with image classification. This attack method uses a decision map surrogate model and a binary weight map to precisely add adversarial perturbations to the effective information parts of multi-focus images. It also incorporates attention mechanisms and Grad-CAM technology to optimize the perturbation areas, aiming to disrupt the key features of the fused image to improve the transferability of the attack. Comprehensive experimental results show that this method significantly improves the efficiency of attacks on downstream classification tasks while maintaining the effectiveness of the fusion model. Xin Jin 0005, Xueshuai Gao, Puming Wang, Shaowen Yao 0001, Wei Zhou 0011 |
Cybersecur. | 2 |
| 2025 | Polyhedral representations with high-frequency for three-dimensional point cloud classification
Xiaoxin Mao, Xue Li 0009, Puming Wang, Xin Jin 0005, Shengfa Miao, Shaowen Yao 0001, Siwang Yang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | SR_ColorNet: Multi-path attention aggregated and mask enhanced network for the super resolution and colorization of panchromatic image
Qianqian Wang 0013, Shengfa Miao, Xin Jin 0005, Shin-Jye Lee, Michal Wozniak 0001, Shaowen Yao 0001 |
Expert Syst. Appl. | 4 |
| 2025 | IDAD: An improved tensor train based distributed DDoS attack detection framework and its application in complex networksabstractWith the vigorous development of Internet technology, the scale of systems in the network has increased sharply, which provides a great opportunity for potential attacks, especially the Distributed Denial of Service (DDoS) attack. In this case, detecting DDoS attacks is critical to system security. However, current detection methods exhibit limitations, leading to compromises in accuracy and efficiency. To cope with it, three key strategies are implemented in this paper: (i) Using tensors to model large-scale and heterogeneous data in complex networks; (ii) Proposing a denoising algorithm based on the improved and distributed tensor train (IDTT) decomposition, which optimizes the tensor train(TT) decomposition in terms of parallel computation and low-rank estimation; (iii) Combining (i), (ii) and Light Gradient Boosting Machine (LightGBM) classification model, an efficient DDoS attack detection framework is proposed. Datasets CIC-DDoS2019 and NSL-KDD are used to evaluate the framework, and results demonstrate that accuracy can reach 99.19% while having the characteristics of low storage consumption and well speedup ratio. Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao, Min An |
Future Gener. Comput. Syst. | 4 |
| 2025 | DMNet: A Dense Multiscale Feature Extraction Network With Two-Stage Training for Infrared-Visible Image FusionabstractWith the increasing need for intelligent and secure multimedia systems, infrared and visible image fusion (IVIF) has garnered a lot of attention due to its ability to overcome the limitations of a single sensor and integrate unique information from different modalities. However, it is common to overlook how the spatial frequency information of visible and infrared images differs. A less thorough feature extraction may result from many approaches’ inability to reconcile the extraction of both global and local information. To solve the aforementioned difficulties, we propose a dense multiscale fusion network DMNet. Through a dual-stream collaborative feature decoupling, the proposed network optimizes both the encoder–decoder network and the diffusion model to extract multimodal information more comprehensively. Specifically, the three-stage progressive encoder sequentially integrates dense transformer block (DTB) and dense invertible neural network block (DIB) to achieve global feature extraction and multimodal feature decoupling. Our proposed channel and spatial attention block (CSAB) selectively focuses on the important feature maps to better capture the critical information. Additionally, multiscale latent features are extracted by the diffusion module (DM) to enhance the representation of cross-modal latent features. As demonstrated by extensive experiments, DMNet outperforms representative state-of-the-art methods. Furthermore, we conduct sufficient ablation experiments to validate each module’s effectiveness, and we demonstrate that DMNet can enhance downstream infrared-visible object detection performance. Our fused results and code will be accessible athttps://github.com/Pancy9476/DMNet. Chengyi Pan, Huangqimei Zheng, Hongyue Huang, Xin Jin 0005, Keqin Li 0001, Wei Zhou 0011 |
IEEE Internet Things J. | 5 |
| 2025 | Robust manipulated media localization and detection based on high frequency and texture featuresabstractAdvances in facial manipulation techniques have resulted in the increasing trend of realistic and indistinguishable identity swap media, which mislead the viewers and accompanied by severe security concerns. While current deepfake detectors demonstrate strong performance under high-quality conditions, they still face notable limitations. This article proposes a novel framework mining high frequency and degraded texture features for locating manipulated traces and improving the generalization ability. To improve the universality of the proposed detector, we design the Multi-feature Mining Stream for capturing the global and subtle discrepancies of undegraded images. Moreover, the Encoder-Decoder Structure is introduced for gaining high localization accuracy and full resolution manipulated regions. This work attempts to solve the tampered region localization issue and achieve face forgery image detection at the meantime. This contributes to help the model perform a more effective differentiation between real and fake content when confronted with high- or low-quality compressed images. Comprehensive experiments on the popular FaceForensics++, Celeb-DF, and DFDC datasets demonstrate the superior performance and robustness of our proposed framework, in particular, achieving performance improvements ranging from 1% to 10% in comparison with the most recent related work. Shuai Liu 0009, Shengfa Miao, Huasong Yi, Xin Jin 0005, Yuru Kou, Hanxian Duan |
Discov. Comput. | 5 |
| 2025 | HPM-GMN: Hierarchical Pooling Multi-level Graph Matching Network
Shengfa Miao, YongKang Mu, Yuling Tian, Yesen Liu, Kuang Li, Puming Wang, Xin Jin 0005, Shaowen Yao 0001 |
Knowl. Based Syst. | 10 |
| 2025 | Mutli-focus image fusion based on guided filter and image matting network
Puchao Zhu, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001 |
Multim. Tools Appl. | 4 |
| 2025 | Learning dual aggregate features for face forgery detection
Yuru Kou, Xin Jin 0005, Shengfa Miao, Xing Chu |
Neural Comput. Appl. | 4 |
| 2025 | Food3D: Text-Driven Customizable 3D Food Generation With Gaussian SplattingabstractRealistic 3D food creation generation plays a critical role in applications such as nutritional assessment, advertising, and virtual content creation. The existing text-to-3D models typically begin by initializing a 3D representation, which is subsequently refined using supervision from a text-to-image model to obtain the final 3D output. In this work, we present Food3D, a novel framework for 3D food generation designed to address two main limitations of current models. First, the limitation of initialization in 3D generation: poor initialization can result in the generated 3D food lacking crucial details and realism, thereby reducing its quality. To address this issue, we propose a generalized method named Food3D-G, which uses Mamba-based initialization to improve the starting point of the initialization process, thereby enhancing the visual fidelity and quality of the generated 3D food. Second, the limitation of text-to-image models: current text-to-3D models often rely on text-to-image models for supervision. However, a considerable gap persists between the generated images and real-world visuals, particularly when modeling complex food structures. These models fail to accurately capture the fine details and textures, which negatively impacts the quality and realism of the generated 3D food models. To address this limitation, we propose a customizable method for personalized 3D food generation, termed Food3D-C. This method employs a dual-branch diffusion model that effectively captures intricate details, particularly in complex food structures. Within the Food3D framework, both proposed methods incorporate 3D Gaussian splatting (3D GS) and a schedulable interval score matching (S-ISM) algorithm to enhance shape and texture generation. Extensive experiments demonstrate that Food3D achieves state-of-the-art performance, with substantial improvements in detail, shape accuracy, and overall visual realism. Project page and source codes: https://yudongjian.github.io/Food3D/. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shaowen Yao 0001, Shuqiang Jiang |
IEEE Trans. Image Process. | 3 |
| 2025 | GDRNet: a channel grouping based time-slice dilated residual network for long-term time-series forecasting
Qingda Bao, Shengfa Miao, Xin Jin 0005, Puming Wang, Shaowen Yao 0001, Da Hu, Ruoshu Wang |
J. Supercomput. | 4 |
| 2025 | SDHNet: a sampling-based dual-stream hybrid network for long-term time series forecasting
Shengfa Miao, Shaowen Yao 0001, Xin Jin 0005, Xing Chu, Yuling Tian, Ruoshu Wang |
J. Supercomput. | 4 |
| 2025 | BDIP: An Efficient Big Data-Driven Information Processing Framework and Its Application in DDoS Attack DetectionabstractWith the rapid advancement of 5G communication technology in the era of big data, massive terminal devices connected to the Internet have dramatically increased the scale of network, generating a large amount of high-dimensional and heterogeneous information. This not only enhances the difficulty of information processing in the network, but also poses a severe challenge to data storage and calculation, which has become a big data problem to be solved urgently. To cope with it, this paper proposes an efficient information processing framework and applies it to Distributed Denial of Service (DDoS) attack detection. Overall, three major highlights are made: (i) Tensor is used to represent multi-modal information in large-scale networks; (ii) A novel denoising algorithm based on tensor train(TT) decomposition is proposed, focused on optimizing both computation and correlation; (iii) A big data-driven information processing framework is developed, which includes information preprocessing, denoising and classification. Results in case study indicate that the framework can achieve an accuracy of 99.19%, all while maintaining the great storage advantage, well speedup ratio and strong computing capabilities under the same computational complexity. It can also be generalized to other network data processing scenarios. Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao, Sizhang Li, Min An |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Diverse and High-Quality Food Image Generation from Only Food NamesabstractFood image generation holds promising application prospects in food design, advertising, and food education. However, the existing methods rely on information such as recipes, ingredients, or food names, which leads to generated food images with less intra-class diversity. When recipes, ingredients, and food names are identical for the same food, the real-world images may vary significantly in appearance. The question of how to simultaneously ensure the quality and diversity of the generated images is a key issue. To this end, we employ pre-trained diffusion model and Transformer to propose a method for generating diverse and high-quality images of both Chinese and Western food, named CW-Food. Different from previous works that utilize an overall food feature to generate new images, CW-Food first decouples the food images to obtain common intra-class features and private instance features. Additionally, we design a Transformer-based feature fusion module to integrate the common and private features, in order to avoid the shortcomings of conventional methods. Moreover, we also utilize a pre-trained diffusion model as our backbone, which is fine-tuned using LoRA with the fused multi-variate features. Extensive experiments on four datasets demonstrate the advantages of our proposed method, producing diverse and high-quality food images encompassing both Chinese and Western cuisines. To the best of our knowledge, our work is the first attempt to generate Chinese food images using only food names. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Crafting imperceptible and transferable adversarial examples: leveraging conditional residual generator and wavelet transforms to deceive deepfake detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011 |
Vis. Comput. | 2 |
| 2025 | DS-GAN: a dual sub-structure GAN for thermal infrared image colorization using U-Net with ConvNeXt and multi-scale large kernel attention
Guoliang Yao, Xin Jin 0005, Michal Wozniak 0001, Shengfa Miao, Shaowen Yao 0001, Wei Zhou 0011 |
Vis. Comput. | 2 |
| 2024 | RIHNet: A Robust Image Hiding Method for JPEG Compression
Xin Jin 0005, Zien Cheng, Weiping Ding 0001, Yunyun Dong, Liwen Wu, Shengfa Miao |
ICIC (10) | 1 |
| 2024 | Boosting the Transferability of Adversarial Examples via Adaptive Attention and Gradient Purification MethodsabstractDeep neural networks are shown to be vulnerable to adversarial examples. Recently, various methods have been proposed to improve the transferability of adversarial examples. However, most of the existing methods add perturbations to the whole image without discrimination, causing the visual quality of the adversarial examples to degrade drastically. In addition, existing attack methods ignore the gradient information of secondary features, which affects the accuracy of generating adversarial perturbations. In this work, we propose Adaptive Attention and Gradient Purification Attack (AAGP) to address such issues. Specifically, we judge the mean and standard deviation of the gradient values to find out where the model is interested. Since different models share similar regions of attention, adding perturbations only to such areas can reduce the addition of adversarial perturbation and can also lead to better transferability of adversarial examples to other models. In addition, we disrupt the correlation of pixels at the distribution of secondary features by random discarding pixels in low-attention areas, generating more transferable perturbations through more accurate gradient information. Experimental results on ImageNet show that our method enhances the visibility of the adversarial examples and their transferability compared with several advanced baselines. Liwen Wu, Lei Zhao 0013, Bin Pu, Xin Jin 0005, Shaowen Yao 0001 |
IJCNN | 5 |
| 2024 | Enhanced YOLOv7 Model for Aerial Drone Detection in Complex EnvironmentsabstractWith the growing popularity of Unmanned Aerial Vehicles (UAVs) in civilian, commercial, and military applications, the need for robust drone detection systems has become increasingly urgent. However, due to the small size of drones when observed from a distance, most traditional machine learning and two-stage deep learning detection methods currently struggle to capture effective feature information of drones in complex image backgrounds, and often fail to meet the requirements for real-time performance. To address this challenge, we have introduced an improved version of the YOLOv7-Tiny model, known as YOLOv7-ADD, which significantly enhances the detection performance of small target drones at various distances and under complex backgrounds. The model integrates the Scylla-IoU (SIoU) loss function to improve the accuracy of bounding box regression and employs the BiFormer attention mechanism, which dynamically focuses on key features of drones within the detection scene, enhancing the model’s recognition capabilities for drones. Furthermore, the introduction of the Diverse Branch Block (DBB) helps the model capture multi-scale features, optimizing the detection effect for drones of various sizes. Extensive experiments on drones dataset have demonstrated the superior performance of YOLOv7-ADD. It achieved a 2.7% improvement in [email protected] and a 1.1% increase in [email protected]:0.95, while maintaining high FPS detection performance. This provides an efficient drone detection solution for aerial surveillance. The code is available on https://github.com/FuChanglong/ADD.git. Changlong Fu, Yasu Wu, Xin Jin 0005, Puming Wang |
ISPA | 3 |
| 2024 | Dif-GAN: A Generative Adversarial Network with Multi-Scale Attention and Diffusion Models for Infrared-Visible Image FusionabstractTo obtain fused images with rich information, visible and infrared images are combined. Most current fusion techniques provide decent results. However, they have shortcomings in extracting information of the source images. This limitation prevents the fused images from adequately considering thermal radiation regions and texture details. As a result, the detailed texture information of the source visible image in the final fusion image is much more than the thermal target information of the source infrared image, or vice versa. Since features at a single scale fail to adequately capture the spatial details of complex scenes, a multi-scale attention network is used to extract the deep feature information of source images. For latent variable issues, the Expectation Maximization (EM) technique can yield maximum likelihood estimates. This not only stabilizes the training of the Generative Adversarial Network (GAN) but also aids in addressing the issue of labels lacking in the fusion of visible and infrared images. Although the EM algorithm framework can greatly enhance the training stability of GAN models, the improvement in fusion quality is not large. Therefore, a diffusion model is introduced into the generator to capture the potential joint structure information between infrared and visible images. Massive experiments show that Dif-GAN outperforms the state-of-the-art. Chengyi Pan, Xiuliang Xi, Xin Jin 0005, Huangqimei Zheng, Puming Wang, Qiang Jiang |
ISPA | 3 |
| 2024 | Online public opinion time series prediction based on improved N-Beat and multimodal hybrid fusionabstractSocial media offers a promising way to analyze online public opinion, which has drawn extensive attention from various sectors. In academia, most studies focus on predicting public opinion using unimodal time series methods, paying little attention to multimodal approaches. However, public opinion may be affected by various complex social factors, so it is necessary to explore multimodal elements. Based on the N-Beats model, we propose a novel model, HFN-BeatsConv, which employs a powerful modal alignment strategy, 3D-TCN. Most fuses multimodal data for public opinion prediction. The model employs a component, 3D-TCN, for modal alignment, differs from other research in that it focuses on the time at which the text appears. Subsequently, the model, HFN-BeatsConv, the N-Beats model is enhanced through the utilization of 3D-TCN, which enables the processing of multivariate time series data and the reduction of multimodal time series forecast error. To increase the usability and sustainability of research, this study provides a valuable social media dataset, as a supplementary feature of time series prediction. Through extensive experiments, the proposed method outperforms the existing methods. Yuling Tian, Shengfa Miao, Shaowen Yao 0001, Puming Wang, Xin Jin 0005 |
ISPA | 5 |
| 2024 | Lighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View SynthesisabstractVolumetric rendering-based methods, like NeRF, excel in HDR view synthesis from RAW images, especially for nighttime scenes. They suffer from long training times and cannot perform real-time rendering due to dense sampling requirements. The advent of 3D Gaussian Splatting (3DGS) enables real-time rendering and faster training. However, implementing RAW image-based view synthesis directly using 3DGS is challenging due to its inherent drawbacks: 1) in nighttime scenes, extremely low SNR leads to poor structure-from-motion (SfM) estimation in dis- tant views; 2) the limited representation capacity of the spherical harmonics (SH) function is unsuitable for RAW linear color space; and 3) inaccurate scene structure hampers downstream tasks such as refocusing. To address these issues, we propose LE3D (Lighting Every darkness with 3DGS). Our method proposes Cone Scatter Initialization to enrich the estimation of SfM and replaces SH with a Color MLP to represent the RAW linear color space. Additionally, we introduce depth distortion and near-far regularizations to improve the accuracy of scene structure for down- stream tasks. These designs enable LE3D to perform real-time novel view synthesis, HDR rendering, refocusing, and tone-mapping changes. Compared to previous vol- umetric rendering-based methods, LE3D reduces training time to 1% and improves rendering speed by up to 4,000 times for 2K resolution images in terms of FPS. Code and viewer can be found in https://srameo.github.io/projects/le3d. Xin Jin 0005, Pengyi Jiao, Zheng-Peng Duan, Xingchao Yang, Chongyi Li, Chunle Guo, Bo Ren 0003 |
NeurIPS | 1 |
| 2024 | DTA: distribution transform-based attack for query-limited scenarioabstractAbstract In generating adversarial examples, the conventional black-box attack methods rely on sufficient feedback from the to-be-attacked models by repeatedly querying until the attack is successful, which usually results in thousands of trials during an attack. This may be unacceptable in real applications since Machine Learning as a Service Platform (MLaaS) usually only returns the final result (i.e., hard-label) to the client and a system equipped with certain defense mechanisms could easily detect malicious queries. By contrast, a feasible way is a hard-label attack that simulates an attacked action being permitted to conduct a limited number of queries. To implement this idea, in this paper, we bypass the dependency on the to-be-attacked model and benefit from the characteristics of the distributions of adversarial examples to reformulate the attack problem in a distribution transform manner and propose a distribution transform-based attack (DTA). DTA builds a statistical mapping from the benign example to its adversarial counterparts by tackling the conditional likelihood under the hard-label black-box settings. In this way, it is no longer necessary to query the target model frequently. A well-trained DTA model can directly and efficiently generate a batch of adversarial examples for a certain input, which can be used to attack un-seen models based on the assumed transferability. Furthermore, we surprisingly find that the well-trained DTA model is not sensitive to the semantic spaces of the training dataset, meaning that the model yields acceptable attack performance on other datasets. Extensive experiments validate the effectiveness of the proposed idea and the superiority of DTA over the state-of-the-art. Renyang Liu 0001, Wei Zhou 0011, Xin Jin 0005, Yuanyu Wang, Ruxin Wang 0002 |
Cybersecur. | 3 |
| 2024 | SIHNet: A safe image hiding method with less information leakingabstractAbstract Image hiding is a task that hides secret images into cover images. The purposes of image hiding are to ensure the secret images are invisible to the human and the secret images can be recovered. The current state‐of‐the‐art steganography methods run the risk of secret information leakage. A safe image hiding network (SIHNet) is presented to reduce the leakage of secret information. Based on some phenomena of image hiding methods which use invertible neural network, a reversible secret image processing (SIP) module is proposed to make the secret images suitable for hiding and make the stego images leak less secret information. Besides, a reversible lost information hiding (LIH) module is used to hide the lost information into the cover images, thus the method can recover the secret images better than the method that uses random noise to replace the lost information. Experimental results show that SIHNet outperforms other state‐of‐the‐art methods on the PSNR and SSIM values of the recovered secret images and the stego images. Besides, residual images of other state‐of‐the‐art methods all contain information about secret images while residual images of SIHNet leak almost no secret information. Thus the method can prevent the listener of transmission channel from obtaining the information of the secret image through the residual image, which means SIHNet performs better in security than other state‐of‐the‐art methods. Zien Cheng, Xin Jin 0005, Liwen Wu, Yunyun Dong, Wei Zhou 0011 |
IET Image Process. | 2 |
| 2024 | MCDC-Net: Multi-scale forgery image detection network based on central difference convolutionabstractAbstract Generative Adversarial Networks (GANs) emerged thanks to the development of deep neural networks. Forgery images generated by various variants of GANs are widely spread on the Internet, which may be damage personal credibility and cause huge property losses. Thus, numerous methods are proposed to detect forgery images, but most of them are designed to detect forgery faces. Therefore, a method to detect forgery images of various scenes is proposed. In this work, central difference convolution and vanilla convolution (CDC‐Mix) are mixed after considering the depth and width features of neural networks and analyzing the influence of attention on network performance. Based on CDC‐Mix, a separable convolution (SeparableCDC‐Mix) is proposed. The proposed method consists of three parts: (1) CDC‐Mix and SeparableCDC‐Mix are used to extract the gradient information and texture features; (2) CDCM is used to extract the multi‐scale information of the image; (3) multi‐scale fusion module (MS‐Fusion) is used to fuse the multi‐scale information from different locations of the network. A large number of experiments have been carried out on several datasets generated by GAN, and the experimental results show that the proposed method has a great improvement compared with the existing advanced methods. Defen He, Xin Jin 0005, Zien Cheng, Shuai Liu 0009, Shaowen Yao 0001, Wei Zhou 0011 |
IET Image Process. | 3 |
| 2024 | RIHINNet: A robust image hiding method against JPEG compression based on invertible neural networkabstractAbstract Image hiding is a task that embeds secret images in digital images without being detected. The performance of image hiding has been greatly improved by using the invertible neural network. However, current image hiding methods are less robust in the face of Joint Photographic Experts Group (JPEG) compression. The secret image cannot be extracted from the stego image after JPEG compression of the stego image. Some methods show good robustness for some certain JPEG compression quality factors but poor robustness for other common JPEG compression quality factors. An image‐hiding network (RIHINNet) that is robust to all common JPEG compression quality factors is proposed. First of all, the loss function is redesigned; thus, the secret image is hidden as much as possible in the area that is less likely to be changed after JPEG compression. Second, the classifier is designed, which can help the model to select the extractor according to the range of JPEG compression degree. Finally, the interval robustness of the secret image extraction is improved through the design of a denoising module. Experimental results show that this RIHINNet outperforms other state‐of‐the‐art image‐hiding methods in the face of JPEG compressed noise with random compression quality factors, with more than 10 dB peak signal‐to‐noise ratio improvement in secret image recovery on ImageNet, COCO and DIV2K datasets. Xin Jin 0005, Chengyi Pan, Zien Cheng, Yunyun Dong |
IET Image Process. | 1 |
| 2024 | A novel multi-modal incremental tensor decomposition for anomaly detection in large-scale networks
Rongqiao Fan, Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001 |
Inf. Sci. | 6 |
| 2023 | Underwater Ranker: Learn Which Is Better and How to Be BetterabstractIn this paper, we present a ranking-based underwater image quality assessment (UIQA) method, abbreviated as URanker. The URanker is built on the efficient conv-attentional image Transformer. In terms of underwater images, we specially devise (1) the histogram prior that embeds the color distribution of an underwater image as histogram token to attend global degradation and (2) the dynamic cross-scale correspondence to model local degradation. The final prediction depends on the class tokens from different scales, which comprehensively considers multi-scale dependencies. With the margin ranking loss, our URanker can accurately rank the order of underwater images of the same scene enhanced by different underwater image enhancement (UIE) algorithms according to their visual quality. To achieve that, we also contribute a dataset, URankerSet, containing sufficient results enhanced by different UIE algorithms and the corresponding perceptual rankings, to train our URanker. Apart from the good performance of URanker, we found that a simple U-shape UIE network can obtain promising performance when it is coupled with our pre-trained URanker as additional supervision. In addition, we also propose a normalization tail that can significantly improve the performance of UIE networks. Extensive experiments demonstrate the state-of-the-art performance of our method. The key designs of our method are discussed. Our code and dataset are available at https://li-chongyi.github.io/URanker_files/. Chunle Guo, Xin Jin 0005, Linghao Han, Weidong Zhang 0007, Chongyi Li |
AAAI | 3 |
| 2023 | DNF: Decouple and Feedback Network for Seeing in the DarkabstractThe exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clean and RAW-to-sRGB, misleads the single-stage methods due to the domain ambiguity. The multi-stage methods propagate the information merely through the resulting image of each stage, neglecting the abundant features in the lossy image-level dataflow. In this paper, we probe a generalized solution to these bottlenecks and propose a Decouple aNd Feedback framework, abbreviated as DNF. To mitigate the domain ambiguity, domain-specific subtasks are decoupled, along with fully utilizing the unique properties in RAW and sRGB domains. The feature propagation across stages with a feedback mechanism avoids the information loss caused by image-level dataflow. The two key insights of our method resolve the inherent limitations of RAW data-based low-light image enhancement satisfactorily, empowering our method to outperform the previous state-of-the-art method by a large margin with only 19% parameters, achieving 0.97dB and 1.30dB PSNR improvements on the Sony and Fuji subsets of SID. Xin Jin 0005, Linghao Han, Zhen Li 0031, Chunle Guo, Chongyi Li |
CVPR | 1 |
| 2023 | Multi-Layer Seasonal Perception Network for Time Series ForecastingabstractSeasonal time series contain rich long-term dependencies. How to make good use of the seasonal information to predict the future is still a challenging problem. In this paper, we propose a neural network model called Multilayer Seasonal Perception Network (MSPNet) to predict seasonal time series. Firstly, we propose the idea of seasonal alignment, which converts univariate time series into multivariate time series, in order to capture seasonal features more effectively. Secondly, we extract the seasonal features and historical dependencies, using the Multi-layer Seasonal Perception Attention. Finally, we combine the obtained nonlinear features with linear features to conduct the final prediction. Experimental verification shows that the proposed MSPNet model is significantly superior to the baseline methods on multiple public datasets. The source code and datasets are available at https://github.com/MasterofEating/MSPNet Ruoshu Wang, Shengfa Miao, Di Liu 0002, Xin Jin 0005, Weisheng Zhang |
ICASSP | 4 |
| 2023 | Lighting Every Darkness in Two Pairs : A Calibration-Free Pipeline for RAW DenoisingabstractCalibration-based methods have dominated RAW image denoising under extremely low-light environments. However, these methods suffer from several main deficiencies: 1) the calibration procedure is laborious and time-consuming, 2) denoisers for different cameras are difficult to transfer, and 3) the discrepancy between synthetic noise and real noise is enlarged by high digital gain. To overcome the above shortcomings, we propose a calibration-free pipeline for Lighting Every Drakness (LED), regardless of the digital gain or camera sensor. Instead of calibrating the noise parameters and training repeatedly, our method could adapt to a target camera only with few-shot paired data and fine-tuning. In addition, well-designed structural modification during both stages alleviates the domain gap between synthetic and real noise without any extra computational cost. With 2 pairs for each additional digital gain (in total 6 pairs) and 0.5% iterations, our method achieves superior performance over other calibration-based methods. Xin Jin 0005, Jia-Wen Xiao, Linghao Han, Chunle Guo, Ruixun Zhang, Xialei Liu, Chongyi Li |
ICCV | 1 |
| 2023 | TAN-GFD: generalizing face forgery detection based on texture information and adaptive noise mining
Xin Jin 0005, Liwen Wu, Shaowen Yao 0001 |
Appl. Intell. | 2 |
| 2023 | Adversarial attacks on multi-focus image fusion models
Xin Jin 0005, Xin Jin 0021, Ruxin Wang 0002, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011 |
Comput. Secur. | 1 |
| 2023 | A theoretical analysis of continuous firing condition for pulse-coupled neural networks with its applications
Xin Jin 0005, Pingfan Zhang, Youwei He, Puming Wang, Jingyu Hou 0001, Wei Zhou 0011, Shaowen Yao 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | DBCT-Net:A dual branch hybrid CNN-transformer network for remote sensing image fusion
Quanli Wang, Xin Jin 0005, Liwen Wu, Yunchun Zhang, Wei Zhou 0011 |
Expert Syst. Appl. | 2 |
| 2023 | An unsupervised multi-focus image fusion method based on Transformer and U-NetabstractAbstract This work presents a multi‐focus image fusion method based on Transformer and U‐Net with an unsupervised training fashion. In this work, the authors introduce Transformer into image fusion because it has great ability to capture the global dependencies and low‐frequency features. In image processing, convolutional neural network (CNN) has good performance of detailed feature extraction but a weakness for global feature extraction, and Transformer has limited power in local or detailed information extraction but a strong capacity for global feature extraction. Thus, this work combines the advantages of CNN and Transformer to propose an unsupervised decision map making model for image fusion joint U‐Net. The authors construct a model including feature extraction and feature reconstruction modules which correspond to the encoder and decoder network of U‐Net, respectively. In addition, perceptual loss is introduced on the basis of structural similarity loss because the combination of these two loss functions can achieve better performance with lower training cost. Experiments show that the proposed image fusion method performs better fusion performance compared with the existing methods. Xin Jin 0005, Xiuliang Xi, Xiaoxuan Ren, Jie Yang 0052 |
IET Image Process. | 1 |
| 2023 | Image colorization using deep convolutional auto-encoder with multi-skip connections
Xin Jin 0005, Yide Di, Xing Chu, Qing Duan, Shaowen Yao 0001, Wei Zhou 0011 |
Soft Comput. | 1 |
| 2023 | Type-I Generative Adversarial AttackabstractDeep neural networks are vulnerable to adversarial attacks either by examples with indistinguishable perturbations which produce incorrect predictions, or by examples with noticeable transformations that are still predicted as the original label. The latter case is known as the Type I attack which, however, has achieved limited attention in literature. We advocate that the vulnerability comes from the ambiguous distributions among different classes in the resultant feature space of the model, which is saying that the examples with different appearances may present similar features. Inspired by this, we propose a novel Type I attack method called generative adversarial attack (GAA). Specifically, GAA aims at exploiting the distribution mapping from the source domain of multiple classes to the target domain of a single class by using generative adversarial networks. A novel loss and a U-net architecture with latent modification are elaborated to ensure the stable transformation between the two domains. In this way, the generated adversarial examples have similar appearances with examples of the target domain, yet obtaining the original prediction by the model being attacked. Extensive experiments on multiple benchmarks demonstrate that the proposed method generates adversarial images that are more visually similar to the target images than the competitors, and the state-of-the-art performance is achieved. Shenghong He, Ruxin Wang 0002, Tongliang Liu, Chao Yi, Xin Jin 0005, Renyang Liu 0001, Wei Zhou 0011 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | Multiple Feature Mining Based on Local Correlation and Frequency Information for Face Forgery DetectionabstractAs facial image manipulation techniques developed, deep fake detection attracted extensive attentions. Although researchers have made remarkable progresses in deepfake detection recently, which is still suffering from two limitations: a) current detectors achieve high accuracy in the high-quality videos and images, but it is hard to capture local and subtle artifacts in the low-quality and high-compression media; b) few of deep fake detection methods gain satisfying performance under cross-database scenario, because detector overfit to specific color textures producing by same manipulation algorithm. Inspired the above issues, this paper proposes a novel framework fusing local related features and frequency information to mine the forgery patterns. Firstly, we design multi-feature enhancement module, which amplifies implicit local disc repancies and capture spatial correlation from three shallow feature layers and high-level semantic layer guided by attention maps. Secondly, dual frequency decomposition module is proposed for disassembling high-frequency and low-frequency features, the forgery artifacts are exposed after dual cross attention block processing in the frequency spectrum. Features from the two streams are fused to the classification for the final result. Comprehensive experiments demonstrate the superior performance of our proposed approach in the low-quality benchmark database and cross-dataset sce-nario. Shuai Liu 0009, Xin Jin 0005, Zhenli He, Wei Zhou 0011, Shaowen Yao 0001, Qiannian Wang |
ICTAI | 3 |
| 2022 | Deepfake Detection Using Multiple Feature Fusion
Xin Jin 0005, Yunyun Dong, Shaowen Yao 0001, Wei Zhou 0011 |
IFIP Int. Conf. Digital Forensics | 2 |
| 2022 | DDF-GAN: A Generative Adversarial Network with Dual-Discriminator for Multi-Focus Image FusionabstractMulti-focus image fusion can overcome the issues that optical lens imaging cannot focus multiple targets simultaneously due to the depth of field limitation. In this paper, we propose a generative adversarial network (DDF-GAN) which consists of a generator and two discriminators to directly generate fused images without decision maps and post-processing. In training process, the source all-in-focus image and the fused image generated by generator are used as input to one of the dual discriminators. Meanwhile, the gradient map of the fused image and the source gradient map of the source all-in-focus image are used as input to another discriminator. An adversarial relationship is established to enhance texture details of fused image. In addition, we create a data set and use it as the main training set for the proposed model. Abundant experiments were carried out to verify the availability of our method. Experimental results prove that our method has advantages in subjective visual perception of human and quantitative measurement. Xin Jin 0005, Jie Yang 0052, Xiuliang Xi, Yunyun Dong |
MSN | 2 |
| 2022 | Deep learning for image colorization: Current and future prospects
Shanshan Huang 0001, Xin Jin 0005, Li Liu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | CASR-Net: A color-aware super-resolution network for panchromatic image
Ling Liu 0010, Xin Jin 0005, Jianan Feng, Ruxin Wang 0002, Hangying Liao, Shin-Jye Lee, Shaowen Yao 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | MCRD-Net: An unsupervised dense network with multi-scale convolutional block attention for multi-focus image fusionabstractAbstract Multi‐focus image fusion technology solves the problem of limited depth of field of the optical lens. It can extract different focus parts under the same target to synthesize a full‐focus image. This paper proposes an unsupervised dense network for multi‐focus image fusion. In the network, a multi‐scale feature extraction module is employed to extract the spatial details of source images from different scales, and a convolutional block attention module is used to select the useful deep features, and a residual module is used to effectively optimize the performance of the network. By introducing these three modules, the proposed network can effectively extract the shallow and deep features of the source images. Besides, Gaussian‐based Sum‐Modified‐Laplacian (GSML) is used to calculate the activity level of the feature map to generate a decision map. The performance of the proposed method is analyzed from two aspects: visual quality and objective metrics. Experimental results show that compared with nine image fusion methods, the performance of this algorithm is better. Xin Jin 0005, Shin-Jye Lee, Shaowen Yao 0001 |
IET Image Process. | 2 |
| 2022 | FFR_FD: Effective and fast detection of DeepFakes via feature point defects
Gaojian Wang, Xin Jin 0005, Xiaohui Cui |
Inf. Sci. | 3 |
| 2022 | MC-LCR: Multimodal contrastive classification by locally correlated representations for effective face forgery detection
Gaojian Wang, Xin Jin 0005, Wei Li 0121, Xiaohui Cui |
Knowl. Based Syst. | 3 |
| 2022 | How to Analyze the Neurodynamic Characteristics of Pulse-Coupled Neural Networks? A Theoretical Analysis and Case Study of Intersecting Cortical ModelabstractThe intersecting cortical model (ICM), initially designed for image processing, is a special case of the biologically inspired pulse-coupled neural-network (PCNN) models. Although the ICM has been widely used, few studies concern the internal activities and firing conditions of the neuron, which may lead to an invalid model in the application. Furthermore, the lack of theoretical analysis has led to inappropriate parameter settings and consequent limitations on ICM applications. To address this deficiency, we first study the continuous firing condition of ICM neurons to determine the restrictions that exist between network parameters and the input signal. Second, we investigate the neuron pulse period to understand the neural firing mechanism. Third, we derive the relationship between the continuous firing condition and the neural pulse period, and the relationship can prove the validity of the continuous firing condition and the neural pulse period as well. A solid understanding of the neural firing mechanism is helpful in setting appropriate parameters and in providing a theoretical basis for widespread applications to use the ICM model effectively. Extensive experiments of numerical tests with a common image reveal the rationality of our theoretical results. Xin Jin 0005, Dongming Zhou 0001, Xing Chu, Shaowen Yao 0001, Keqin Li 0001, Wei Zhou 0011 |
IEEE Trans. Cybern. | 1 |
| 2022 | A Deep Multitask Convolutional Neural Network for Remote Sensing Image Super-Resolution and ColorizationabstractRemote sensing data have become increasingly vital in target detection, disaster monitoring, and military surveillance. Abundant pan-sharpening and super-resolution (SR) methods based on deep learning have been proposed and have achieved remarkable performance. However, pan-sharpening requires paired panchromatic (PAN) and multispectral (MS) images, and SR cannot increase the spectral resolution of PAN. Thus, we introduce a computational imaging-based method to recover or produce the incomplete data of single PAN or MS. This work also explores the integration of multiple tasks by a single neural network. We start with SR and colorization, study the feasibility of simultaneously finishing SR colorization, and use a model trained in SR colorization to finish pan-sharpening without MS. A generic neural network, remote sensing image improvement network (RSI-Net), is designed for remote sensing image SR, colorization, simultaneous SR colorization, and pan-sharpening. To verify its performance, RSI-Net is compared with the state-of-the-art SR and colorization methods. Experiments show that RSI-Net can be competitive in visual effects and evaluation indexes, and it performs well at simultaneous SR colorization, and RSI-Net finishes pan-sharpening and only needs to input PAN. Our experiments confirm the effect of integrating multiple tasks. Jianan Feng, Ching-Hsun Tseng, Xin Jin 0005, Ling Liu 0010, Wei Zhou 0011, Shaowen Yao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | CSRDNN: An Integrated Scheme for Single Satellite Image Colorization and Super-Resolution Using Deep Neural NetworksabstractDeep convolutional neural networks have respectively achieved significant success in image super-resolution and colorization. The DNN has a strong capability to generate high quality images. Both colorization and super-resolution (SR) can be regarded as an independent pixel mapping problem, and this work combines these two visual problems into an integrated task. In this work, we propose an end-to-end model for accomplishing single satellite image colorization and SR simultaneously. Our model comprises two phases: features extraction network and recovery network. First, the residual receptive field block structure is introduced in features extraction network to learn better feature representations for image colorization and SR. Residual Receptive Field Block(RRFB) is improved by expanding the receptive field and enhancing the context connection from inception model. Second, the extracted features are transformed to a color high-resolution image by a recovery architecture. In this work, U-net is employed as the key structure of the recovery architecture. Besides, the squeeze-and-excitation blocks and complex residual blocks are incorporated into the proposed model to increase the reconstruction performance. To verify the performance, our method is compared with the state-of-the-art methods of SR and colorization. The experiments show that proposed method can get competitive in visual effect and evaluation index compared with the existing methods. In the end, the panchromatic dataset is also used to validate our model, and a good color high-resolution image can be obtained by giving a gray and low-resolution panchromatic image. Jianan Feng, Xin Jin 0005, Ching-Hsun Tseng, Shin-Jye Lee, Shaowen Yao 0001 |
IJCNN | 3 |
| 2021 | Using Grayscale Frequency Statistic to Detect Manipulated Faces in Wavelet-DomainabstractManipulating facial images results in negative influences on the social association, with deep generative models. Although many detection methods have been proposed, they have either designed sophisticated neural networks that lack enough interpretability, or found defects specific to one manipulation method. To address this issue, we propose a new approach to explore the defects of fake facial images after wavelet transform and call it GFS (Grayscale Frequency Statistics). First, we utilize Haar wavelet transformation to decompose the image into low-frequency approximation, horizontal detail, vertical detail, and diagonal detail. The GFS of real and fake images exhibit different distribution and forms in these four subbands. We qualitatively analyze these differences and quantify them as weights. Then, these four subband images are used to train four CNNs respectively, and the obtained detection results also verify the differences in GFS. After that, we combine the prediction results of the four CNNs and the corresponding weights to further improve the detection performance. We conduct extensive experiments on 11 datasets generated by various facial manipulation methods, and the superior results show the effectiveness of our proposed approach. Our findings indicate that the fake images generated by the current facial manipulation methods cannot simulate real images in wavelet-domain. Gaojian Wang, Wei Li 0121, Xin Jin 0005, Xiaohui Cui |
SMC | 4 |
| 2021 | Color-UNet++: A resolution for colorization of grayscale images using improved UNet++
Yide Di, Xiaoke Zhu, Xin Jin 0005, Qiwei Dou, Wei Zhou 0011, Qing Duan |
Multim. Tools Appl. | 3 |
| 2021 | A fully-automatic image colorization scheme using improved CycleGAN with skip connections
Shanshan Huang 0001, Xin Jin 0005, Jie Li 0023, Shin-Jye Lee, Puming Wang, Shaowen Yao 0001 |
Multim. Tools Appl. | 2 |
| 2021 | Remote sensing image colorization using symmetrical multi-scale DCGAN in YUV color space
Xin Jin 0005, Shin-Jye Lee, Wentao Liang, Shaowen Yao 0001 |
Vis. Comput. | 2 |
| 2020 | New Entropy and Distance Measures of Intuitionistic Fuzzy SetsabstractIn fuzzy set theory, the distance and entropy measure of intuitionistic fuzzy sets (IFSs) have received extensive concern because of the capability for handling imprecise or uncertain problems. However, most of the existing modeling methods for distance and entropy measure are imperfect in teams of intelligibility and performance. In this work, we proposed a new geometric modeling method that can be simultaneously used for distance and fuzzy entropy modeling of IFSs. We used rigorously mathematical derivation to prove that the proposed distance and fuzzy entropy measures satisfy the properties of the definitions. In the experiments, we applied the proposed distance and fuzzy entropy measure into pattern recognition, medical diagnosis, and multi-attribute decision making to examine the usability of the two measures in practical situations. Jinfang Huang, Xin Jin 0005, Dianwu Fang, Shin-Jye Lee, Shaowen Yao 0001 |
FUZZ-IEEE | 2 |
| 2020 | Two-scale decomposition-based multifocus image fusion framework combined with image morphology and fuzzy set theory
Xin Jin 0005, Shin-Jye Lee, Xiaohui Cui, Shaowen Yao 0001, Liwen Wu |
Inf. Sci. | 2 |
| 2019 | A new similarity/distance measure between intuitionistic fuzzy sets based on the transformed isosceles triangles and its applications to pattern recognition
Xin Jin 0005, Shin-Jye Lee, Shaowen Yao 0001 |
Expert Syst. Appl. | 2 |
| 2019 | Multi-focus image fusion combining focus-region-level partition and pulse-coupled neural network
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie, Xin Jin 0005 |
Soft Comput. | 5 |
| 2018 | The Advance of Support Tensor MachineabstractIn recent years, tensor-based machine learning methods, in which the Support Tensor Machine (STM) is a typical technology, have gradually attracted the attention of researchers. Compared with Support Vector Machine (SVM), STM has superior generalization ability that can make full use of the structural information of data. However, it still faces many challenges due to the imperfection of its theoretical basis and model. In order to study the further development of STM, this paper provides a survey about the potential and existing problems in STM. Jing He 0012, Xin Jin 0005, Liwen Wu, Shaowen Yao 0001 |
SERA | 4 |
| 2018 | A lightweight scheme for multi-focus image fusion
Xin Jin 0005, Jingyu Hou 0001, Rencan Nie, Shaowen Yao 0001, Dongming Zhou 0001, Kangjian He |
Multim. Tools Appl. | 1 |
| 2018 | Multimodal sensor medical image fusion based on nonsubsampled shearlet transform and S-PCNNs in HSV space
Xin Jin 0005, Jingyu Hou 0001, Dongming Zhou 0001, Shaowen Yao 0001 |
Signal Process. | 1 |
| 2018 | Multi-focus image fusion method using S-PCNN optimized by particle swarm optimization
Xin Jin 0005, Dongming Zhou 0001, Shaowen Yao 0001, Rencan Nie, Kangjian He |
Soft Comput. | 1 |