EDBT 2026 Demo / reviewers in the wild / expert
Dan Xu 0001
dblp:16/3823-1
· DBLP profile ↗
112ranked-venue papers
0as first author
85since 2021 · last 2026
0000-0003-4602-3550ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 44 since 2021Artificial intelligence and machine learning · 43 · 37 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StrCCL: Structure-aware Contrastive Consistency Loss for Artistic Style Transfer
Shuyu Pan, Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Multimodal progressive contrastive learning for sentiment analysis
Lianmin Zhou, Zhengpeng Zhao, Jue Feng, Dan Xu 0001, Jinjing Gu |
Neurocomputing | 5 |
| 2026 | Enhancing infrared-visible image fusion via text-guided adaptive feature integration
Jundong Zhang, Kangjian He, Dan Xu 0001, Songhan Zheng, Wencheng Mei |
Multim. Syst. | 4 |
| 2026 | CUDiff: Consistency and uncertainty guided conditional diffusion for infrared and visible image fusion
Yueying Luo, Kangjian He, Dan Xu 0001 |
Pattern Recognit. | 3 |
| 2026 | Prior knowledge driven dynamic fusion network for infrared and visible images
Yueying Luo, Kangjian He, Dan Xu 0001, Yiqiao Zhou |
Pattern Recognit. | 3 |
| 2026 | CoMPLe: Cross-Modal Hybrid Prompt Learning for End-to-End Multimodal Emotion RecognitionabstractThe quality of features directly affects the accuracy of Multimodal Emotion Recognition (MER). A key challenge in this context is the effective extraction of dynamically interactive multimodal features to enrich conversational emotion representations. However, existing approaches are often constrained by non-end-to-end architectures, overlooking the significance of feature extraction in MER. To address the problem of dynamic interaction in emotion feature extraction, this paper introduces an end-to-end network based on Cross-Modal Hybrid Prompt Learning (CoMPLe). The model takes raw video as input and leverages three prompt mechanisms to guide large-scale pre-trained encoders in extracting emotionally salient features with latent correlations. Specifically, we design a cross-modal soft prompt learning strategy to mine complementary information across modalities and dynamically adjust the cross-modal semantic space. To capture stage-dependent characteristics, deep feature prompts are incorporated to progressively learn intra-modal contextual representations. Furthermore, a label prompt mechanism is proposed to construct hard prompt templates from emotion labels. Finally, the highest cosine similarity is computed between each unimodal feature and the label prompt templates to activate factual knowledge relevant to emotion recognition. Experiments on three public datasets show that the end-to-end network proposed in this paper surpasses the existing State-Of-The-Art baselines. Jue Feng, Zhengpeng Zhao, Lianmin Zhou, Jiale Ye, Dan Xu 0001, Jinjing Gu |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | Test-Time Domain-Agnostic Meta-Prompt Learning for Multi-Source Few-Shot Domain Adaptation
Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain AdaptationabstractConventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptation (UMFDA) schema. It belongs to a decentralized edge collaborative learning in the edge-side models that must maintain a low computational load. And only a limited amount of annotations in source domain data is provided, with most of the data being unannotated. Further, this paper proposes a vision-aware multimodal prompt tuning framework (VAMP) under the decentralized schema, where the vision-aware prompt guides the text domain-specific prompt to maintain semantic discriminability and perceive the domain information. The cross-modal semantic and domain distribution alignment losses optimize each edge-side model, while text classifier consistency and semantic diversity losses promote collaborative learning among edge-side models. Extensive experiments were conducted on OfficeHome and DomainNet datasets to demonstrate the effectiveness of the proposed VAMP in the UMFDA, which outperformed the previous prompt tuning methods. Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 4 |
| 2025 | Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian PerspectiveabstractCurrent neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods affect pre-trained language models (PLMs) remains unclear. This study revisits the assumption that non-IID information enhances PLMs to achieve performance improvements from a Bayesian perspective, which unearths and integrates non-IID and IID features. Furthermore, we proposed a multi-attribute multi-grained framework for PLM adaptations (M2A), which combines multi-attribute and multi-grained views to mitigate uncertainty in a lightweight manner. We evaluate M2A through prevalent text-understanding datasets and demonstrate its superior performance, mainly when data are implicitly non-IID, and PLMs scale larger. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 4 |
| 2025 | CoFormer: Coupling Attentive Model for Visual Sentiment Analysis with Hierarchical Emotion Loss
Gaifang Luo, Hao Zhang 0110, Zhaoyu Xiong, Dan Xu 0001 |
CGI (1) | 5 |
| 2025 | PCM-Net: A Hierarchical Medical Image Registration Framework Integrating Channel Adaptability and Multi-scale Awareness
Zihang Sun, Dan Xu 0001, Kangjian He, Zilong Xue, Yijie He |
CGI (1) | 2 |
| 2025 | A Brain-Inspired Multimodal Sentiment Analysis Framework via Rationale-Guided Representation
Gaifang Luo, Hao Zhang 0110, Haomin Tan, Zhijing Wu 0011, Dan Xu 0001 |
CogSci | 6 |
| 2025 | 3SmLSTM: Symmetric Feature Mixing MatrixLSTM with Specific Skeletal Connectivity Encoding for Skeleton-Based Action Recognition
Dan Xu 0001 |
ICIC (14) | 4 |
| 2025 | Non-uniform Degradation Aware and Content Complexity Adaptive Optimization for Blind Super-Resolution
Hongzhen Shi, Dan Xu 0001 |
ICIG (3) | 3 |
| 2025 | Construct a Powerful Discriminative Relationship for Few-Shot Action RecognitionabstractLearning discriminative features from very few labeled samples has gradually become a hot issue in the task of human skeleton action recognition. Most of the existing works follow the paradigms of meta-learning or contrastive learning. However, we argue that the discriminative relationship established in this way is rather simple, because the temporal and spatial features are complex and it is difficult to find the relationships of the same category. In this paper, we propose a Multi-Level Semantic Prompting Joint Contrastive Learning Head (MSJCL-Head), which consists of Joint Hard-Soft Contrastive Learning module (JCL) and Action Semantic Prompts (ASP), to obtain the discriminative representations of texts and skeletons with effect strength and discover and calibrate ambiguous samples in the feature space. A large number of experiments have been conducted on the NTU-T, NTU-S and Kinetics datasets, and the results show that our model has achieved competitive results in few-shot tasks. Qianhan Tang, Ningxin Wang, Kangjian He, Hao Zhang 0110, Dan Xu 0001 |
ICME | 6 |
| 2025 | An AI-Enhanced VR Metaverse for Ethnic Festival Culture Protection and Inheritance
Tingyu Zhu, Qianhan Tang, Hao Zhang 0110, Dan Xu 0001 |
ICXR | 8 |
| 2025 | Qwen-Gender: A Chain-of-Thought Based Multi-task Gender Bias Mitigation System
You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
NLPCC (4) | 4 |
| 2025 | FCReg: Medical Image Registration Network with Image-Text Feature Coupling
Yijie He, Dan Xu 0001, Yueying Luo, Zihang Sun, Kangjian He |
PRCV (14) | 2 |
| 2025 | DPA-SAM: Enhancing Medical Image Segmentation with 3D-DCAF and PGAttention
Liye Li, Kangjian He, Gaifang Luo, Hao Zhang 0110, Yijie He, Dan Xu 0001 |
PRCV (14) | 6 |
| 2025 | Bridging the Degradation Gap in Real Super-Resolution: A Transfer-Based Paired Dataset Construction
Yinghui Zhu, Congcong Zeng, Dan Xu 0001, Jiangang Pan, Kangjian He, Hongzhen Shi |
PRCV (9) | 3 |
| 2025 | TSSA-Net: Transposed Sparse Self-Attention-based network for image super-resolution
Guanhao Chen, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | FNContra: Frequency-domain Negative Sample Mining in Contrastive Learning for limited-data image generation
Qiuxia Yang, Zhengpeng Zhao, Shuyu Pan, Jinjing Gu, Dan Xu 0001 |
Expert Syst. Appl. | 6 |
| 2025 | Medical Image Registration via Spatial Feature Extraction Mamba and Substrate Iterative RefinementabstractABSTRACT One of the major challenges in medical image registration is balancing computational efficiency with the ability to capture large deformations in complex anatomical structures. Existing methods often struggle with high computational costs due to the need for extensive feature extraction and attention computations at various levels of the network. Moreover, some methods do not take into account the spatial relationships of the feature images during registration, and the loss of these spatial relationships leads to suboptimal results for these methods. To this end, we introduce a novel medical image registration network, PSMamba‐Net, which leverages optimized iteration and the Mamba framework within a dual‐stream pyramid architecture. The network reduces the computational burden by narrowing attention computations at each decoding level, while an optimized iterative registration module at the bottom of the pyramid captures large deformations. This approach eliminates the need for repeated feature extraction, significantly accelerating the registration process. Additionally, the SMB module is incorporated as a decoder to enhance spatial relationship modelling and leverage Mamba's strengths in long‐sequence processing. PSMamba‐Net balances efficiency and accuracy, surpassing state‐of‐the‐art methods across LPBA40, Mindboggle, and Abdomen CT datasets. Our source code is available at: https://github.com/VCMHE/PSMamba . Zilong Xue, Kangjian He, Dan Xu 0001 |
IET Image Process. | 3 |
| 2025 | Multimodal hypergraph network with contrastive learning for sentiment analysis
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001 |
Neurocomputing | 7 |
| 2025 | Heterogeneous federated distillation with mutual information maximization for medical relation extraction
Jin Wang 0008, Jiaxu Dao, You Zhang 0002, Dan Xu 0001, Xuejie Zhang 0002 |
Inf. Sci. | 4 |
| 2025 | Infrared and visible image fusion based on hybrid multi-scale decomposition and adaptive contrast enhancement
Yueying Luo, Kangjian He, Dan Xu 0001, Hongzhen Shi, Wenxia Yin |
Signal Process. Image Commun. | 3 |
| 2025 | MDH-Net: advancing 3D brain MRI registration with multi-stage transformer and dual-stream feature refinement hybrid network
Chenou Liu, Kangjian He, Dan Xu 0001, Hongzhen Shi |
J. Supercomput. | 3 |
| 2025 | Multi-modality medical image fusion by edge supervising and multi-scale attention features extraction
Wencheng Mei, Kangjian He, Dan Xu 0001, Siqi Xie, Yiqiao Zhou |
J. Supercomput. | 3 |
| 2024 | Personalized LoRA for Human-Centered Text UnderstandingabstractEffectively and efficiently adapting a pre-trained language model (PLM) for human-centered text understanding (HCTU) is challenging since user tokens are million-level in most personalized applications and do not have concrete explicit semantics. A standard and parameter-efficient approach (e.g., LoRA) necessitates memorizing numerous suits of adapters for each user. In this work, we introduce a personalized LoRA (PLoRA) with a plug-and-play (PnP) framework for the HCTU task. PLoRA is effective, parameter-efficient, and dynamically deploying in PLMs. Moreover, a personalized dropout and a mutual information maximizing strategies are adopted and hence the proposed PLoRA can be well adapted to few/zero-shot learning scenarios for the cold-start issue. Experiments conducted on four benchmark datasets show that the proposed method outperforms existing methods in full/few/zero-shot learning scenarios for the HCTU task, even though it has fewer trainable parameters. For reproducibility, the code for this paper is available at: https://github.com/yoyo-yun/PLoRA. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 4 |
| 2024 | Improving Personalized Sentiment Representation with Knowledge-enhanced and Parameter-efficient Layer NormalizationabstractExisting studies on personalized sentiment classification consider a document review as an overall text unit and incorporate backgrounds (i.e., user and product information) to learn sentiment representation. However, it is difficult when these methods meet the current pretrained language models (PLMs) owing to quadratic costs that increase with text length and heterogeneous mixes of randomly initialized background information and textual information initialized from well-pretrained checkpoints during information incorporation. To address these problems, we propose a knowledge-enhanced and parameter-efficient layer normalization (E2LN) for efficient and effective review modeling via leveraging LN in transformer structures. Initially, a knowledge base is introduced that stores well-pretrained checkpoints, structured text information, and background information. Based on such a knowledge base, the ability of LN can be magnified as being a crucial component of transformer structure and then improve the performance of PLMs in downstream tasks. Moreover, the proposed E2LN can make PLMs capable of modeling long document reviews and incorporating background information with parameter-efficient fine-tuning and knowledge injecting. Extensive experimental results were obtained for three document-level sentiment classification benchmark datasets. By comparing the results, the effectiveness and efficiency of the proposed model was demonstrated. Code and Data are released at https://github.com/yoyo-yun/E2LN. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
LREC/COLING | 4 |
| 2024 | Multimodal Sentiment Analysis Based on 3D Stereoscopic AttentionabstractIn the multimodal (text, audio, and visual) sentiment analysis, the current methods generally consider the bi-modal sentiment interaction, resulting in inadequate mining and fusion of relations between modalities. In this paper, we propose the concept of multimodal 3D (3-Dimensional) stereoscopic attention for the first time, which constructs the tri-modal stereoscopic attention with temporal sequences simultaneously to adequately structure the sentiment interaction. To solve the problems of stereoscopic attention construction such as the increased complexity of algorithms caused by rising dimensions, we propose a progressive construction method with 2D attention as an intermediate process. To implement sentiment relations based on stereoscopic attention to integrating modal information sufficiently, a forward propagation mechanism is proposed, which optimizes the representations of each modality with multimodal modulation. The results on two public datasets confirm the superiority of the proposed method in all metrics to the baselines. Dongming Zhou 0001, Zhengpeng Zhao, Dan Xu 0001, Jinde Cao |
ICASSP | 6 |
| 2024 | Learning Defendant-aware Label Representation for Multi-Defendant Charge PredictionabstractAutomatic charge prediction based on deep learning methods is a crucial task in legal judgment prediction, aiming to predict the charges based on the fact description for a criminal case. While existing methods focus on multi-class cases with a single defendant, they fail to account for situations involving multiple defendants and labels, limiting their real-world application. To address these limitations, we propose a multi-defendant charge prediction approach that learns defendant-aware label representations (DLR). To handle complex circumstances for diverse defendants in a case, we extract defendant-specific representation by a machine reading comprehension approach via prompting the defendant’s name. In comparison with traditional text classifications that use discrete one-hot label representations, labels in charge predictions require clear definitions such as textual descriptions for determining the exact classified target. Therefore, we resort to a label description encoder to facilitate the charge predictions via understanding defendant-specific representations. Accordingly, we empower an efficient low-rank adaption module as a feature fuser that incorporates dependent-specific representations into label encoders. The proposed method is evaluated on both multi- and single-dependent-based charge prediction datasets, showing its comparable performances in real-world scenarios. The codes and collected datasets for our study are available at: https://github.com/cy330874054/LDLRMDCP. You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
IJCNN | 4 |
| 2024 | LoRA-Enhanced Language Alignments for Robust Code-Mixed Text RepresentationabstractThe utilization of code-mixed texts allows individuals the opportunity to express sentiments flexibly in international and multilingual contexts. However, the diversity of languages and pragmatic writing styles can lead to semantic shifts at both word and sentence levels, resulting in a degraded comprehension of code-mixed texts by machines. To tackle this issue, we propose a method for language alignments, which aligns both word- and sentence-level semantic representations via a low-rank injection (LoRI) and a data augmentation strategy (DA), dubbed LoRIDA. LoRI integrates linguistic features into textual representations as a feature fusion mechanism. To further bridge the gaps between sentence-level semantics, we augment code-mixed data into individual source languages and apply a knowledge distillation method for joint alignments. We evaluate the performance of the proposed method on four code-mixed sentiment analysis datasets, demonstrating its superiority over existing methods. Our code is publicly available at https://github.com/linsongisgood/LELA. Xuqiao Ran, You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
IJCNN | 4 |
| 2024 | Multimodal Medical Image Registration Using Optimized Phase Consistency Within Joint Frequency-Space Domain
Dan Xu 0001, Kangjian He |
PRCV (5) | 2 |
| 2024 | Hyperspectral Image Super-Resolution Based on Dual-Domain Gated Attention Network
Songhan Zheng, Dan Xu 0001, Kangjian He |
PRCV (13) | 2 |
| 2024 | Dual-path hypernetworks of style and text for one-shot domain adaptation
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Yupan Li, Dan Xu 0001 |
Appl. Intell. | 7 |
| 2024 | LP-BFGS attack: An adversarial attack based on the Hessian with limited pixels
Jiebao Zhang, Wenhua Qian, Jinde Cao, Dan Xu 0001 |
Comput. Secur. | 4 |
| 2024 | FCLFusion: A frequency-aware and collaborative learning for infrared and visible image fusion
Chengchao Wang 0002, Zhengpeng Zhao, Rencan Nie, Jinde Cao, Dan Xu 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | ASFusion: Adaptive visual enhancement and structural patch decomposition for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Dapeng Tao, Chengzhou Li |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | PSANet: Automatic colourisation using position-spatial attention for natural imagesabstractAbstract Due to the richness of natural image semantics, natural image colourisation is a challenging problem. Existing methods often suffer from semantic confusion due to insufficient semantic understanding, resulting in unreasonable colour assignments, especially at the edges of objects. This phenomenon is referred to as colour bleeding. The authors have found that using the self‐attention mechanism benefits the model's understanding and recognition of object semantics. However, this leads to another problem in colourisation, namely dull colour. With this in mind, a Position‐Spatial Attention Network(PSANet) is proposed to address the colour bleeding and the dull colour. Firstly, a novel new attention module called position‐spatial attention module (PSAM) is introduced. Through the proposed PSAM module, the model enhances the semantic understanding of images while solving the dull colour problem caused by self‐attention. Then, in order to further prevent colour bleeding on object boundaries, a gradient‐aware loss is proposed. Lastly, the colour bleeding phenomenon is further improved by the combined effect of gradient‐aware loss and edge‐aware loss. Experimental results show that this method can reduce colour bleeding largely while maintaining good perceptual quality. Peng-Jie Zhu, Qiuxia Yang, Zhengpeng Zhao, Hao Wu 0010, Dan Xu 0001 |
IET Comput. Vis. | 7 |
| 2024 | Multi-stage image inpainting using improved partial convolutionsabstractAbstract In recent years, deep learning models have dramatically influenced image inpainting. However, many existing studies still suffer from over‐smoothed or blurred textures when missing regions are large or contain rich visual details. To restore textures at a fine‐grained level, a multi‐stage inpainting approach is proposed, which applies a series of partial inpainting modules as well as a progressive inpainting module to inpaint missing areas from their boundaries to the centre successively. Some improvements are made on the partial convolutions to reduce artifacts like blurriness, which require a convolution kernel to contain known pixels more than a certain proportion. Towards photorealistic inpainting results, the intermediate outputs from each stage are used to compute the loss. Finally, to facilitate the training process, a multi‐step training is designed that progressively adds inpainting modules to optimize the model. Experiments show that this method outperforms the current excellent techniques on the publicly available datasets: CelebA, Places2 and Paris StreetView. Cheng Li 0021, Dan Xu 0001, Hao Zhang 0110 |
IET Image Process. | 2 |
| 2024 | Dynamic hypergraph convolutional network for multimodal sentiment analysis
Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001 |
Neurocomputing | 7 |
| 2024 | Fidelity based visual compensation and salient information rectification for infrared and visible image fusion
Yueying Luo, Dan Xu 0001, Kangjian He, Hongzhen Shi |
Knowl. Based Syst. | 2 |
| 2024 | Co-space Representation Interaction Network for multimodal sentiment analysis
Zhengpeng Zhao, Dongming Zhou 0001, Dan Xu 0001, Jinde Cao |
Knowl. Based Syst. | 6 |
| 2024 | Affective image recognition with multi-attribute knowledge in deep neural networks
Hao Zhang 0110, Gaifang Luo, Yingying Yue, Kangjian He, Dan Xu 0001 |
Multim. Tools Appl. | 5 |
| 2024 | A multi-weight fusion framework for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110 |
Multim. Tools Appl. | 3 |
| 2024 | Exploring adversarial examples and adversarial robustness of convolutional neural networks by mutual information
Jiebao Zhang, Wenhua Qian, Jinde Cao, Dan Xu 0001 |
Neural Comput. Appl. | 4 |
| 2024 | Layerwised multimodal knowledge distillation for vision-language pretrained model
Jin Wang 0008, Dawei Liao, You Zhang 0002, Dan Xu 0001, Xuejie Zhang 0002 |
Neural Networks | 4 |
| 2024 | Decoupled Knowledge Embedded Graph Convolutional Network for Skeleton-Based Human Action RecognitionabstractSkeleton-based action recognition has broad prospects owing to the fact that skeleton data is more robust to scene noise and camera view changes. Recently, researchers mainly aim to explore deep-learning feature engineering with competitive recognition accuracy for skeleton actions. However, a high-performance recognition network is usually stacked by complex feature extraction modules introducing massive computational costs. In this work, we designed a powerful and universal action knowledge distillation paradigm based on decoupled knowledge distillation for transferring action knowledge from heavy teachers to lightweight students more robustly. We constructed a network architecture space consisting of the shrinking versions of outdated 2s-AGCN and searched for several robust students. On this basis, this paradigm is further developed into a powerful decoupled knowledge embedded graph convolutional network (DKE-GCN), which outperforms the teacher significantly on three public datasets and achieves the state-of-the-art. In addition, a light-DKE-GCN is designed to achieve comparable performance with teacher with 16× less parameters, 26× less FLOPs and 8× FPS. Hao Zhang 0110, Xuejie Zhang 0002, Dan Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | RegFSC-Net: Medical Image Registration via Fourier Transform With Spatial Reorganization and Channel Refinement NetworkabstractMedical image registration is crucial in medical image analysis applications. Recently, U-Net-style networks have been commonly used for unsupervised image registration, predicting dense displacement fields in full-resolution space. However, this process is resource-intensive and time-consuming for high-resolution volumetric image data. To address this challenge, this paper proposes a novel model named RegFSC-Net, which utilizes Fourier transform with spatial reorganization (SR) and channel refinement (CR) network for registration. We embed efficient feature extraction modules SR and CR modules into the encoder, and adopt a parameter-free model to drive the decoder to improve the U-shaped network. Precisely, RegFSC-Net does not directly predict the full-resolution displacement field in space but learns the low-dimensional representation of the displacement field in the bandlimited Fourier domain, which is beneficial in reducing network parameters, memory usage, and computational costs. Experimental results show that RegFSC-Net outperforms various state-of-the-art methods. Specifically, in comparison to the widely recognized Transformer-based method TransMorph, RegFSC-Net utilizes only around 8.2% of its parameters, resulting in a 1.95% higher Dice score and significantly faster inference speeds of 126.67% and 419.99% on GPU and CPU, respectively. Furthermore, we also designed three variants of RegFSC-Net and demonstrated their potential applications in computer-aided diagnosis. Chenou Liu, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110, Kunyuan Zhao |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | MVSFusion: infrared and visible image fusion method for multiple visual scenarios
Chengzhou Li, Kangjian He, Dan Xu 0001, Yueying Luo, Yiqiao Zhou |
Vis. Comput. | 3 |
| 2023 | Skeleton-Based Human Action Recognition via Multi-Knowledge Flow Embedding Hierarchically Decomposed Graph Convolutional Network
Hao Zhang 0110, Shouzheng Sun, Dan Xu 0001 |
CAD/Graphics | 5 |
| 2023 | BIT: Improving Image-text Sentiment Analysis via Learning Bidirectional Image-text InteractionabstractExploring the interaction between image and text has a great strength for image-text sentiment analysis. However, most methods only focus on learning forward interaction in forward image-text features and fail to capture the backward interaction in backward image-text features, which leads to the loss of necessary information embedded in backward interaction. In this paper, Bidirectional Interaction Transformer (BIT) that models both forward and backward image-text interactions is proposed for image-text sentiment analysis. Specifically, we first encode image and text to forward and backward features. Then, these features are fed into Bidirectional Interaction Encoder (BIE) with Forward Interaction and Back Interaction branches to model bidirectional (i.e., forward and backward) image-text interaction. Finally, Two-scale Adaptive Gating Fusion (TAGF) is designed to adaptively fuse the forward and backward interactions learned by BIE. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the proposed model. Xingwang Xiao, Zhengpeng Zhao, Jinjing Gu, Dan Xu 0001 |
IJCNN | 5 |
| 2023 | Siamese conditional generative adversarial network for multi-focus image fusion
Huaguang Li, Wenhua Qian, Rencan Nie, Jinde Cao, Dan Xu 0001 |
Appl. Intell. | 5 |
| 2023 | Generate adversarial examples by adaptive moment iterative fast gradient sign method
Jiebao Zhang, Wenhua Qian, Rencan Nie, Jinde Cao, Dan Xu 0001 |
Appl. Intell. | 5 |
| 2023 | W2GAN: Importance Weight and Wavelet feature guided Image-to-Image translation under limited data
Qiuxia Yang, Zhengpeng Zhao, Dan Xu 0001 |
Comput. Graph. | 4 |
| 2023 | GAGCN: Generative adversarial graph convolutional network for non-homogeneous texture extension synthesisabstractAbstract In the non‐homogeneous texture synthesis task, the overall visual characteristics should be consistent when extending the local patterns of the exemplar. The existing methods mainly focus on the local visual features of patterns but ignore the relative position features that are important for non‐homogeneous texture synthesis. Although these methods have achieved success on homogeneous textures, they cannot perform well on non‐homogeneous textures. Thus, it is desirable to model the dependence between pixels to improve the synthesis performance. To ensure synthesis results from both the local detail structure and the overall structure, this paper proposes a non‐homogeneous texture extended synthesis model (GAGCN) combining the generate adversarial network (GAN) and the graph convolutional network (GCN). The GAN learns the internal distribution of image patches, which makes the synthetic image have rich local details. The GCN learns the latent dependence between pixels according to the statistical characteristics of the image. Based on this, a novel graph similarity loss is proposed. This loss describes the latent spatial differences between the sample image and the generated image, which helps the model to better capture global features. Experiments show that our method outperforms existing methods on non‐homogeneous textures. Shasha Xie, Wenhua Qian, Rencan Nie, Dan Xu 0001, Jinde Cao |
IET Image Process. | 4 |
| 2023 | Low-light image enhancement for infrared and visible image fusionabstractAbstract Infrared and visible image fusion (IVIF) is an essential branch of image fusion, and enhancing the visible image of IVIF can significantly improve the fusion performance. However, many existing low‐light enhancement methods are unsuitable for the visible image enhancement of IVIF. In order to solve this problem, this paper proposes a new visible image enhancement method for IVIF. Firstly, the colour balance and contrast enhancement‐based self‐calibrated illumination estimation (CCSCE) is proposed to improve the input image's brightness, contrast, and colour information. Then, the method based on Mutually Guided Image Filtering (muGIF) is adopted to design a strategy to extract details adaptively from the original visible image, which can keep details without introducing additional noise effectively. Finally, the proposed visible image enhancement technique is used for IVIF tasks. In addition, the proposed method can be used for the visible image enhancement of IVIF and other low‐light images. Experiment results on different public datasets and IVIF demonstrate the authors’ method's superiority from both qualitative and quantitative comparisons. The authors’ code will be publicly available at https://github.com/yiqiao666/low‐light‐enhancement‐for‐IVIF/tree/master . Yiqiao Zhou, Lisiqi Xie, Kangjian He, Dan Xu 0001, Dapeng Tao |
IET Image Process. | 4 |
| 2023 | Soft multimodal style transfer via optimal transport
Jie Li 0023, Liwen Wu, Dan Xu 0001, Shaowen Yao 0001 |
Knowl. Based Syst. | 3 |
| 2023 | Collaborative fine-grained interaction learning for image-text sentiment analysis
Xingwang Xiao, Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001 |
Knowl. Based Syst. | 7 |
| 2023 | Superpixel-based adaptive salient region analysis for infrared and visible image fusion
Chengzhou Li, Kangjian He, Dan Xu 0001, Dapeng Tao, Hongzhen Shi, Wenxia Yin |
Neural Comput. Appl. | 3 |
| 2023 | Image-Text Sentiment Analysis Via Context Guided Adaptive Fine-Tuning Transformer
Xingwang Xiao, Zhengpeng Zhao, Rencan Nie, Dan Xu 0001, Wenhua Qian, Hao Wu 0010 |
Neural Process. Lett. | 5 |
| 2023 | Fidelity-driven Optimization Reconstruction and Details Preserving Guided Fusion for Multi-Modality Medical ImageabstractBy integrating effective features of multi-modality medical images to provide richer information, multi-modality medical image fusion has been substantially used in computer-aided diagnosis applications. However, many existing fusion schemes do not consider how to eliminate the effects of the noise in source medical images and cannot provide enough details and textures for disease diagnosis. To address the problems above, we propose a new fidelity-driven optimization (FDO) reconstruction and details preserving guided-based fusion method for multi-modality medical images. To overcome the influence of noise in multi-modality medical images, a rank coefficient optimization method of low-rank approximation based on weighted mean curvature is proposed to reconstruct multi-modality medical image. Moreover, we propose an iterative detail preserving guided fusion (DPGF) method to integrate more textures and detail information of source multi-modality medical images, while ensuring high signal-to noise ratios. The experimental results show that the proposed method outperforms some of the state-of-the-art fusion methods. Specifically, the extensive experiments prove that our method has high robustness for noisy medical images, which also indicates the application prospects in diagnosis applications. Kangjian He, Xuejie Zhang 0002, Dan Xu 0001, Lisiqi Xie |
IEEE Trans. Multim. | 3 |
| 2023 | Optimal Transport-Based Patch Matching for Image Style TransferabstractState-of-the-art image style transfer methods have achieved impressive results by using neural networks. However, neural style transfer (NST) methods either ignore the local details of the style image by using the global statistics for style modeling or cannot fully use shallow features of neural networks, leading to the synthesized image having fewer details. In this study, we proposed a new patch-based style transfer method that directly operates in the image pixel domain without using any neural networks, achieving fascinating style transfer results with rich image details. The proposed method was derived from classic texture synthesis methods. Most previous methods rely on nearest neighbor search (NNS) for patch matching. However, this greedy strategy cannot guarantee the similarity of patch distributions between the synthesized image and the style image, which limits the expressiveness of textures. We solved this problem by proposing an optimal patch matching algorithm formed on the Optimal Transport (OT) theory, which theoretically guarantees the similarity of the patch distributions and gives a flexible style modeling method. Various qualitative and quantitative experiments demonstrated that the proposed method achieves better synthesized results than state-of-the-art style transfer methods, including NST and classic methods based on texture synthesis. Jie Li 0023, Yong Xiang 0001, Hao Wu 0010, Shaowen Yao 0001, Dan Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Unpaired Artistic Portrait Style Transfer via Asymmetric Double-Stream GANabstractWith the development of image style transfer technologies, portrait style transfer has attracted growing attention in this research community. In this article, we present an asymmetric double-stream generative adversarial network (ADS-GAN) to solve the problems that caused by cartoonization and other style transfer techniques when they are applied to portrait photos, such as facial deformation, contours missing, and stiff lines. By observing the characteristics between source and target images, we propose an edge contour retention (ECR) regularized loss to constrain the local and global contours of generated portrait images to avoid the portrait deformation. In addition, a content-style feature fusion module is introduced for further learning of the target image style, which uses a style attention mechanism to integrate features and embeds style features into content features of portrait photos according to the attention weights. Finally, a guided filter is introduced in content encoder to smooth the textures and specific details of source image, thereby eliminating its negative impact on style transfer. We conducted overall unified optimization training on all components and got an ADS-GAN for unpaired artistic portrait style transfer. Qualitative comparisons and quantitative analyses demonstrate that the proposed method generates superior results than benchmark work in preserving the overall structure and contours of portrait; ablation and parameter study demonstrate the effectiveness of each component in our framework. Fanmin Kong, Ivan Lee 0001, Rencan Nie, Zhengpeng Zhao, Dan Xu 0001, Wenhua Qian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Skeleton-based Human Action Recognition via Large-kernel Attention Graph Convolutional NetworkabstractThe skeleton-based human action recognition has broad application prospects in the field of virtual reality, as skeleton data is more resistant to data noise such as background interference and camera angle changes. Notably, recent works treat the human skeleton as a non-grid representation, e.g., skeleton graph, then learns the spatio-temporal pattern via graph convolution operators. Still, the stacked graph convolution plays a marginal role in modeling long-range dependences that may contain crucial action semantic cues. In this work, we introduce a skeleton large kernel attention operator (SLKA), which can enlarge the receptive field and improve channel adaptability without increasing too much computational burden. Then a spatiotemporal SLKA module (ST-SLKA) is integrated, which can aggregate long-range spatial features and learn long-distance temporal correlations. Further, we have designed a novel skeleton-based action recognition network architecture called the spatiotemporal large-kernel attention graph convolution network (LKA-GCN). In addition, large-movement frames may carry significant action information. This work proposes a joint movement modeling strategy (JMM) to focus on valuable temporal interactions. Ultimately, on the NTU-RGBD 60, NTU-RGBD 120 and Kinetics-Skeleton 400 action datasets, the performance of our LKA-GCN has achieved a state-of-the-art level. Hao Zhang 0110, Kangjian He, Dan Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Cascade connection-based channel attention network for bidirectional medical image registration
Lingxiang Kong, Lisiqi Xie, Dan Xu 0001, Kangjian He |
Vis. Comput. | 4 |
| 2023 | Adaptive low light visual enhancement and high-significant target detection for infrared and visible image fusion
Wenxia Yin, Kangjian He, Dan Xu 0001, Yingying Yue, Yueying Luo |
Vis. Comput. | 3 |
| 2022 | An Optimized Material Point Method for Soil-Water Coupled Simulation
Zhaoyu Xiong, Hao Zhang 0110, Dan Xu 0001 |
CGI | 4 |
| 2022 | Sliced Wasserstein Distance for Neural Style Transfer
Jie Li 0023, Dan Xu 0001, Shaowen Yao 0001 |
Comput. Graph. | 2 |
| 2022 | Multi-view stereo for large-scale scene reconstruction with MRF-based depth inference
Shang Sun, Dan Xu 0001, Hao Wu 0015, Haocong Ying, Yurui Mou |
Comput. Graph. | 2 |
| 2022 | Abstract Painting Synthesis via Decremental optimizationabstractAbstract Existing stroke‐based painting synthesis methods usually fail to achieve good results with limited strokes because these methods use semantically irrelevant metrics to calculate the similarity between the painting and photo domains. Hence, it is hard to see meaningful semantical information from the painting. This paper proposes a painting synthesis method that uses a CLIP (Contrastive‐Language‐Image‐Pretraining) model to build a semantically‐aware metric so that the cross‐domain semantic similarity is explicitly involved. To ensure the convergence of the objective function, we design a new strategy called decremental optimization. Specifically, we define painting as a set of strokes and use a neural renderer to obtain a rasterized painting by optimizing the stroke control parameters through a CLIP‐based loss. The optimization process is initialized with an excessive number of brush strokes, and the number of strokes is then gradually reduced to generate paintings of varying levels of abstraction. Experiments show that our method can obtain vivid paintings, and the results are better than the comparison stroke‐based painting synthesis methods when the number of strokes is limited. Zhengpeng Zhao, Dan Xu 0001, Qiuxia Yang, Ruxin Wang 0002 |
Comput. Graph. Forum | 4 |
| 2022 | Arbitrary style transfer with attentional networks via unbalanced optimal transportabstractAbstract Arbitrary style transfer aims to stylize the content image with the style image. The key problem of style transfer is how to balance the global content structure and the local style patterns. A promising method to solve this problem is the attentional style transfer method, where a learnable embedding of image features enables style patterns to be flexibly recombined with the content image, so local style patterns will be well preserved in the stylized image. However, current attentional style transfer methods cannot well preserve the global content structure. To solve this problem, a novel attentional style transfer network is proposed, that relies on Optimal Transport (OT) for computing the attention map. The proposed OT‐based attention ensures the similarity between global distributions of the synthesized image and its corresponding style image. For the optimal transport computation, a regularized formulation is used, which not only allows an unbalanced optimal transport to address the deviational distributions but also improves the robustness of stylized results. The proposed method finds a well balance between the global content structure and local style patterns. Various experiments are conducted to demonstrate the superiority of the proposed method over state‐of‐the‐art methods. Jie Li 0023, Liwen Wu, Dan Xu 0001, Shaowen Yao 0001 |
IET Image Process. | 3 |
| 2022 | OsaMOT: Occlusion and scale-aware multi-object tracking algorithm for low viewpointabstractAbstract Multi‐object tracking (MOT), which uses the context information of image sequences to locate, maintain identities and generate trajectories of multiple targets in each frame, is key technology in the field of computer vision. To address the problems of occlusion and scale variation in low‐viewpoint MOT, OsaMOT is proposed here. First, according to the global occlusion state of each frame, OsaMOT proposes the adaptive anti‐occlusion feature to enhance the awareness and adaptability for occlusion. At the same time, OsaMOT uses the cascade screening mechanism to reduce the “virtual new target” phenomenon due to the dramatic change in target features caused by scale variation and occlusion. Finally, considering that the occluded templates will affect the tracking performance, OsaMOT proposes an adaptive anti‐noise template update mechanism according to the partial occlusion state of the target, which improves the purity of the template library and further enhances the applicability to occlusion. The experimental results show that OsaMOT can weaken the influence of scale variation, partial occlusion, short‐term full occlusion and long‐term full occlusion in the low‐viewpoint tracking scenes. Most evaluation indexes of OsaMOT under low‐viewpoint tracking scenario are superior to those of some typical algorithms proposed in recent years, and the tracking robustness is improved. Yingying Yue, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
IET Image Process. | 2 |
| 2022 | MAM: A multipath attention mechanism for image recognitionabstractAbstract Attention mechanism has shown excellent performance in many computer vision tasks, while the previous literature may not adequately consider different types of attention mechanisms or is individual elaborate designed for a certain network. In this paper, a general yet effective multipath attention mechanism (MAM) to explore the effect of visual attention for image recognition is proposed. In contrast with other attentions that leverage global pooling, the main advantage is that the MAM considers both the correlation of featuremaps and different scale structural information into account. The backbone representations are enhanced by adding MAM laterally along independent and separate dimensions, channel and spatial. Due to only a simple and unified calculation block is generated, MAM can be flexibly integrated into various CNNs within few parameters and trained together end‐to‐end. Furthermore, the topology structures of attention path arrangement are investigated using different connection schemes. Experimental results on several image recognition datasets show that the model outperforms various existing models. Finally, performance improvement through visualisation is intuitively discussed. The source code for the proposed attention module is publicly available. Hao Zhang 0110, Guoqin Peng, Dan Xu 0001, Hongzhen Shi |
IET Image Process. | 5 |
| 2022 | RGB-D mutual guidance for semi-supervised defocus blur detection
Huaguang Li, Wenhua Qian, Rencan Nie, Jinde Cao, Peng Liu 0056, Dan Xu 0001 |
Knowl. Based Syst. | 6 |
| 2022 | Graph transformer network with temporal kernel attention for skeleton-based action recognitionabstractSkeleton-based human action recognition has caused wide concern, as skeleton data can robustly adapt to dynamic circumstances such as camera view changes and background interference thus allowing recognition methods to focus on robust features. In recent studies, the human body is modeled as a topological graph, and the graph convolution network (GCN) is used to extract features of actions. Although GCN has a strong ability to learn spatial modes, it ignores the varying degrees of higher-order dependencies that are captured by message passing. Moreover, the joints represented by vertices are interdependent, and hence incorporating an attention mechanism to weigh dependencies is beneficial. In this work, we propose a kernel attention adaptive graph transformer network (KA-AGTN), which models the higher-order spatial dependencies between joints by the graph transformer operator based on multihead self-attention. In addition, the Temporal Kernel Attention (TKA) block in KA-AGTN generates a channel-level attention score using temporal features, which can enhance temporal motion correlation. After combining the two-stream framework and adaptive graph strategy, KA-AGTN outperforms the baseline 2s-AGCN by 1.9% and by 1% under X-Sub and X-View on the NTU-RGBD 60 dataset, by 3.2% and 3.1% under X-Sub and X-Set on the NTU-RGBD 120 dataset, and by 2% and 2.3% under Top-1 and Top-5 and achieves the state-of-the-art performance on the Kinetics-Skeleton 400 dataset. Hao Zhang 0110, Dan Xu 0001, Kangjian He |
Knowl. Based Syst. | 3 |
| 2022 | Focus-pixel estimation and optimization for multi-focus image fusionabstractAbstract To integrate the effective information and improve the quality of multi-source images, many spatial or transform domain-based image fusion methods have been proposed in the field of information fusion. The key purpose of multi-focus image fusion is to integrate the focused pixels and remove redundant information of each source image. Theoretically, if the focused pixels and complementary information of different images are detected completely, the fusion image with best quality can be obtained. For this goal, we propose a focus-pixel estimation and optimization based multi-focus image fusion framework in this paper. Because the focused pixels of an image are in the same depth of field (DOF), we propose a multi-scale focus-measure algorithm for the focused pixels matting to integrate the focused region firstly. Then, the boundaries of focused and defocused regions are obtained accurately by the proposed optimizing strategy. And the boundaries are also fused to reduce the influence of insufficient boundary precision. The experimental results demonstrate that the proposed method outperforms some previous typical methods in both objective evaluations and visual perception. Kangjian He, Dan Xu 0001 |
Multim. Tools Appl. | 3 |
| 2022 | Adaptive enhanced infrared and visible image fusion using hybrid decomposition and coupled dictionary
Wenxia Yin, Kangjian He, Dan Xu 0001, Yueying Luo |
Neural Comput. Appl. | 3 |
| 2022 | Learning multi-level representations for affective image recognitionabstractAbstract Images can convey intense affective experiences and affect people on an affective level. With the prevalence of online pictures and videos, evaluating emotions from visual content has attracted considerable attention. Affective image recognition aims to classify the emotions conveyed by digital images automatically. The existing studies using manual features or deep networks mainly focus on low-level visual features or high-level semantic representation without considering all factors. To better understand how deep networks are working for affective recognition tasks, we investigate the convolutional features by visualization them in this work. Our research shows that the hierarchical CNN model mainly relies on deep semantic information while ignoring the shallow visual details, which are essential to evoke emotions. To form a more general and discriminative representation, we propose a multi-level hybrid model that learns and integrates the deep semantics and shallow visual representations for sentiment classification. In addition, this study shows that class imbalance would affect performance as the main category of the affective dataset will overwhelm training and degenerate the deep networks. Therefore, a new loss function is introduced to optimize the deep affective model. Experimental results on several affective image recognition datasets show that our model outperforms various existing studies. The source code is publicly available. Hao Zhang 0110, Dan Xu 0001, Gaifang Luo, Kangjian He |
Neural Comput. Appl. | 2 |
| 2021 | Image Emotion Analysis Based on the Distance Relation of Emotion Categories via Deep Metric Learning
Guoqin Peng, Hao Zhang 0110, Dan Xu 0001 |
CGI | 3 |
| 2021 | Multi-modal image synthesis combining content-style adaptive normalization and attentive normalization
Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian |
Comput. Graph. | 4 |
| 2021 | Adaptive colour restoration and detail retention for image enhancementabstractAbstract Computer vision‐based crowd understanding and analysis technology has been widely used in public safety due to the rapid growth of population and the frequent occurrence of various accidents. Improving imaging quality is the key to improve the performance of crowd analysis, density estimation, target recognition, segmentation, and detection in computer vision tasks. Due to the complex imaging environment such as fog and low illumination, some images taken in outdoor environment often have the problems of colour distortion, lack of details, and the poor imaging quality, which affect the subsequent visual tasks. To improve the imaging quality and visual effect, an adaptive colour restoration and detail retention‐based method is proposed for image enhancement. First, to overcome the problem of colour distortion caused by low illumination and fog, a multi‐channel fusion based adaptive image colour restoration method is proposed. To make the enhancement result more consistent with human observation, the detail retention‐based method is applied to enhance the details. Experimental results demonstrate that the authors' results are effective and outperform the compared methods both in visual and objective evaluations. Kangjian He, Dapeng Tao, Dan Xu 0001 |
IET Image Process. | 3 |
| 2021 | Virtual Try-on Network With Attribute Transformation and Local RenderingabstractA virtual try-on network has gradually become a popular topic in recent years. It aims to transfer images of in-shop clothes onto the image of a target person. Owing to the diversity of clothing attributes, developing an image-based virtual try-on network is a complicated task for computers to perform and requires significant effort. Existing methods are unsatisfactory as they cannot preserve the characteristics of the clothes or the target person's identity well, thereby affecting the perception of the generated images; therefore, further research is required. To address this problem, we propose a novel try-on method that combines attribute transformation and local rendering. First, we employ pixel-level semantic segmentation to identify the try-on area and provide implementation conditions for local rendering. Second, we construct a learnable attribute transformation module to complete the try-on task for different attributes. Third, we use a learnable clothing warping module to fit the pose and figure of the target person well and establish a novel loss function, called modified style loss (M-SL), to handle clothes with rich details. Finally, we adopt a local rendering strategy, using which only renders the clothing area to ensure that the details of the non-target area are not lost. Extensive experiments are performed to test our method. The results demonstrate that our method outperforms other state-of-the-art methods. Jun Xu 0028, Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian |
IEEE Trans. Multim. | 4 |
| 2021 | Contrastive learning for a single historical painting's blind super-resolutionabstractMost of the existing blind super-resolution(SR) methods explicitly estimate the kernel in pixel space, which usually has a large deviation and results in poor SR performance. As a seminal work, DASR learns abstract representations to distinguish various degradations in the feature space, which effectively reduces degradation estimation bias. Therefore, we also employ the feature space to extract degradation representations for an ancient painting. However, most of the blind SR mehods, including DASR, are committed to removing degradations introduced by kernels, downsampling and additive noise. Among them, downsampling degradation is often accompanied by unpleasant artifacts. To address this issue, the paper designs a high-resolution(HR) representation encoder EHR based on contrastive learning to distinguish artifacts introduced by downsampling. Moreover, to optimize the ill-posed nature of blind SR, we propose a contrastive regularization(CR) to minimize the contrastive loss based on VGG-19. With the help of CR, the SR images are pulled closer to the HR images and pushed far away from bicubic LR observations. Benefiting from these improvements, our method consistently achieves higher quantitative performance and better visual quality with more natural textures than state-of-the-art approaches on a specialized painting dataset. Hongzhen Shi, Dan Xu 0001, Kangjian He, Hao Zhang 0110, Yingying Yue |
Vis. Informatics | 2 |
| 2020 | CNN-Based Embroidery Style RenderingabstractNonphotorealistic rendering (NPR) techniques are used to transform real-world images into high-quality aesthetic styles automatically. NPR mainly focuses on transfer hand-painted styles to other content images, and simulates pencil drawing, watercolor painting, sketch painting, Chinese monochromes, calligraphy and, so on. However, digital simulation of Chinese embroidery style has not attracted researcher’s much attention. This study proposes an embroidery style transfer method from a 2D image on the basis of a convolutional neural network (CNN) and evaluates the relevant rendering features. The primary novelty of the rendering technique is that the strokes and needle textures are produced by the CNN and the results can display embroidery styles. The proposed method can not only embody delicate strokes and needle textures but also realize stereoscopic effects to achieve real embroidery features. First, using conditional random fields (CRF), the algorithm segments the target content and the embroidery style images through a semantic segmentation network. Then, the binary mask image is generated to guide the embroidery style transfer for different regions. Next, CNN is used to extract the strokes and texture features from the real embroidery images, and transfer these features to the content images. Finally, the simulating image is generated to show the features of the real embroidery styles. To demonstrate the performance of the proposed method, the simulations are compared with real embroidery artwork and other methods. In addition, the quality evaluation method is used to evaluate the quality of the results. In all the cases, the proposed method is found to achieve needle visual quality of the embroidery styles, thereby laying a foundation for the research and preservation of embroidery works. Wenhua Qian, Jinde Cao, Dan Xu 0001, Rencan Nie |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2019 | Multi-Feature Fusion for Multimodal Attentive Sentiment AnalysisabstractSentiment analysis has been an interesting and challenging task, researchers mostly pay attention to single-modal (image or text) emotion recognition, less attention is paid to joint analysis of multi-modal data. Most existing multi-modal sentiment analysis algorithms combined with attention mechanism focus only on local area of images, ignore the emotional information provided by the global features of the image. Motivated by the research status quo, in this paper, we proposed a novel multi-modal sentiment analysis model, which focuses on local attentive feature also on the global contextual feature from image, then a novel feature fusion mechanism is utilized to fuse features from different modal. In our proposed model, we use a convolutional neural network (CNN) to extract the region maps of images, and use the attention mechanism to acquire attention coefficient, then use a CNN with fewer hidden layers to extract the global feature, a long-short term memory model (LSTM) is utilized to extract textual feature. Finally, a tensor fusion network (TFN) is utilized to fuse all features from different modal. Extensive experiments are conducted on both weakly labeled and manually labeled datasets, and the results demonstrate the superiority of the proposed method. Man A, Dan Xu 0001, Wenhua Qian, Zhengpeng Zhao, Qiuxia Yang |
MMAsia | 3 |
| 2019 | Aesthetic art simulation for embroidery style
Wenhua Qian, Dan Xu 0001, Jinde Cao |
Multim. Tools Appl. | 2 |
| 2019 | Adaptive fast local Laplacian filters and its edge-aware application
Zhenping Qiang, Libo He, Yaqiong Chen, Dan Xu 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Non-Local Texture Optimization With Wasserstein Regularization Under Convolutional Neural NetworkabstractExample-based texture synthesis aims to generate a new texture from an exemplar texture and has long been drawing attention in the fields of computer graphics, computer vision, and image processing. Nevertheless, synthesizing structured textures remains a challenging task. Most previous methods rely on additional guidance channels, which encode the structured features of textures. However, estimating the guidance channel is very difficult, and often fails when a texture has unpronounced features. In this paper, we propose a novel texture synthesis method, based on non-local operators, which captures the long-range structure of a texture without the additional guidance channel. The synthesized texture is generated by minimizing non-local texture energy through an expectation-maximization like optimization algorithm. A statistical constraint based on the Wasserstein distance is also proposed to ensure that the synthesized texture preserves the global statistics of the exemplar texture. Extensive experiments show that the proposed method can stably handle textures with different scale structures. Jie Li 0023, Yong Xiang 0001, Jingyu Hou 0001, Dan Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Region covariance based total variation optimization for structure-texture decomposition
Hao Wu 0015, Dan Xu 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Simulating Chalk Art Style PaintingabstractDifferent kinds of illustrations and artistic imagery can be generated or simulated through the nonphotorealistic rendering (NPR) technique. However, designing and simulating new NPR artistic styles remains extremely challenging. Chalk art style is a very famous artistic work all over the world, and few algorithms have been put forward to illustrate this style. This paper presents a novel NPR technique which generates a chalk art drawing from a 2D photograph automatically. We aim at obtaining a set of lines surface with coarse appearance and generating stroke textures of the real chalk painting. Firstly, the edge of the source image is extracted by difference-of-Gaussian filter method. To simulate chalk painting’s lines, image diffusion and enhancement techniques are proposed to produce coarse and rough lines. Secondly, we developed an improved line integral convolution and dilation operation methods to produce the chalk stroke texture. Finally, the edge image, stroke texture image and color image will be mapped to another background image to generate the chalk art drawing. Experimental results are presented to show the effectiveness of our method in producing the color chalk stylistic illustrations, and the methods can simulate the characters of the real chalk art painting. The proposed method of this paper will enlarge the research and application fields of NPR. Meanwhile, it provides a tool for the user to create chalk art paintings via computers even without painting skill. Wenhua Qian, Dan Xu 0001, Kun Yue |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2017 | Gourd pyrography art simulating based on non-photorealistic rendering
Wenhua Qian, Dan Xu 0001, Kun Yue, Yongjie Shi |
Multim. Tools Appl. | 2 |
| 2015 | A Scores Based Rendering for Yunnan Out-of-Print WoodcutabstractWoodcut is a painting, in which knife and woodblock are used instead of pen and paper. Yunnan out-of-print woodcut is an extension of the traditional woodcut. In this paper, we introduce an automatic image-based approach with three steps. First, image is sliced out by color clustering. Then, scores collected from real woodcuts are placed automatically in segmenting regions. Finally, a virtual printing method combining all regions with their colors is proposed to synthesize multicolor woodcut. This method is simple and efficient, and the simulated result looks good. Jie Li 0023, Dan Xu 0001 |
CAD/Graphics | 2 |
| 2015 | Fast Multi-band Blending Using Run-Length EncodingabstractThis paper presents a fast implementation of multi-band blending for combining a set of registered images into a composite mosaic with no visible seams and minimal texture distortion. We first compute a unique seam image using two-pass nearest distance transform, which is independent on the order of input images and has good scalability. Each individual mask can be extracted from this seam image quickly. To promote execution speed and reduce memory usage in building large area mosaics, the seam image and masks are compressed using run-length encoding, and all the following mask operations are built on run-length encoding scheme. We apply our fast blending system to large scale data sets and present detailed quantitative results compared with Open CV and Enblend to demonstrate the speed and memory improvements. Wenhua Qian, Dan Xu 0001 |
CAD/Graphics | 3 |
| 2014 | Image compositing using dominant patch transformations
Hao Wu 0015, Dan Xu 0001 |
Comput. Graph. | 2 |
| 2013 | Feature Extraction and Analysis for Scientific Understanding of Visual ArtabstractIn this paper, the research of visual art based on the computer and information of paintings, which can be called scientific understanding of visual art, has been brought out. Four features, multi-scale amplitude, non-stationarity of artworks, anisotropy of artworks and correlation of coefficients of the curve let transform between scales, are extracted based on the curve let transform to evaluate the style of different artists. The relations between the style of visual art and these features are also stated, and the similarities of these styles are also qualified by comparing these features. It is apparent that each feature reflects different characteristics of the different school paintings. Yaqun Huang, Dan Xu 0001, Wenhua Qian |
CAD/Graphics | 4 |
| 2011 | A New Background Subtraction Method Using Texture and Color Information
Dan Xu 0001, Mu-Rong Jiang |
ICIC (2) | 3 |
| 2010 | Gradual Sampling and Mutual Information Maximisation for Markerless Motion Capture
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (2) | 5 |
| 2010 | Compressive Evaluation in Human Motion Tracking
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (4) | 5 |
| 2010 | A content-aware image resizing method with prominent object size adjustedabstractA novel method that prominent object size can be controlled during image resizing is proposed in this paper. By a simple parameter adjustment, this new method can change the primary object size according to user preference. To accomplish this, we present a new quad distortion energy criterion by considering both the shape and the size of a quad. Moreover, we improve the single resolution visual attention model based on the rarity of features to a multiresolution saliency model. Then, redefine the significance map as the weighted average of this multi-resolution saliency result and gradient magnitude. Meiling Shi, Guoqin Peng, Dan Xu 0001 |
VRST | 4 |
| 2010 | Automatic eye detection using intensity filtering and K-means clustering
Zhiming Qian, Dan Xu 0001 |
Pattern Recognit. Lett. | 2 |
| 2009 | Interactive local color transfer based on Coupled Map LatticesabstractThe paper presents a new interactive local color transformation method base on GrowCut image segmentation algorithm. The method improves Maslennikova et al's CIM based color transformation algorithm by using coupled map lattices (CML). An intuitive and easy user interface is also supported. User just needs to draft some trails both in reference and target image to make the local color matched. The color transformation method represents some good characteristics auto-repair and segmentation with the usage of GrowCut. Dan Xu 0001 |
CAD/Graphics | 2 |
| 2009 | Automatic Face Recognition Systems Design and Realization
Zhiming Qian, Chaoqun Huang, Dan Xu 0001 |
ISNN (4) | 3 |
| 2007 | Convolution Filter Based Pencil Drawing and Its Implementation on GPU
Dang-en Xie, Dan Xu 0001, Xiaochuan Yang |
APPT | 3 |
| 2007 | An Efficient Approach for Generating Pencil Filter and Its Implementation on GPUabstractTraditional pencil drawing methods have their own drawbacks, such as modeling complexity and higher time-consuming. Thus, they are difficult to be suitable for the real-time rendering. In this paper, we present a new pencil texture generating method based on the pencil filter. This approach can conveniently generate the pencil drawing effect by convoluting the input image with the pencil filter. With the help of accelerating hardware GPU (graphics processing unit), the method accomplished real-time video stylization. Dang-en Xie, Dan Xu 0001 |
CAD/Graphics | 3 |
| 2007 | Automatic and Real-Time Video StylizationabstractIn this paper, we present an automatic, real-time video stylization framework that artistically stylizes video stream by using bilateral filter and soft color quantization. In this framework, firstly, we use a separable implementation of bilateral filter as an adaptive and iterative smoothing operation that selectively simplifies image color, leading to an abstracted look. Secondly, we perform a soft color quantization step on the abstracted video. Successively, some optional approaches are designed to generate different artistic styles. The framework is designed for a GPU-based, real-time implementation. We evaluate the effectiveness of our stylization framework with the experiment results. Dan Xu 0001 |
CAD/Graphics | 2 |
| 2005 | Non-photorealistic Tour into Panorama
Ya-Ping Zhang, Dan Xu 0001 |
ICCSA (3) | 3 |
| 2005 | Authorization Based on Palmprint
Xiaoyong Wei, Dan Xu 0001 |
ICIC (1) | 2 |
| 2000 | Multiresolution Image Morphing in Wavelet DomainabstractThis paper presents a new view synthesis technique using the 2D discrete wavelet-based view morphing. View morphing is completely based on pairwise images without camera calibration and depth information of images. First, the fundamental matrix related to any pair of images is estimated. Then, with the fundamental matrix, the pair of image planes is rectified to be parallel and to have their corresponding points lie on the same scanline, which gives an opportunity to generate new views with linear interpolating techniques. The pre-warped images are then decomposed into a hierarchical structure with the wavelet transform. Corresponding coefficients between two decomposed images are therefore linearly interpolated to form the multiresolution representation of an intermediate view. Quantization techniques can be embedded here to compress the coefficients to reduce the morphing complexity. Finally, when displaying, compressed images are decoded and inverse wavelet transformed. A post-warping procedure is employed to transform the interpolated views to its desired position. Paul Bao, Dan Xu 0001 |
IV | 2 |
| 1999 | Complex wavelet-based image mosaics using edge-preserving visual perception modeling
Paul Bao, Dan Xu 0001 |
Comput. Graph. | 2 |
| 1998 | Panoramic image mosaics via complex wavelet pyramidabstractImage mosaics is one of the scene encoding approaches and is a very active research area in computer vision community and computer graphics community. This paper presents a robust panoramic image mosaicing scheme which employs complex wavelet pyramid and the human visual perceptive thresholding techniques. Complex wavelet transform guarantees not only a global optimal solution, but also scale and translation invariance for image alignment. Applying HVS (Human Visual System) thresholding for wavelet coefficients shows that image can be compressed significantly (15-30.1) while the reconstructed images are visually indistinguishable. The stitched images can be progressively refined on the multiresolution decomposition. This guarantees that the scheme has higher performance than traditional mosaicing techniques. Additionally, the scheme registers images taken directly by hand-held digital camera without knowing camera motion and any intrinsic parameters of camera. The present paper addresses the problems of both image matching and registration automatically and simultaneously. Paul Bao, Dan Xu 0001 |
SMC | 2 |