EDBT 2026 Demo / reviewers in the wild / expert
Wenhua Qian
dblp:50/8387
· DBLP profile ↗
46ranked-venue papers
4as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIPTrack: Time-series information prompt network for RGBT tracking
Kaixiang Yan, Wenhua Qian |
Expert Syst. Appl. | 2 |
| 2026 | Dynamic weighted knowledge distillation for skin lesion diagnosis
Wenhua Qian |
Inf. Process. Manag. | 3 |
| 2026 | Cross-modality masked autoencoder for infrared and visible image fusion
Cong Bi, Wenhua Qian, Qiuhan Shao, Jinde Cao, Xue Wang 0011, Kaixiang Yan |
Pattern Recognit. | 2 |
| 2026 | UWRGBD1k: A large-scale RGBD dataset of underwater object tracking
Kaixiang Yan, Wenhua Qian, Cong Bi, Peng Liu 0056 |
Pattern Recognit. | 2 |
| 2026 | Leveraging Global Context for Improved Image-Text Sentiment AnalysisabstractMultimodal sentiment analysis refers to predicting the emotions conveyed by the joint input of multiple modalities. The primary challenge of this task is to effectively integrate emotional features across different modalities to improve the precision of emotion recognition. For visual–textual sentiment analysis, existing methods typically employ two encoders to extract text and image features separately and enhance analysis performance through cross-modal fusion. However, previous methods often overlook the fact that due to the significant noise in social media data, the correlation between text and images may be weak, and forced alignment could lead to semantic distortion. To address the problem, we propose a context-based alignment method that employs a global modality-aware attention mechanism. This mechanism allows each image or text to reference relevant information from other samples in the entire training set during the alignment process. In addition, to further enhance global alignment, we integrate self-supervised contrastive learning to refine feature representations. By leveraging global semantic information, our approach enhances the robustness of local alignment, effectively reducing alignment bias caused by data noise and improving multimodal fusion performance. Experimental results demonstrate that the proposed approach better understands and integrates information from different modalities in multimodal sentiment analysis, leading to improved accuracy and robustness in multimodal sentiment analysis. Peng Liu 0056, Wenhua Qian |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Visual Object Tracking via Integrating Images of Visible, Thermal, and Depth ModalitiesabstractMulti-modal object tracking leverages the complementary attributes of different modalities, such as visible+ thermal (RGBT) or visible+depth (RGBD) data, to enhance tracking performance. However RGBT or RGBD image can not obtain reliable object features for limitations derived from dense or low-illumination scenarios, respectively. This work explores improving tracking reliability by integrating the advantages and complementarity of visible, depth, and thermal (RGBDT) data. To achieve this goal, a triple-modal object tracking dataset (VDT1k) and a triple-modal object tracking network (TOTrack) are proposed in this work. VDT1k includes 1,325 RGBDT image sequences (totaling 240k RGBDT image pairs) that cover 12 common scenarios of downtown streets. Meanwhile, 7 challenging attributes are annotated with binary labels to evaluate the detailed performance of the tested methods under these extreme scenarios. Considering the specific attributes of different modalities, TOTrack adopts asymmetric feature extraction and multi-stage feature fusion to fuse RGBDT information and complete the object tracking task. The asymmetric feature extraction treats the depth modality as a distance perception information that divides the visible and thermal images into three parts (object, background, and foreground). The multi-stage feature fusion employs different fusion strategies for the early, middle, and late fusion of visible and thermal modalities. Experimental results show that TOTrack outperforms current RGBD/RGBT methods on the test data of VDT1k. TOTrack demonstrates that integrating the complementary information of RGBDT data is a more effective approach to further improve tracking reliability and performance. Realted sources can be accessed at the link. Kaixiang Yan, Wenhua Qian, Jinde Cao, Cong Bi, Rufei Gan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | AGVOT: Visual Object Tracking via Cooperation of Aerial and Ground ViewsabstractVisual object tracking (VOT) is a fundamental task in computer vision that underpins many higher-level tasks. Previous VOT methods have explored various approaches to improve the performance of VOT tasks and achieve predetermined goals. However, due to the limitations of a single view, issues such as visual blind spots and glitches in photographic devices present insurmountable challenges, often causing the visual features of the tracked object to vanish entirely or partially. Providing supplementary images from another view is an effective way to improve tracking performance by eliminating blind spots and enhancing the reliability of image sources. In this work, we propose a novel benchmark called the Aerial and Ground Cooperation Dataset for VOT (AGVOT). AGVOT consists of 643 sequences from aerial and ground views, comprising 119,859 image pairs. These image pairs were collected from 15 scenes and span 26 object categories. Every sequence is strictly time-aligned and annotated with 4-point bounding boxes. For comprehensive evaluation, AGVOT is further categorized into 12 challenging attributes, each labeled at the frame level. Moreover, we introduce an Aerial and Ground Cooperation Tracking Network (AGTrack) to demonstrate dual-view feature integration and cross-view tracking cooperation. Extensive experiments show that AGTrack outperforms 10 state-of-the-art (SOTA) methods in the computer vision community when tested on AGVOT’s dataset and its challenging attributes. These results demonstrate that implementing VOT tasks with cooperative views achieves more accurate and reliable tracking performance. The project source is available at the link. Kaixiang Yan, Wenhua Qian, Jinde Cao, Cong Bi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Gap-KD: Bridging the Significant Capacity Gap Between Teacher and Student Model
Wenhua Qian |
CVM (3) | 2 |
| 2025 | Residual Prior-driven Frequency-aware Network for Image FusionabstractImage fusion aims to integrate complementary information across modalities to generate high-quality fused images, thereby enhancing the performance of high-level vision tasks. While global spatial modeling mechanisms show promising results, constructing long-range feature dependencies in the spatial domain incurs substantial computational costs. Additionally, the absence of ground-truth exacerbates the difficulty of capturing complementary features effectively. To tackle these challenges, we propose a Residual Prior-driven Frequency-aware Network, termed as RPFNet. Specifically, RPFNet employs a dual-branch feature extraction framework: the Residual Prior Module (RPM) extracts modality-specific difference information from residual maps, thereby providing complementary priors for fusion; the Frequency Domain Fusion Module (FDFM) achieves efficient global feature modeling and integration through frequency-domain convolution. Additionally, the Cross Promotion Module (CPM) enhances the synergistic perception of local details and global structures through bidirectional feature interaction. During training, we incorporate an auxiliary decoder and saliency structure loss to strengthen the model's sensitivity to modality-specific differences. Furthermore, a combination of adaptive weight-based frequency contrastive loss and SSIM loss effectively constrains the solution space, facilitating the joint capture of local details and global features while ensuring the retention of complementary information. Extensive experiments validate the fusion performance of RPFNet, which effectively integrates discriminative features, enhances texture details and salient objects, and can effectively facilitate the deployment of the high-level vision task. The source code can be available at https://github.com/wang-x-1997/RPFNet. Xue Wang 0011, Wenhua Qian, Peng Liu 0056, Runzhuo Ma |
ACM Multimedia | 3 |
| 2025 | SPMCGS: Sharpness Prior Enhanced Markov Chain Monte Carlo for Deblurring 3D Gaussian Splatting
Can Luo, Huaguang Li, Qihao Huang, Wenhua Qian |
PRCV (10) | 4 |
| 2025 | GEAST-RF: Geometry Enhanced 3D Arbitrary Style Transfer Via Neural Radiance Fields
Wenhua Qian, Jinde Cao |
Comput. Graph. | 2 |
| 2025 | Knowledge distillation via teacher-modeled sample relationships for skin cancer diagnosis
Peng Liu 0056, Wenhua Qian, Shan Tang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Art creator: Steering styles in diffusion model
Shan Tang, Wenhua Qian, Peng Liu 0056, Jinde Cao |
Neurocomputing | 2 |
| 2025 | VT2V: A benchmark of object tracking with aerial & ground cooperation and multi-modal images
Kaixiang Yan, Wenhua Qian, Cong Bi, Xue Wang 0011 |
Knowl. Based Syst. | 2 |
| 2025 | PID Controller-Driven Network for Image FusionabstractWith its well-designed network architecture, the deep learning-based infrared and visible image fusion (IVIF) method shows its efficiency and effectiveness by realizing a fine feature extraction and fusion mechanism. However, disparities in cross-modal features often result in an imbalance between texture details and contextual information, causing detailed features to be overshadowed by prevailing contextual information. To tackle this issue, this study introduces PIDFusion, a fusion model driven by a PID controller, designed to dynamically optimize cross-modal feature fusion deviations. The core of PIDFusion is the dynamic adaptation capability of the PID controller, which facilitates real-time corrections for deviations encountered during the fusion process, thereby maintaining a harmonious balance between texture details and contextual information. Additionally, we introduced the Cyclic Self-Supervised Feature Refinement (CSSFR), which under the constraint of self-supervised loss, minimizes redundant information within the feature flow and ensures the preservation of salient feature through the cyclic input of decoupled features. Concurrently, we developed the Iterative Attention Module (IAM), utilizing the unique gating mechanism of LSTM to capture feature changes across successive iterations, thereby driving the model to cultivate more discriminative feature representations. Extensive experiments revealed that PIDFusion outperforms SOTA methods in terms of both efficiency and cost-effectiveness, through static statistics and high-level vision tasks. Our code is available athttps://github.com/wang-x-1997/PIDFusion. Xue Wang 0011, Wenhua Qian, Jinde Cao, Runzhuo Ma |
IEEE Trans. Multim. | 3 |
| 2025 | STFuse: Infrared and Visible Image Fusion via Semisupervised Transfer LearningabstractInfrared and visible image fusion (IVIF) aims to obtain an image that contains complementary information about the source images. However, it is challenging to define complementary information between source images in the lack of ground truth and without borrowing prior knowledge. Therefore, we propose a semisupervised transfer learning-based method for IVIF, termed STFuse, which aims to transfer knowledge from an informative source domain to a target domain, thus breaking the above limitations. The critical aspect of our method is to borrow supervised knowledge from the multifocus image fusion (MFIF) task and to filter out task-specific attribute knowledge by using a guidance loss , which motivates its cross-task use in IVIF tasks. Using this cross-task knowledge effectively alleviates the limitation of the lack of ground truth on fusion performance, and the complementary expression ability under the constraint of supervised knowledge is more instructive than prior knowledge. Moreover, we designed a cross-feature enhancement module (CEM) that utilizes self-attention and mutual-attention features to guide each branch to refine features and then facilitate the integration of cross-modal complementary features. Extensive experiments demonstrate that our method has good advantages in terms of visual quality and statistical metrics, as well as the docking of high-level vision tasks, compared with other state-of-the-art methods. Xue Wang 0011, Wenhua Qian, Jinde Cao, Chengchao Wang 0002, Runzhuo Ma |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Overlay Mantle-Free for Semi-supervised Medical Image Segmentation
Wenhua Qian, Jinde Cao, Peng Liu 0056 |
MICCAI (10) | 2 |
| 2024 | EGSRNet: Emotion-Label Guiding and Similarity Reasoning Network for Multimodal Sentiment Analysis
Chunlan Zhan, Wenhua Qian, Peng Liu 0056 |
PRCV (5) | 2 |
| 2024 | LightingFormer: Transformer-CNN hybrid network for low-light image enhancement
Cong Bi, Wenhua Qian, Jinde Cao, Xue Wang 0011 |
Comput. Graph. | 2 |
| 2024 | LP-BFGS attack: An adversarial attack based on the Hessian with limited pixels
Jiebao Zhang, Wenhua Qian, Jinde Cao, Dan Xu 0001 |
Comput. Secur. | 2 |
| 2024 | Exploring adversarial examples and adversarial robustness of convolutional neural networks by mutual information
Jiebao Zhang, Wenhua Qian, Jinde Cao, Dan Xu 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Semi-Supervised Dimensional Media Sentiment Analysis via Exploring Sample RelationshipsabstractDimensional sentiment analysis (DSA) aims to recognize continuous real-valued annotations in multidimensional spaces such as valence-arousal space. It can serve as a sentiment analysis that is more fine-grained than traditional polarity classification. Existing methods primarily focused on supervised learning, which requires a large number of training samples. To address this issue, recent studies have suggested the use of a generative model for semi-supervised DSA. However, these methods lack the exploration of relationships between samples, which may limit their effectiveness. In this study, we proposed a sample relationships (SRs) model for semi-supervised DSA. Compared with the conventional methods, the SR module can integrate the interactive information from different samples, effectively capturing their interdependencies. The results of experiments conducted on three datasets indicated that the performance of the proposed method was better than those of the existing models and that the performance was effective and competitive even with insufficient data. Peng Liu 0056, Wenhua Qian, Huaguang Li, Jinde Cao |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Necessary and Sufficient Conditions for Event-Triggered Set Stabilizability of Markovian Jump Logical NetworksabstractThis technical paper utilizes the Lyapunov theory to characterize the event-triggered set stabilizability of Markovian jump logical control networks (MJLCNs). Whereas the existing result for checking the set stabilizability of MJLCNs is only sufficient, this technical paper further establishes its necessary and sufficient condition. First, the Lyapunov function is established to describe the set stabilizability of MJLCNs necessarily and sufficiently by combining recurrent switching modes and desired state set. Then, the triggering condition and the input updating mechanism are designed regarding the value change of the Lyapunov function. Finally, the effectiveness of theoretical results is demonstrated by a biological example concerning the lac operon in Escherichia coli. Lin Lin 0012, Jinde Cao, Jie Zhong 0005, Yang Liu 0040, Wenhua Qian |
IEEE Trans. Cybern. | 5 |
| 2023 | Siamese conditional generative adversarial network for multi-focus image fusion
Huaguang Li, Wenhua Qian, Rencan Nie, Jinde Cao, Dan Xu 0001 |
Appl. Intell. | 2 |
| 2023 | Generate adversarial examples by adaptive moment iterative fast gradient sign method
Jiebao Zhang, Wenhua Qian, Rencan Nie, Jinde Cao, Dan Xu 0001 |
Appl. Intell. | 2 |
| 2023 | Arbitrary style transfer based on Attention and Covariance-Matching
Haiyuan Peng, Wenhua Qian, Jinde Cao, Shan Tang |
Comput. Graph. | 2 |
| 2023 | GAGCN: Generative adversarial graph convolutional network for non-homogeneous texture extension synthesisabstractAbstract In the non‐homogeneous texture synthesis task, the overall visual characteristics should be consistent when extending the local patterns of the exemplar. The existing methods mainly focus on the local visual features of patterns but ignore the relative position features that are important for non‐homogeneous texture synthesis. Although these methods have achieved success on homogeneous textures, they cannot perform well on non‐homogeneous textures. Thus, it is desirable to model the dependence between pixels to improve the synthesis performance. To ensure synthesis results from both the local detail structure and the overall structure, this paper proposes a non‐homogeneous texture extended synthesis model (GAGCN) combining the generate adversarial network (GAN) and the graph convolutional network (GCN). The GAN learns the internal distribution of image patches, which makes the synthetic image have rich local details. The GCN learns the latent dependence between pixels according to the statistical characteristics of the image. Based on this, a novel graph similarity loss is proposed. This loss describes the latent spatial differences between the sample image and the generated image, which helps the model to better capture global features. Experiments show that our method outperforms existing methods on non‐homogeneous textures. Shasha Xie, Wenhua Qian, Rencan Nie, Dan Xu 0001, Jinde Cao |
IET Image Process. | 2 |
| 2023 | Improving defocus blur detection via adaptive supervision prior-tokens
Huaguang Li, Wenhua Qian, Jinde Cao, Peng Liu 0056 |
Image Vis. Comput. | 2 |
| 2023 | Image-Text Sentiment Analysis Via Context Guided Adaptive Fine-Tuning Transformer
Xingwang Xiao, Zhengpeng Zhao, Rencan Nie, Dan Xu 0001, Wenhua Qian, Hao Wu 0010 |
Neural Process. Lett. | 6 |
| 2023 | Unpaired Artistic Portrait Style Transfer via Asymmetric Double-Stream GANabstractWith the development of image style transfer technologies, portrait style transfer has attracted growing attention in this research community. In this article, we present an asymmetric double-stream generative adversarial network (ADS-GAN) to solve the problems that caused by cartoonization and other style transfer techniques when they are applied to portrait photos, such as facial deformation, contours missing, and stiff lines. By observing the characteristics between source and target images, we propose an edge contour retention (ECR) regularized loss to constrain the local and global contours of generated portrait images to avoid the portrait deformation. In addition, a content-style feature fusion module is introduced for further learning of the target image style, which uses a style attention mechanism to integrate features and embeds style features into content features of portrait photos according to the attention weights. Finally, a guided filter is introduced in content encoder to smooth the textures and specific details of source image, thereby eliminating its negative impact on style transfer. We conducted overall unified optimization training on all components and got an ADS-GAN for unpaired artistic portrait style transfer. Qualitative comparisons and quantitative analyses demonstrate that the proposed method generates superior results than benchmark work in preserving the overall structure and contours of portrait; ablation and parameter study demonstrate the effectiveness of each component in our framework. Fanmin Kong, Ivan Lee 0001, Rencan Nie, Zhengpeng Zhao, Dan Xu 0001, Wenhua Qian |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Non-Fragile H∞ Synchronization for Markov Jump Singularly Perturbed Coupled Neural Networks Subject to Double-Layer Switching RegulationabstractThis work explores the$H_{\infty }$synchronization issue for singularly perturbed coupled neural networks (SPCNNs) affected by both nonlinear constraints and gain uncertainties, in which a novel double-layer switching regulation containing Markov chain and persistent dwell-time switching regulation (PDTSR) is used. The first layer of switching regulation is the Markov chain to characterize the switching stochastic properties of the systems suffering from random component failures and sudden environmental disturbances. Meanwhile, PDTSR, as the second-layer switching regulation, is used to depict the variations in the transition probability of the aforementioned Markov chain. For systems under double-layer switching regulation, the purpose of the addressed issue is to design a mode-dependent synchronization controller for the network with the desired controller gains calculated by solving convex optimization problems. As such, new sufficient conditions are established to ensure that the synchronization error systems are mean-square exponentially stable with a specified level of the$H_{\infty }$performance. Eventually, the solvability and validity of the proposed control scheme are illustrated through a numerical simulation. Hao Shen 0001, Jing Wang 0071, Jinde Cao, Wenhua Qian |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | RGB-D mutual guidance for semi-supervised defocus blur detection
Huaguang Li, Wenhua Qian, Rencan Nie, Jinde Cao, Peng Liu 0056, Dan Xu 0001 |
Knowl. Based Syst. | 2 |
| 2021 | Multi-modal image synthesis combining content-style adaptive normalization and attentive normalization
Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian |
Comput. Graph. | 6 |
| 2021 | Performance analysis of polling-based MAC protocol with retrial for Internet of ThingsabstractSummary We consider and analyze a single‐server multiqueue polling model with inner arrivals. Customers arriving at the queue before polling instant could receive service in the current polling round; furthermore, each one could be retried (turns into an inner arrival) a given number of times with a specified probability. Such polling model can be used to study the performance of certain scheduling data transmission in the Internet of Things (IoT) and the relationship between data retransmission and delay. We obtain the closed‐form expression for the generating function of the amount of customers, which are presented at polling instants. Then, it is used to derive the precise closed‐form formula of mean queue length and mean waiting time in symmetric system. Wenhua Qian, Zhaoxu Zhou |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | H∞ synchronization of delayed neural networks via event-triggered dynamic output control
Yachun Yang, Zhengwen Tu, Liangwei Wang 0002, Jinde Cao, Lei Shi 0003, Wenhua Qian |
Neural Networks | 6 |
| 2021 | Multi-Source Information Exchange Encoding With PCNN for Medical Image FusionabstractMultimodal medical image fusion (MMIF) is to merge multiple images for better imaging quality with preserving different specific features, which could be more informative for efficient clinical diagnosis. In this paper, a novel fusion framework is proposed for multimodal medical images based on multi-source information exchange encoding (MIEE) by using Pulse Coupled Neural Network (PCNN). We construct an MIEE model by using two types of PCNN, such that the information of an image can be exchanged and encoded to another image. Then the fusion contributions for each pixel are estimated qualitatively according to a logical comparison of exchanged information. Further, the exchanged information is nonlinearly transformed using an exponential function with a functional parameter. Finally, quantitative fusion contributions are produced through a reverse-proportional operator to the exchanged information. Also, particle swarm optimization-based derivative-free optimization and a total vibration-based derivative optimization are used to optimize the PCNN and functional transform parameters, respectively. Experiments demonstrate that our method gives the best results than other state-of-the-art fusion approaches. Rencan Nie, Jinde Cao, Dongming Zhou 0001, Wenhua Qian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Virtual Try-on Network With Attribute Transformation and Local RenderingabstractA virtual try-on network has gradually become a popular topic in recent years. It aims to transfer images of in-shop clothes onto the image of a target person. Owing to the diversity of clothing attributes, developing an image-based virtual try-on network is a complicated task for computers to perform and requires significant effort. Existing methods are unsatisfactory as they cannot preserve the characteristics of the clothes or the target person's identity well, thereby affecting the perception of the generated images; therefore, further research is required. To address this problem, we propose a novel try-on method that combines attribute transformation and local rendering. First, we employ pixel-level semantic segmentation to identify the try-on area and provide implementation conditions for local rendering. Second, we construct a learnable attribute transformation module to complete the try-on task for different attributes. Third, we use a learnable clothing warping module to fit the pose and figure of the target person well and establish a novel loss function, called modified style loss (M-SL), to handle clothes with rich details. Finally, we adopt a local rendering strategy, using which only renders the clothing area to ensure that the details of the non-target area are not lost. Extensive experiments are performed to test our method. The results demonstrate that our method outperforms other state-of-the-art methods. Jun Xu 0028, Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian |
IEEE Trans. Multim. | 6 |
| 2020 | CNN-Based Embroidery Style RenderingabstractNonphotorealistic rendering (NPR) techniques are used to transform real-world images into high-quality aesthetic styles automatically. NPR mainly focuses on transfer hand-painted styles to other content images, and simulates pencil drawing, watercolor painting, sketch painting, Chinese monochromes, calligraphy and, so on. However, digital simulation of Chinese embroidery style has not attracted researcher’s much attention. This study proposes an embroidery style transfer method from a 2D image on the basis of a convolutional neural network (CNN) and evaluates the relevant rendering features. The primary novelty of the rendering technique is that the strokes and needle textures are produced by the CNN and the results can display embroidery styles. The proposed method can not only embody delicate strokes and needle textures but also realize stereoscopic effects to achieve real embroidery features. First, using conditional random fields (CRF), the algorithm segments the target content and the embroidery style images through a semantic segmentation network. Then, the binary mask image is generated to guide the embroidery style transfer for different regions. Next, CNN is used to extract the strokes and texture features from the real embroidery images, and transfer these features to the content images. Finally, the simulating image is generated to show the features of the real embroidery styles. To demonstrate the performance of the proposed method, the simulations are compared with real embroidery artwork and other methods. In addition, the quality evaluation method is used to evaluate the quality of the results. In all the cases, the proposed method is found to achieve needle visual quality of the embroidery styles, thereby laying a foundation for the research and preservation of embroidery works. Wenhua Qian, Jinde Cao, Dan Xu 0001, Rencan Nie |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2019 | Multi-Feature Fusion for Multimodal Attentive Sentiment AnalysisabstractSentiment analysis has been an interesting and challenging task, researchers mostly pay attention to single-modal (image or text) emotion recognition, less attention is paid to joint analysis of multi-modal data. Most existing multi-modal sentiment analysis algorithms combined with attention mechanism focus only on local area of images, ignore the emotional information provided by the global features of the image. Motivated by the research status quo, in this paper, we proposed a novel multi-modal sentiment analysis model, which focuses on local attentive feature also on the global contextual feature from image, then a novel feature fusion mechanism is utilized to fuse features from different modal. In our proposed model, we use a convolutional neural network (CNN) to extract the region maps of images, and use the attention mechanism to acquire attention coefficient, then use a CNN with fewer hidden layers to extract the global feature, a long-short term memory model (LSTM) is utilized to extract textual feature. Finally, a tensor fusion network (TFN) is utilized to fuse all features from different modal. Extensive experiments are conducted on both weakly labeled and manually labeled datasets, and the results demonstrate the superiority of the proposed method. Man A, Dan Xu 0001, Wenhua Qian, Zhengpeng Zhao, Qiuxia Yang |
MMAsia | 4 |
| 2019 | Aesthetic art simulation for embroidery style
Wenhua Qian, Dan Xu 0001, Jinde Cao |
Multim. Tools Appl. | 1 |
| 2018 | Fully Convolutional Network-Based Multifocus Image FusionabstractAs the optical lenses for cameras always have limited depth of field, the captured images with the same scene are not all in focus. Multifocus image fusion is an efficient technology that can synthesize an all-in-focus image using several partially focused images. Previous methods have accomplished the fusion task in spatial or transform domains. However, fusion rules are always a problem in most methods. In this letter, from the aspect of focus region detection, we propose a novel multifocus image fusion method based on a fully convolutional network (FCN) learned from synthesized multifocus images. The primary novelty of this method is that the pixel-wise focus regions are detected through a learning FCN, and the entire image, not just the image patches, are exploited to train the FCN. First, we synthesize 4500 pairs of multifocus images by repeatedly using a gaussian filter for each image from PASCAL VOC 2012, to train the FCN. After that, a pair of source images is fed into the trained FCN, and two score maps indicating the focus property are generated. Next, an inversed score map is averaged with another score map to produce an aggregative score map, which take full advantage of focus probabilities in two score maps. We implement the fully connected conditional random field (CRF) on the aggregative score map to accomplish and refine a binary decision map for the fusion task. Finally, we exploit the weighted strategy based on the refined decision map to produce the fused image. To demonstrate the performance of the proposed method, we compare its fused results with several start-of-the-art methods not only on a gray data set but also on a color data set. Experimental results show that the proposed method can achieve superior fusion performance in both human visual quality and objective assessment. Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Wenhua Qian |
Neural Comput. | 5 |
| 2017 | Simulating Chalk Art Style PaintingabstractDifferent kinds of illustrations and artistic imagery can be generated or simulated through the nonphotorealistic rendering (NPR) technique. However, designing and simulating new NPR artistic styles remains extremely challenging. Chalk art style is a very famous artistic work all over the world, and few algorithms have been put forward to illustrate this style. This paper presents a novel NPR technique which generates a chalk art drawing from a 2D photograph automatically. We aim at obtaining a set of lines surface with coarse appearance and generating stroke textures of the real chalk painting. Firstly, the edge of the source image is extracted by difference-of-Gaussian filter method. To simulate chalk painting’s lines, image diffusion and enhancement techniques are proposed to produce coarse and rough lines. Secondly, we developed an improved line integral convolution and dilation operation methods to produce the chalk stroke texture. Finally, the edge image, stroke texture image and color image will be mapped to another background image to generate the chalk art drawing. Experimental results are presented to show the effectiveness of our method in producing the color chalk stylistic illustrations, and the methods can simulate the characters of the real chalk art painting. The proposed method of this paper will enlarge the research and application fields of NPR. Meanwhile, it provides a tool for the user to create chalk art paintings via computers even without painting skill. Wenhua Qian, Dan Xu 0001, Kun Yue |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2017 | Gourd pyrography art simulating based on non-photorealistic rendering
Wenhua Qian, Dan Xu 0001, Kun Yue, Yongjie Shi |
Multim. Tools Appl. | 1 |
| 2015 | Fast Multi-band Blending Using Run-Length EncodingabstractThis paper presents a fast implementation of multi-band blending for combining a set of registered images into a composite mosaic with no visible seams and minimal texture distortion. We first compute a unique seam image using two-pass nearest distance transform, which is independent on the order of input images and has good scalability. Each individual mask can be extracted from this seam image quickly. To promote execution speed and reduce memory usage in building large area mosaics, the seam image and masks are compressed using run-length encoding, and all the following mask operations are built on run-length encoding scheme. We apply our fast blending system to large scale data sets and present detailed quantitative results compared with Open CV and Enblend to demonstrate the speed and memory improvements. Wenhua Qian, Dan Xu 0001 |
CAD/Graphics | 2 |
| 2015 | Qualitative probabilistic network-based fusion of time-series uncertain knowledge
Kun Yue, Wenhua Qian, Xiaodong Fu, Jin Li 0007 |
Soft Comput. | 2 |
| 2013 | Feature Extraction and Analysis for Scientific Understanding of Visual ArtabstractIn this paper, the research of visual art based on the computer and information of paintings, which can be called scientific understanding of visual art, has been brought out. Four features, multi-scale amplitude, non-stationarity of artworks, anisotropy of artworks and correlation of coefficients of the curve let transform between scales, are extracted based on the curve let transform to evaluate the style of different artists. The relations between the style of visual art and these features are also stated, and the similarities of these styles are also qualified by comparing these features. It is apparent that each feature reflects different characteristics of the different school paintings. Yaqun Huang, Dan Xu 0001, Wenhua Qian |
CAD/Graphics | 5 |