VLDB 2026 Research / reviewers in the wild / expert
Aiwen Jiang
dblp:52/1645
· DBLP profile ↗
44ranked-venue papers
6as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CG-SAM2: Confidence-Guided Pseudo-label Refinement for Weakly Supervised Camouflaged Object Detection
Shengmin Zhao, Aiwen Jiang |
ICIC (19) | 5 |
| 2026 | Codebook Knowledge with Mamba-Transformer For Low-Light Image EnhancementabstractLow-light image enhancement is a critical task which aims to improve the quality of images captured in bad lighting conditions, and contribute to more robust and reliable computer vision systems. Existing methods failed to account for multifaceted and intertwined degradations typically encountered in low-light scenarios. In this paper, we have reconsider the application of vector-quantized codebook in low-light image enhancement task as a domain adaptation paradigm and proposed an effective method called CodeMTNet to solve aforementioned issues. Specifically, we leverage codebook learning from collections of norm-light images to provide unified high-quality knowledge guidance. We have further developed two learning schemes, namely Domain Adaptation Encoder with implicit neural representation regularization across multiple scales, and hybrid Mamba-Transformer blocks for nearest neighbor matching, to tackle distribution mismatch between features of low-quality low-light images and high-quality normal-light images. Additionally, to solve structural information loss during codebook retrieval, we have introduced a controllable feature fusion modules for well texture detail preservation. Experiments conducted on public datasets have demonstrated that CodeMTNet consistently outperforms many state-of-the-art methods and restore images better in line with human perception. Related source codes and pretrained parameters are in https://github.com/YunHDan/CodeMTNet.git. Runhua Deng, Aiwen Jiang, Long Peng 0003, Qiuhai Yan |
WACV | 2 |
| 2026 | FARF-Net: Frequency-guided Adaptive Receptive Field Network for Edge-enhanced Polyp SegmentationabstractAccurate segmentation of colorectal polyps plays a vital role in the early diagnosis and prevention of colorectal cancer. Despite notable progress, existing methods struggle with limited region adaptability due to fixed receptive fields, lack explicit boundary modeling, and are prone to interference from background noise, leading to suboptimal segmentation results. To address these issues, we propose FARF-Net, a novel edge-aware segmentation framework that leverages frequency-domain adaptive receptive fields. Built upon the Pyramid Vision Transformer v2 backbone, FARF-Net introduces three tailored components. Specifically, EdgeKAN module applies Kolmogorov–Arnold Networks for channel-wise nonlinear modeling, enhancing local edge semantics and boundary detail representation. Adaptive Receptive Field module adjusts spatial receptive fields based on localized frequency energy, boosting sensitivity to high-frequency boundaries. Frequency-Guided Dual-Supervision Decoder integrates high-frequency structural features and boundary priors to refine edge predictions and suppress irrelevant high-frequency background noise. Extensive experiments on five public polyp segmentation benchmarks demonstrate that FARF-Net consistently surpasses state-of-the-art methods. Notably, it achieves superior boundary reconstruction and robustness in challenging cases such as blurred contours and small polyps. Aiwen Jiang, Hongqian Yu |
WACV | 2 |
| 2026 | An industrial informatics-oriented multi-scale convolutional Mamba with multi-frequency attention for robust medical image segmentation
Yugen Yi, Wei Zhou 0003, Qiangqiang Zhou, Aiwen Jiang, Naixue Xiong, Yingkui Du, Xiaomei Huang |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | DaLPSR: Leverage degradation-aligned language prompt for real-world image super-resolution
Aiwen Jiang, Zhi Wei 0004, Long Peng 0003, Feiqiang Liu, Mingwen Wang 0001 |
Image Vis. Comput. | 1 |
| 2026 | Dual-stage network combining transformer and hybrid convolutions for stereo image super-resolution
Jintao Zeng, Aiwen Jiang, Feiqiang Liu |
Image Vis. Comput. | 2 |
| 2026 | Relational reasoning image captioning via multi-agent retrieval-augmented generation
Aiwen Jiang, Duan Wang |
Knowl. Based Syst. | 1 |
| 2026 | PV Hosting Capacity Assessment of Storage-Integrated Distribution Systems With Constrained PV Curtailment and Load Shedding Under Fault ConditionsabstractIn recent years, distributed photovoltaic (PV) systems have experienced accelerated deployment in distribution networks, driven by inherent advantages. However, high-penetration PV integration introduces critical operational challenges, including voltage violations and reverse power flows, which become progressively more pronounced as penetration levels increase and may ultimately harm user interests. To comprehensively assess the network's PV hosting capacity (PVHC) while safeguarding the interests of both PV owners and load consumers, this article proposes a novel stochastic biscenario (prefault and postfault) assessment framework, which explicitly quantifies and constrains postfault PV curtailment and load shedding, mediated by battery energy storage systems. We derive an analytical linearized expression for the nonlinear PV curtailment under all N-1 contingencies, transforming the original PVHC assessment model into a tractable mixed-integer linear programming model. Furthermore, our analysis reveals that, an inherent tradeoff exists between reducing fault-induced PV curtailment and enhancing load reliability in storage-integrated grids due to state of charge constraints, which holds significant guiding importance for the practical industrial application of PV integration. Finally, the effectiveness, and scalability of the proposed method are verified through different test systems. Yi Wang 0057, Shida Zhang, Aiwen Jiang, Yaoqiang Wang, Venkata Dinavahi, Jun Liang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | DA-Mamba: A Dual-domain Attention Mamba-based Image Deraining NetworkabstractCNN and Transformer-based methods have achieved remarkable performance in image deraining. However, the former often struggles to effectively capture long-range spatial dependencies due to limited receptive fields, while the latter faces quadratic growth in computational complexity as the number of tokens increases. Recently, Mamba-based image restoration methods have garnered significant attention. Despite their potential, these methods face several challenges. Existing models convert 2D images into 1D sequences, disrupting inherent local features. Moreover, as models process 1D pixel sequences, interactions between distant tokens weaken, affecting Mamba’s ability to capture global information. In this work, we propose an efficient DA-Mamba model. Specifically, we introduce a Dual-domain Attention Coupling Module (DACM). First, we use Fast Fourier Transform to extract frequency features, aiding the model in comprehensively understanding image degradation through global modeling. Then, we employ the Sobel operator to obtain gradient features, modeling local detail features of the image. Finally, we use an attention mechanism to weight and fuse information from different domains, enhancing the model’s sensitivity to local structures and improving long-range information extraction. Benchmark experiments on synthetic and real datasets indicate that the proposed method achieves superior performance compared to state-of-the-art methods. Qiuhai Yan, Aiwen Jiang |
IJCNN | 3 |
| 2025 | Enhancing Low-Light Object Detection with Zero-Shot Dual-Branch Illumination-Invariant NetworkabstractObject detection in low-light environments remains challenging due to severe image degradation, limited annotations, and the high cost of manual labeling. While traditional enhancement methods improve visual appearance, they often hinder detection performance. In this work, we propose a lightweight zero-shot enhancement network designed specifically for object detection, enabling effective illumination correction without requiring real low-light data. Unlike pixel-level restoration approaches, our method operates at the feature level, leveraging a physics-inspired model that extracts illumination-invariant representations via Lambertian reflectance and cross-channel chrominance ratios. To enhance global feature perception with minimal computational overhead, we introduce a frequency-domain branch based on complex convolutions, and fuse it adaptively with spatial-domain features through a dual-branch architecture. The proposed module is compact and detector-agnostic, and can be seamlessly integrated into existing frameworks. Extensive experiments show that our approach significantly improves detection performance under low-light conditions, achieving a 4.6% mAP gain on ExDark and a 2.7% improvement on DarkFace. Aiwen Jiang, Jiatian Miao |
MMAsia | 2 |
| 2025 | PB-SAM: A Point-Based Framework for Weakly Supervised Camouflaged Object Detection
Shengmin Zhao, Yiran Nie, Aiwen Jiang |
PRCV (17) | 5 |
| 2025 | Dual-Branch Cross-Scale Texture Feature Fusion for Low-Resolution Face RecognitionabstractABSTRACT Face images captured often suffer from low resolution and significant information loss. Traditional methods struggle to effectively extract local key features, leading to suboptimal recognition accuracy. To address these challenges, this paper introduces a novel approach based on dual‐branch cross‐scale texture feature fusion for low‐resolution face recognition (DCSF‐LR). The proposed method enhances the focus on facial details through local texture feature fusion and a dual‐branch cross‐scale attention module, enabling the extraction of richer facial features. Additionally, knowledge distillation is utilized to transfer knowledge from high‐resolution face images to the low‐resolution face recognition model. A newly designed loss function is introduced to facilitate effective knowledge transfer, better adapting the model to low‐resolution face recognition tasks in uncontrolled environments. Moreover, a degradation module is developed to generate realistic low‐resolution face images for training the student model, thereby improving its adaptability in real‐world scenarios. Extensive experiments on the TinyFace and AgeDB‐30 data sets demonstrate the effectiveness of the proposed method. It achieves 90.04% accuracy at resolution on AgeDB‐30 and 57.73% (ACC@5) on TinyFace, surpassing existing methods in both accuracy and generalization. Jihua Ye, Wentao Geng, Youcai Zou, Aiwen Jiang |
Concurr. Comput. Pract. Exp. | 7 |
| 2025 | Textual prompt guided image restoration
Qiuhai Yan, Aiwen Jiang, Long Peng 0003, Qiaosi Yi |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | CNLA: Collaborative noisy label adaptive learning for facial expression recognition
Jihua Ye, Dong Liu 0060, Huiyuan Huang, Liang Ying, Aiwen Jiang |
Inf. Sci. | 7 |
| 2025 | ConMSDMamba: Multi-Scale Dilated Mamba Based on Conformer for Speech Emotion RecognitionabstractAlthough the Conformer model excels in speech processing, its core self-attention mechanism is limited in capturing multi-scale temporal dynamics and lacks explicit modeling of frequency-domain features, both crucial for Speech Emotion Recognition (SER). To address this, we propose ConMSDMamba, a novel Conformer-based architecture for SER. Specifically, to overcome the single-scale limitation of the original self-attention, we introduce a multi-scale dilated structure with parallel dilated convolutions to capture diverse temporal contexts. We further find that combining this structure with bidirectional Mamba models long-range temporal dependencies more efficiently than multi-head self-attention. Furthermore, to complement the Conformer's time-domain focus, we design a time-frequency convolution module that incorporates a wavelet-based branch for joint time-frequency perception. Experimental results on the widely used IEMOCAP and MELD datasets demonstrate that ConMSDMamba outperforms state-of-the-art methods. Guangyuan Qian, Zhenchun Lei, Sihong Liu, Changhong Liu, Aiwen Jiang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Open-ended Autoregressive Visual Storytelling via Parameter Efficient Instruction TuningabstractVisual storytelling (VIST) involves generating coherent, creative, and vivid narrative for a collection of images. It remains an immense challenge within cross-modal domain. Traditional mainstream storytelling work were less proficient in handling long sequential relationships. Though large-scale visual-language pre-training (VLP) models demonstrated promising prospect on cross-modal tasks. So far they still were not particularly adept at handling tasks involving image sequences. Moreover, the reference descriptions in the available VIST benchmark dataset are short and simplistic, which constrains model’s potential capabilities. Current models struggle to produce truly rich and vivid narratives. Therefore, in this article, we will address these deficiencies, and contribute from both dataset and innovative model aspects. Firstly, by leveraging large language model (LLM), we have constructed a new dataset VIST++ which can enrich vivid narratives for open-ended image sequences. The dataset has potential on providing beneficial support for future model learning. Secondly, we have proposed an innovative auto-regressive story generation model named ReStoryGen. It can be applied to image sequences of varying lengths in an open-ended way. We have performed extensive experiments and evaluations in terms of visual grounding, coherence and non-redundancy. The experiment results have convincingly demonstrated ReStoryGen achieves impressive outcomes through utilizing parameter-efficient instruction-tuning. Related source codes and models are distributed on Github https://github.com/lixinliu1995/story_gen . Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Dual-Path Coupled Image Deraining Network Via Spatial-Frequency InteractionabstractTransformers have recently emerged as a significant force in the field of image deraining. Existing image deraining methods utilize extensive research on self-attention. Though showcasing impressive results, they tend to neglect critical frequency information, as self-attention is generally less adept at capturing high-frequency details. To overcome this shortcoming, we have developed an innovative Dual-Path Coupled Deraining Network (DPCNet) that integrates information from both spatial and frequency domains through Spatial Feature Extraction Block (SFEBlock) and Frequency Feature Extraction Block (FFEBlock). We have further introduced an effective Adaptive Fusion Module (AFM) for the dual-path feature aggregation. Extensive experiments on six public deraining benchmarks and downstream vision tasks have demonstrated that our proposed method not only outperforms the existing state-of-the-art deraining method but also achieves visually pleasuring results with excellent robustness on downstream vision tasks. The source code is available at https://github.com/Madeline-hyh/DPCNet. Aiwen Jiang, Lingfang Jiang, Long Peng 0003, Zhifeng Wang 0006, Lu Wang 0001 |
ICIP | 2 |
| 2024 | Music-driven Character Dance Video Generation based on Pre-trained Diffusion ModelabstractLarge-scale pre-trained models have shown significant progress in cross-modal generation tasks, especially in the text-to-image generation task. However, the pre-trained models for audio-guided video are rare. ControlNet [1] provides a new architecture to enhance the pre-trained diffusion models with task-specific conditions. Following the ControlNet [1] architecture, we propose a music-driven character dance video generation model based on the pre-trained diffusion model by taking the text prompt, music, and character image as the additional guidance conditions to generate dance videos. In this model, multimodal semantic correspondence between text, music, and video is exploited to generate character dance videos better by incorporating the pre-trained CLIP [2] and Wav2CLIP [3] models. Additionally, we design a text prompt to improve the appearance quality of the generated character images. Extensive experiments on the AIST++ dataset show the effectiveness of our method and its ability to generate character dance videos effectively. Changhong Liu, Juan Cai, Ji Ye, Zhenchun Lei, Aiwen Jiang |
IJCNN | 6 |
| 2024 | Latent Diffusion-based Data Augmentation for Continuous-Time Dynamic Graph ModelabstractContinuous-Time Dynamic Graph (CTDG) precisely models evolving real-world relationships, drawing heightened interest in dynamic graph learning across academia and industry. However, existing CTDG models encounter challenges stemming from noise and limited historical data. Graph Data Augmentation (GDA) emerges as a critical solution, yet current approaches primarily focus on static graphs and struggle to effectively address the dynamics inherent in CTDGs. Moreover, these methods often demand substantial domain expertise for parameter tuning and lack theoretical guarantees for augmentation efficacy. To address these issues, we propose Conda, a novel latent diffusion-based GDA method tailored for CTDGs. Conda features a sandwich-like architecture, incorporating a Variational Auto-Encoder (VAE) and a conditional diffusion model, aimed at generating enhanced historical neighbor embeddings for target nodes. Unlike conventional diffusion models trained on entire graphs via pre-training, Conda requires historical neighbor sequence embeddings of target nodes for training, thus facilitating more targeted augmentation. We integrate Conda into the CTDG model and adopt an alternating training strategy to optimize performance. Extensive experimentation across six widely used real-world datasets showcases the consistent performance improvement of our approach, particularly in scenarios with limited historical data. Yuxing Tian, Aiwen Jiang, Jian Guo 0016, Yiyan Qi |
KDD | 2 |
| 2024 | Low-Light Image Enhancement via FourierTMamba: A Hybrid Frequency-Spatial Approach
Shuwei Peng, Xu Zhang 0079, Aiwen Jiang, Changhong Liu, Jihua Ye |
MMAsia | 3 |
| 2024 | MMIDM: Generating 3D Gesture from Multimodal Inputs with Diffusion Models
Ji Ye, Changhong Liu, Haocong Wan, Aiwen Jiang, Zhenchun Lei |
PRCV (6) | 4 |
| 2024 | A multi-scale channel attention network with federated learning for magnetic resonance image super-resolution
Feiqiang Liu, Aiwen Jiang |
Multim. Syst. | 2 |
| 2024 | Physical-prior-guided single image dehazing network via unpaired contrastive learning
Mawei Wu, Aiwen Jiang, Hourong Chen, Jihua Ye |
Multim. Syst. | 2 |
| 2022 | Multi-Scale Cascaded Generator for Music-driven Dance SynthesisabstractDance is a creative performance art and must keep coherent with the rhythm and style of music. To address these issues, most of the existing music-driven dance synthesis methods utilize deep generative models and capture the dynamic characteristics of dance motions. However, we observe that dance motions contain big-scale body part movements and small-scale joint movements that are mutually coordinated and related, and the generated dance motions, particularly affected by$L_{1}$loss, are too restrictive and conservative. In this paper, we propose a multi-scale cascaded music-driven dance synthesis network (MC-MDSN) that first generates big-scale body motions conditioned on music and then further refines local small-scale joint motions. Furthermore, we design a multi-scale feature loss to capture the dynamic characteristics of each scale motions and the relations between different scale motion joints. Experimental results show that our method generates better dance motions than the baselines. Changhong Liu, Aiwen Jiang, Zhenchun Lei, Mingwen Wang 0001 |
IJCNN | 4 |
| 2022 | Learning Hierarchical Semantic Correspondences for Cross-Modal Image-Text RetrievalabstractCross-modal image-text retrieval is a fundamental task in information retrieval. The key to this task is to address both heterogeneity and cross-modal semantic correlation between data of different modalities. Fine-grained matching methods can nicely model local semantic correlations between image and text but face two challenges. First, images may contain redundant information while text sentences often contain words without semantic meaning. Such redundancy interferes with the local matching between textual words and image regions. Furthermore, the retrieval shall consider not only low-level semantic correspondence between image regions and textual words but also a higher semantic correlation between different intra-modal relationships. We propose a multi-layer graph convolutional network with object-level, object-relational-level, and higher-level learning sub-networks. Our method learns hierarchical semantic correspondences by both local and global alignment. We further introduce a self-attention mechanism after the word embedding to weaken insignificant words in the sentence and a cross-attention mechanism to guide the learning of image features. Extensive experiments on Flickr30K and MS-COCO datasets demonstrate the effectiveness and superiority of our proposed method. Sheng Zeng, Changhong Liu, Jun Zhou 0001, Aiwen Jiang |
ICMR | 5 |
| 2022 | A Dense Prediction ViT Network for Single Image Bokeh Rendering
Zhifeng Wang 0006, Aiwen Jiang |
PRCV (4) | 2 |
| 2022 | Semantic-aware automatic image colorization via unpaired cycle-consistent self-supervised networkabstractAutomatic image colorization without manual interventions is an ill-conditioned and inherently ambiguous problem. Most of existing methods focus on formulating colorization as a regression problem and learn parametric mappings from grayscale to color through deep neural networks. Due to the multimodalities of color-grayscale space, in many applications, it is not required to recover exact ground-truth color. Pair-wise pixel-to-pixel learning-based algorithms lack rationality. Techniques such as color space conversion techniques are then proposed to avoid such direct pixel learning. However, the coloring results after color space conversion are blunt and unnatural. In this paper, we hold viewpoints that a reasonable solution is to generate some colorized result that looks natural. No matter what color a region is to be assigned, the colorized region should be semantically and spatially consistent. In this paper, we propose an effective semantic-aware automatic colorization model via unpaired cycle-consistent self-supervised network. Low-level monochrome loss, perceptual identity loss and high-level semantic-consistence loss, together with adversarial loss, are introduced to guide network self-training. We train and test our model on randomly selected subsets from PASCAL VOC 2012. The experimental results including human subjective studies demonstrate that, compared with state-of-the-art methods, our proposed model can achieve more convincing and superior results. Relevant source code is available at https://github.com/YuSuen/ACCycleGAN. Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
Int. J. Intell. Syst. | 2 |
| 2022 | Self-supervised multi-scale pyramid fusion networks for realistic bokeh effect rendering
Zhifeng Wang 0006, Aiwen Jiang, Chunjie Zhang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Efficient and Accurate Multi-Scale Topological Network for Single Image DehazingabstractSingle image dehazing is a challenging ill-posed problem that has drawn significant attention in the last few years. Recently, convolutional neural networks have achieved great success in image dehazing. However, it is still difficult for these increasingly complex models to recover accurate details from the hazy image. In this paper, we pay attention to the feature extraction and utilization of the input image itself. To achieve this, we propose a Multi-scale Topological Network (MSTN) to fully explore the features at different scales. Meanwhile, we design a Multi-scale Feature Fusion Module (MFFM) and an Adaptive Feature Selection Module (AFSM) to achieve the selection and fusion of features at different scales, so as to achieve progressive image dehazing. This topological network provides a large number of search paths that enable the network to extract abundant image features as well as strong fault tolerance and robustness. In addition, ASFM and MFFM can adaptively select important features and ignore interference information when fusing different scale representations. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods. Qiaosi Yi, Juncheng Li 0003, Faming Fang, Aiwen Jiang, Guixu Zhang |
IEEE Trans. Multim. | 4 |
| 2021 | MSNet: A novel end-to-end single image dehazing network with multiple inter-scale dense skip-connectionsabstractAbstract Dehazing is a challenging ill‐posed image restoration task. Various prior‐based and learning‐based methods have been proposed. Among them, end‐to‐end deep models achieve great success on performance improvement. However, most of them are concentrated on feature learning within the same block scale in isolation, and cannot perform associated analysis well on feature characteristics of different scales. Inter‐scale information reuse which is especially beneficial to image restoration is often neglected. Therefore, in this paper, a novel end‐to‐end network with multiple inter‐scale dense skip‐connections for image dehazing is proposed. Sufficient complementary information combination is considered through dense inter‐scale skip‐connections among encoder and decoder block layers. Besides avoiding gradient vanishing, a kind of bottleneck residual block is proposed to control the importance of local gradients at different scales over global learning process. Extensive comparisons and ablation studies on public dehazing datasets and real‐world images have been conducted. The experiment results demonstrate that the proposed novel elements can ensure more stable training process and superior testing performance with great improvements on PSNR and SSIM. Authors' haze‐removal results consistently comply satisfactorily with real situations, having much higher definition and contrast without colour distortion than those from the state‐of‐the‐art methods compared in this paper. Qiaosi Yi, Aiwen Jiang, Xiaolin Deng, Changhong Liu |
IET Image Process. | 2 |
| 2021 | Deliberation on object-aware video style transfer network with long-short temporal and depth-consistent constraints
Aiwen Jiang, Jiancheng Pan, Jihua Ye |
Neural Comput. Appl. | 2 |
| 2021 | Ensemble single image deraining network via progressive structural boosting constraints
Long Peng 0003, Aiwen Jiang, Bo Liu 0006, Mingwen Wang 0001 |
Signal Process. Image Commun. | 2 |
| 2020 | Cumulative Rain Density Sensing Network for Single Image DerainabstractThis paper focuses on single image derain, which aims to restore clear image from single rain image. Through full consideration of different frequency information preservation and the complicated interactions between rain-streaks and background, a novel end-to-end cumulative rain-density sensing network (CRDNet) is proposed for adaptive rain-streaks removal. An effective W-Net with powerful learning ability is proposed as a key component to recover rain-invariant low-frequency signals. A cumulative rain-density classifier with a novel cost-sensitive label encoding strategy is proposed as an auxiliary network to improve discriminative power of extracted high-frequency rain-streaks through multi-task training. The proposed CRDNet has been compared with state-of-the-art methods on two public datasets. The quantitative and visual experimental results demonstrate that it can achieve excellent performance with great improvement. Related source code and models are available on github https://github.com/peylnog/CRDNet. Long Peng 0003, Aiwen Jiang, Qiaosi Yi, Mingwen Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | Single Image Colorization Via Modified CycleganabstractIn this paper, we focus on automatically colorizing single grayscale image without manual interventions. Most of existing methods tried to accurately restore unknown ground-truth colors and require paired training data for model optimization. However, the ideal restoration objective and strict training constraints limited their performance. Inspired by CycleGAN, we formulate the process of colorization as image-to-image translation and propose an effective color-CycleGAN solution. High-level semantic identity loss and low-level color loss are additionally suggested for model optimization. Our method allows using unpaired images for training and direct prediction in rgb color space, which makes training data collection much easier and more general. We train our model on randomly selected PASCAL VOC 2007 images. All ablation study on loss function and comparisons with state-of-the-art methods are performed on grayscale SUN data. The experiment results show that our improvements on training loss could achieve better content consistence and generate better reasonable colors with less artifacts. Moreover, due to the bidirectional nature of our model, our proposed method provides a by-product that gives an excellent alternative way on color image graying. Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
ICIP | 2 |
| 2019 | Static Crowd Scene Analysis via Deep Network with Multi-branch Dilated Convolution BlocksabstractIn this paper, we have proposed a static crowd scene analysis network via multi-branch dilated convolution block, called MDBNet. It focuses on a joint task of estimating crowd count and high-quality density map from static single image. The proposed MDBNet follows one-stage object detection framework, and consists of two parts: pre-trained convolutional layers as the front end for high-level feature extraction and cascaded multi-branch dilated convolution block as the back end for context information aggregation on different ranges. Pixel-wise objectness probabilities are predicted and regressed to generate density map. The proposed MDBNet is an easy training model with strong learning ability. We have tested it on two public datasets (ShanghaiTech dataset and the UFC_CC_50 dataset). On almost all evaluation criterions, the proposed method has achieved superior performance. Especially on structure quality criterions, including our newly introduced spatial adjusted mutual information measurement, the MDBNet reports a new state-of-the-art performance. The source code will be distributed depending on publication of our work. Aiwen Jiang, Qiaosi Yi, Xiaolin Deng, Jianyi Wan, Mingwen Wang 0001 |
IJCNN | 2 |
| 2019 | Deep residual refining based pseudo-multi-frame network for effective single image super-resolutionabstractSingle image super‐resolution (SISR) has gained great attraction and progress in recent years. Since the SISR is an ill‐posed inverse problem, most researchers are concentrated on making efforts to learn effective and reasonable mapping functions from low‐resolution observation to its potential high‐resolution (HR) counterpart. In this study, the authors have proposed a deep residual refining based pseudo‐multi‐frame network for efficient SISR. A channel‐wise attention mechanism is employed for residual refinement. It can ease residual learning process through explicitly modelling non‐linear dependencies between channels by using global information embedding. Multiple potential HRs from different deconvolutional layers are further artificially learned, and then adaptively fused into final desired HR image. The authors call this strategy as pseudo‐multi‐frame SR. It could make full use of available redundant information possessed in hierarchical layers. They have evaluated the proposed network on several popular benchmark datasets. The experimental results have shown that the two highlights proposed can consistently boost final performance. The proposed network can outperform most of the state‐of‐the‐art methods with acceptable less parameters. Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Bo Liu 0006, Jihua Ye, Mingwen Wang 0001 |
IET Image Process. | 2 |
| 2018 | Progressive Feature Fusion Network for Realistic Image Dehazing
Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Mingwen Wang 0001 |
ACCV (1) | 2 |
| 2018 | An Effective Single-Image Super-Resolution Model Using Squeeze-and-Excitation Networks
Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Jihua Ye, Mingwen Wang 0001 |
ICONIP (6) | 2 |
| 2018 | Learning discriminative representations for semantical crossmodal retrieval
Aiwen Jiang, Yi Li 0025, Mingwen Wang 0001 |
Multim. Syst. | 1 |
| 2017 | Deep Multimodal Reinforcement Network with Contextually Guided Recurrent Attention for Image Question Answering
Aiwen Jiang, Bo Liu 0006, Mingwen Wang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2015 | Local similarity preserved hashing learning via Markov graph for efficient similarity search
Aiwen Jiang, Mingwen Wang 0001, Jianyi Wan |
Neurocomputing | 2 |
| 2013 | Listwise Approach to Learning to Rank for Automatic Evaluation of Machine Translation
Maoxi Li, Aiwen Jiang, Mingwen Wang 0001 |
MTSummit | 2 |
| 2010 | A New Biologically Inspired Feature for Scene Image ClassificationabstractScene classification is a hot topic in pattern recognition and computer vision area. In this paper, based on the past research on vision neuroscience, we proposed a new biologically inspired feature method for scene image classification. The new feature accounts for the visual processing from simple cell to complex cell in V1 area, and also the spatial layout for scene gist signature. It provides a different line and model revision to consider some nonlinearities inV1 area. We compare it with traditional HMAX model and recently proposed ScSPM model, and experiment on a popular 15 scenes dataset. We show that our proposed method has many important differences and merits. The experiment results also show that our method outperforms the state-of-the-art like ScSPM and KSPM model. Aiwen Jiang, Chunheng Wang, Baihua Xiao, Ruwei Dai |
ICPR | 1 |
| 2008 | Calibrated Rank-SVM for multi-label image categorizationabstractIn the area of multi-label image categorization, there are two important issues: label classification and label ranking. The former refers to whether a label is relevant or not, and the latter refers to what extent a label is relevant to an image. However, few existing papers have considered them in a holistic way. In this paper we will suggest a concrete improved method, named calibrated RankSVM, to bridge the gap between multi-label classification and label ranking. Through incorporating a virtual label as a calibrated scale [1], the threshold selection stage is embedded into ranking learning stage. This holistic way is essentially different from conventional rank methods, making our proposed method more suitable for multi-label classification task. The experiments on image have demonstrated that our algorithm has better multi-label classification performances than conventional ranksvm while preserving its good ranking characteristics. Aiwen Jiang, Chunheng Wang, Yuanping Zhu |
IJCNN | 1 |