EDBT 2026 Demo / reviewers in the wild / expert
Zhiying Jiang
dblp:21/1662
· DBLP profile ↗
45ranked-venue papers
13as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 21 · 4 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Adaptation Guided Infrared and Visible Image FusionabstractInfrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the visible modality, while the infrared modality remains relatively stable. To overcome domain shifts caused by the imbalance between modalities during image fusion, we propose a Domain Adaptation Guided Infrared and Visible Image Fusion method, termed DAFusion, leveraging a dual-rank domain adapter to enable fast adaptation to diverse adverse conditions during image fusion. Specifically, trainable low-rank and high-rank embedding spaces are respectively used to capture knowledge common across domains (domain-shared) and those unique to target domains (domain-specific). To leverage the dual-rank adapter more effectively, we develop a homeostatic knowledge allotment strategy to integrate the distinct types of knowledge dynamically based on the uncertainty value of target domains. Since domain adaptation typically optimizes for feature alignment across domains and emphasizes invariance rather than preserving specific cues critical for image fusion, while the fusion objective requires retaining discriminative and complementary features, a conflict between the two modules appears. To reconcile this, we further adopt a bi-level optimization framework that structurally decouples the two objectives, enabling the fusion module to steer the adaptation process while benefiting in return from domain-aligned representations. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving both an enhancement in fusion quality and an improvement on subsequent high-level tasks. Tianwei Guan, Haozhen Wei, Zecheng Xu, Zhiying Jiang, Jinyuan Liu 0001, Xingyuan Li 0005 |
AAAI | 6 |
| 2026 | Conditional Prompt Learning via Degradation Perception for Underwater Image EnhancementabstractUnderwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios. Mingze Yao, Zhiying Jiang, Xianping Fu, Huibing Wang |
AAAI | 2 |
| 2026 | HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionabstractInfrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence-induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat-Aware Diffusion for Turbulent InfraRed Video Super-Resolution, which injects heat-aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor-Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence-aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence-Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge-aware feature aggregation via turbulence gating and structure-aware attention. We built FLIR-IVSR, the first dataset for turbulent infrared VSR, comprising paired LR-HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR. Yang Zou 0004, Xingyue Zhu, Kaiqi Han, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 6 |
| 2026 | Versatile Luminosity Tuning: Relighting Illumination via Dual-Prompt Exposure Correction
Jinyuan Liu 0001, Gehui Li, Zhiying Jiang, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Risheng Liu |
Int. J. Comput. Vis. | 3 |
| 2026 | Illumination Refinement via Textual Cues: A Prompt-Driven Approach for Low-Light NeRF EnhancementabstractIn the realm of 3D scene modeling and rendering, the emergence of Neural Radiance Fields (NeRF) represents a significant leap forward. However, NeRF’s rendering performance suffers significantly when rendering images under low-light conditions. Existing approaches are optimized by enhancing low-light input images and combining NeRF models, but still fail to address the issues of multiview consistency and image quality. To address these challenges, our research introduces a textual constraint-prompted enhancement method that facilitates low-light image brightening and new view synthesis in an unsupervised manner. Specifically, we devise a semantic calibration strategy that employs positive and negative prompts to motivate and penalize the network towards attributes associated with high-quality images and exploits the capability of visual language models in semantic parsing to align the generated images with textual descriptors to improve image generation quality. In addition, to address the multiview consistency problem, we propose a two-layer optimization strategy, where the semantic cue optimization in the upper layer and the new view generation in the lower layer interact with each other to achieve a balance between luminance consistency and structural integrity by combining these improved images with text-driven semantic features. Comprehensive tests on two datasets with different resolutions, LOM and LLFF, show that our approach outperforms existing methods by significantly improving the brightness and clarity of low-light images to state-of-the-art while preserving the natural appearance and details. Xinrui Ju, Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SeFENet: Robust Deep Homography Estimation via Semantic-Driven Feature EnhancementabstractImages captured in harsh environments often exhibit blurred details, reduced contrast, and color distortion, which hinder feature detection and matching, thereby affecting the accuracy and robustness of homography estimation. While visual enhancement can improve contrast and clarity, it may introduce visual-tolerant artifacts that obscure the structural integrity of images. Considering the resilience of semantic information against environmental interference, we propose a semantic-driven feature enhancement network for robust homography estimation, dubbed SeFENet. Concretely, in our homography estimation network —— Target Aware Homography Estimation Module(TAHEM), we first introduce an innovative hierarchical scale-aware module to expand the receptive field by aggregating multi-scale information, thereby effectively extracting image’s structural features under diverse harsh conditions. Subsequently, we employ a Semantic Extraction Module to extract multi-scale semantic features from the input images. Combined with a high-level perceptual framework, this enables degradation-tolerant semantic feature extraction. Building upon this, the Semantic-Guide Meta Constraints module leverages a meta-learning training strategy to effectively fuse the semantic features with structural features. By internal-external alternating optimization, the proposed network achieves implicit semantic-wise feature enhancement, thereby improving the robustness of homography estimation in adverse environments by strengthening the local feature comprehension and context information extraction. Experimental results under both normal and harsh conditions demonstrate that SeFENet significantly outperforms SOTA methods, reducing point match error by at least 41% on the large-scale datasets. Zeru Shi, Zengxi Zhang, Kemeng Cui, Ruizhe An, Jinyuan Liu 0001, Zhiying Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D DetectionabstractThe sparse cross-modality detector offers more advantages than its counterpart, the Bird’s-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4% mAP and 74.7% NDS, while running$1.84\times $faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing. The code is available onhttps://github.com/xuehaipiaoxiang/CrossRay3D Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang 0060, Xianzhu Zeng, Li Wang 0092, Zhiwei Li 0011, Zijian Zeng 0001, Zhiying Jiang, Huaping Liu 0001, Kunfeng Wang |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2025 | DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-ResolutionabstractInfrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR Xingyuan Li 0005, Yang Zou 0004, Zhixin Chen, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
CVPR | 6 |
| 2025 | DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image FusionabstractInfrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo. Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
CVPR | 6 |
| 2025 | TextMEF: Text-guided Prompt Learning for Multi-exposure Image FusionabstractMulti-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in unsatisfactory visual effects such as hallucinated details and distorted color tones. With this regard, we propose TextMEF, a prompt-driven fusion method enhanced by prompt learning, for multi-exposure image fusion. Specifically, we learn a set of prompts based on text-image similarity among negative and positive samples (over-exposed, under-exposed images, and well-exposed ones). These learned prompts are seamlessly integrated into the loss function, providing high-level guidance for constraining non-uniform exposure regions. Furthermore, we develop a attention Mamba module effectively translates over-/under- exposed regional features into exposure invariant space and ensure them to build efficient long-range dependency to high dynamic range image. Extensive experimental results on three publicly available benchmarks demonstrate that our TextMEF significantly outperforms state-of-the-art approaches in both visual inspection and objective analysis. Jinyuan Liu 0001, Qianjun Huang, Guanyao Wu, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
IJCAI | 5 |
| 2025 | Physics-Guided Sonar Image Fine-grained Recognition under Scarce AnnotationsabstractSonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy. Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001 |
ACM Multimedia | 5 |
| 2025 | Toward a Training-Free Plug-and-Play Refinement Framework for Infrared and Visible Image Registration and FusionabstractInfrared and Visible Image Fusion (IVIF) under unregistered conditions has been of great interest in various visual tasks under challenging environments. While existing approaches often demonstrate promising results on specific benchmarks, they tend to exhibit performance drops in unseen scenarios and incur high computational overhead when retrained on new datasets. To address these challenges, we propose TRACE, a Training-free Reinforcement-based Alignment method for Cross-modality Enhancement, which incorporates Evaluator, a rewarding network, into an evaluation-driven Reinforcement Learning (RL) framework, enabling efficient and plug-and-play refinement of any existing registration approach. Specifically, TRACE constructs the Evaluator network to assess the alignment quality of the given registration model, generating confidence scores and adjustment masks via spatial and channel attention. Leveraging these cues as RL rewards, TRACE iteratively refines the registration network to mitigate misalignments until the accumulated improvement is satisfied. Due to its training-free and plug-and-play nature, TRACE notably enhances fusion results across diverse and unseen scenarios. TRACE achieves impressive improvements in different methods across diverse datasets with minimal computational cost. The project page is available at https://github.com/pubyLu/TRACE. Yang Zou 0004, Xingyuan Li 0005, Xingyue Zhu, Kaiqi Han, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
ACM Multimedia | 6 |
| 2025 | Depth-Supervised Fusion Network for Seamless-Free Image StitchingabstractImage stitching synthesizes images captured from multiple perspectives into a single image with a broader field of view. The significant variations in object depth often lead to large parallax, resulting in ghosting and misalignment in the stitched results. To address this, we propose a depth-consistency-constrained seamless-free image stitching method. First, to tackle the multi-view alignment difficulties caused by parallax, a multi-stage mechanism combined with global depth regularization constraints is developed to enhance the alignment accuracy of the same apparent target across different depth ranges. Second, during the multi-view image fusion process, an optimal stitching seam is determined through graph-based low-cost computation, and a soft-seam region is diffused to precisely locate transition areas, thereby effectively mitigating alignment errors induced by parallax and achieving natural and seamless stitching results. Furthermore, considering the computational overhead in the shift regression process, a reparameterization strategy is incorporated to optimize the structural design, significantly improving algorithm efficiency while maintaining optimal performance. Extensive experiments demonstrate the superior performance of the proposed method against the existing methods. Code is available at https://github.com/DLUT-YRH/DSFN. Zhiying Jiang, Ruhao Yan, Zengxi Zhang, Jinyuan Liu 0001 |
NeurIPS | 1 |
| 2025 | Enhancing Infrared Vision: Progressive Prompt Fusion Network and BenchmarkabstractWe engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one enhancement methods, commonly applied to RGB sensors, often demonstrate limited effectiveness due to the significant differences in imaging models. In sight of this, we first revisit the imaging mechanism and introduce a Recurrent Prompt Fusion Network (RPFN). Specifically, the RPFN initially establishes prompt pairs based on the thermal imaging process. For each type of degradation, we fuse the corresponding prompt pairs to modulate the model's features, providing adaptive guidance that enables the model to better address specific degradations under single or multiple conditions.In addition, a selective recurrent training mechanism is introduced to gradually refine the model's handling of composite cases to align the enhancement process, which not only allows the model to remove camera noise and retain key structural details, but also enhancing the overall contrast of the thermal image. Furthermore, we introduce the most comprehensive high-quality infrared benchmark covering a wide range of scenarios. Extensive experiments substantiate that our approach not only delivers promising visual results under specific degradation but also significantly improves performance on complex degradation scenes, achieving a notable 8.76% improvement. Jinyuan Liu 0001, Zhu Liu 0004, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu |
NeurIPS | 4 |
| 2025 | Image Stitching in Adverse Condition: A Bidirectional-Consistency Learning Framework and BenchmarkabstractDeep learning-based image stitching methods have achieved promising performance on conventional stitching datasets. However, real-world scenarios may introduce challenges such as complex weather conditions, illumination variations, and dynamic scene motion, which severely degrade image quality and lead to significant misalignment in stitching results. To solve this problem, we propose an adverse condition-tolerant image stitching network, dubbed ACDIS. We first introduce a bidirectional consistency learning framework, which ensures reliable alignment through an iterative optimization paradigm that integrates differentiable image restoration and Gaussian-distribute encoded homography estimation. Subsequently, we incorporate motion constraints into the seamless composition network to produce robust stitching results without interference from moving scenes. We further propose the first adverse scene image stitching dataset, which covers diverse parallax and scenes under low-light, haze, and underwater environments. Extensive experiments show that the proposed method can generate visually pleasing stitched images under adverse conditions, outperforming state-of-the-art methods. Zengxi Zhang, Junchen Ge, Zhiying Jiang, Miao Zhang 0004, Jinyuan Liu 0001 |
NeurIPS | 3 |
| 2025 | HUPE: Heuristic Underwater Perceptual Enhancement with Semantic Collaborative Learning
Zengxi Zhang, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
Int. J. Comput. Vis. | 2 |
| 2025 | Infrared and Visible Image Fusion: From Data Compatibility to Task AdaptionabstractInfrared-visible image fusion (IVIF) is a fundamental and critical task in the field of computer vision. Its aim is to integrate the unique characteristics of both infrared and visible spectra into a holistic representation. Since 2018, growing amount and diversity IVIF approaches step into a deep-learning era, encompassing introduced a broad spectrum of networks or loss functions for improving visual enhancement. As research deepens and practical demands grow, several intricate issues like data compatibility, perception accuracy, and efficiency cannot be ignored. Regrettably, there is a lack of recent surveys that comprehensively introduce and organize this expanding domain of knowledge. Given the current rapid development, this paper aims to fill the existing gap by providing a comprehensive survey that covers a wide array of aspects. Initially, we introduce a multi-dimensional framework to elucidate the prevalent learning-based IVIF methodologies, spanning topics from basic visual enhancement strategies to data compatibility, task adaptability, and further extensions. Subsequently, we delve into a profound analysis of these new approaches, offering a detailed lookup table to clarify their core ideas. Last but not the least, We also summarize performance comparisons quantitatively and qualitatively, covering registration, fusion and follow-up high-level tasks. Beyond delving into the technical nuances of these learning-based fusion approaches, we also explore potential future directions and open issues that warrant further exploration by the community. Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | DRNet: Learning a dynamic recursion network for chaotic rain streak removal
Zhiying Jiang, Risheng Liu, Shuzhou Yang, Zengxi Zhang, Xin Fan 0001 |
Pattern Recognit. | 1 |
| 2025 | Harmonized Domain Enabled Alternate Search for Infrared and Visible Image AlignmentabstractInfrared and visible image alignment is essential and critical to the fusion and multi-modal perception applications. It addresses discrepancies in position and scale caused by spectral properties and environmental variations, ensuring precise pixel correspondence and spatial consistency. Existing manual calibration requires regular maintenance and exhibits poor portability, challenging the adaptability of multi-modal application in dynamic environments. In this paper, we propose a harmonized representation based infrared and visible image alignment, achieving both high accuracy and scene adaptability. Specifically, with regard to the disparity between multi-modal images, we develop an invertible translation process to establish a harmonized representation domain that effectively encapsulates the feature intensity and distribution of both infrared and visible modalities. Building on this, we design a hierarchical framework to correct deformations inferred from the harmonized domain in a coarse-to-fine manner. Our framework leverages advanced perception capabilities alongside residual estimation to enable accurate regression of sparse offsets, while an alternate correlation search mechanism ensures precise correspondence matching. Furthermore, we propose the first ground truth available misaligned infrared and visible image benchmark for evaluation. Extensive experiments validate the effectiveness of the proposed method against the state-of-the-arts, advancing the subsequent applications further. Code and dataset are available at https://github.com/Jzy2017/HR4IR. Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible AttacksabstractImage stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations and distortions which go unnoticed by the human visual system tend to attack the correspondence matching, impairing the performance of image stitching algorithms. In light of this challenge, this paper presents the first attempt to improve the robustness of image stitching against adversarial attacks. Specifically, we introduce a stitching-oriented attack (SoA), tailored to amplify the alignment loss within overlapping regions, thereby targeting the feature matching procedure. To establish an attack resistant model, we delve into the robustness of stitching architecture and develop an adaptive adversarial training (AAT) to balance attack resistance with stitching precision. In this way, we relieve the gap between the routine adversarial training and benign models, ensuring resilience without quality compromise. Comprehensive evaluation across real-world and synthetic datasets validate the deterioration of SoA on stitching performance. Furthermore, AAT emerges as a more robust solution against adversarial perturbations, delivering superior stitching results. Code is available at: https://github.com/Jzy2017/TRIS. Zhiying Jiang, Xingyuan Li 0005, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
AAAI | 1 |
| 2024 | Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View SynthesisabstractNeural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this issue, we propose an innovative approach capable of effectively modeling and rendering images under multiple exposure conditions. Our method adaptively learns the characteristics of images under different exposure conditions through an unsupervised evaluator-simulator structure for HDR (High Dynamic Range) fusion. This approach enhances NeRF's comprehension and handling of light variations, leading to the generation of images with appropriate brightness. Simultaneously, we present a bilevel optimization method tailored for novel view synthesis, aiming to harmonize the luminance information of input images while preserving their structural and content consistency. This approach facilitates the concurrent optimization of multi-exposure correction and novel view synthesis, in an unsupervised manner. Through comprehensive experiments conducted on the LOM and LOL datasets, our approach surpasses existing methods, markedly enhancing the task of novel view synthesis for multi-exposure environments and attaining state-of-the-art results. The source code can be found at https://github.com/Archer-204/AME-NeRF. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 3 |
| 2024 | AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularizationabstract3D object detection plays a crucial role in intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes while most of existing methods fail in these scenes. To address this issue, this paper proposes a monocular 3D detection model, termed AEAM3D, which effectively mitigates the degradation of detection performance in various harsh environments. Additionally, we assemble a new adverse 3D object detection dataset encompassing some challenging scenes, including rainy, foggy, and low light weather conditions. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches by an average of 3.12% in terms of APR40for car category across adverse environments. Yixin Lei, Xingyuan Li 0005, Zhiying Jiang, Xinrui Ju, Jinyuan Liu 0001 |
ICASSP | 3 |
| 2024 | Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders images across a spectrum of exposure conditions. Our approach utilizes an unsupervised classifier-generator structure for HDR fusion, significantly enhancing NeRF’s ability to comprehend and adjust to light variations, leading to the generation of images with appropriate brightness. Extensive evaluations on the LOM[1] and LOL[2] datasets underscore our method’s edge. Our approach significantly improves the task of novel view synthesis for multi-exposure images, attaining state-of-the-art results. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Tiantian Yan, Jinyuan Liu 0001 |
ICASSP | 3 |
| 2024 | Multispectral Image Stitching via Global-Aware Quadrature Pyramid RegressionabstractImage stitching is a critical task in panorama perception that involves combining images captured from different viewing positions to reconstruct a wider field-of-view (FOV) image. Existing visible image stitching methods suffer from performance drops under severe conditions since environmental factors can easily impair visible images. In contrast, infrared images possess greater penetrating ability and are less affected by environmental factors. Therefore, we propose an infrared and visible image-based multispectral image stitching method to achieve all-weather, broad FOV scene perception. Specifically, based on two pairs of infrared and visible images, we employ the salient structural information from the infrared images and the textual details from the visible images to infer the correspondences within different modality-specific features. For this purpose, a multiscale progressive mechanism coupled with quadrature correlation is exploited to improve regression in different modalities. Exploiting the complementary properties, accurate and credible homography can be obtained by integrating the deformation parameters of the two modalities to compensate for the missing modality-specific information. A global-aware guided reconstruction module is established to generate an informative and broad scene, wherein the attentive features of different viewpoints are introduced to fuse the source images with a more seamless and comprehensive appearance. We construct a high-quality infrared and visible stitching dataset for evaluation, including real-world and synthetic sets. The qualitative and quantitative results demonstrate that the proposed method outperforms the intuitive cascaded fusion-stitching procedure, achieving more robust and credible panorama generation. Code and dataset are available at https://github.com/Jzy2017/MSGA. Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
IEEE Trans. Image Process. | 1 |
| 2023 | What the DAAM: Interpreting Stable Diffusion Using Cross AttentionabstractRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Raphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
ACL (1) | 4 |
| 2023 | Multi-Spectral Image Stitching via Spatial Graph ReasoningabstractMulti-spectral image stitching leverages the complementarity between infrared and visible images to generate a robust and reliable wide field-of-view~(FOV) scene. The primary challenge of this task is to explore the relations between multi-spectral images for aligning and integrating multi-view scenes. Capitalizing on the strengths of Graph Convolutional Networks (GCNs) in modeling feature relationships, we propose a spatial graph reasoning based multi-spectral image stitching method that effectively distills the deformation and integration of multi-spectral images across different viewpoints. To accomplish this, we embed multi-scale complementary features from the same view position into a set of nodes. The correspondence across different views is learned through powerful dense feature embeddings, where both inter- and intra-correlations are developed to exploit cross-view matching and enhance inner feature disparity. By introducing long-range coherence along spatial and channel dimensions, the complementarity of pixel relations and channel interdependencies aids in the reconstruction of aligned multi-view features, generating informative and reliable wide FOV scenes. Moreover, we release a challenging dataset named ChaMS, comprising both real-world and synthetic sets with significant parallax, providing a new option for comprehensive evaluation. Extensive experiments demonstrate that our method surpasses the state-of-the-arts. Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 1 |
| 2023 | Fearless Luminance Adaptation: A Macro-Micro-Hierarchical Transformer for Exposure CorrectionabstractPhotographs taken with less-than-ideal exposure settings often display poor visual quality. Since the correction procedures vary significantly, it is difficult for a single neural network to handle all exposure problems. Moreover, the inherent limitations of convolutions, hinder the models ability to restore faithful color or details on extremely over-/under-exposed regions. To overcome these limitations, we propose a Macro-Micro-Hierarchical transformer, which consists of a macro attention to capture long-range dependencies, a micro attention to extract local features, and a hierarchical structure for coarse-to-fine correction. In specific, the complementary macro-micro attention designs enhance locality while allowing global interactions. The hierarchical structure enables the network to correct exposure errors of different scales layer by layer. Furthermore, we propose a contrast constraint and couple it seamlessly in the loss function, where the corrected image is pulled towards the positive sample and pushed away from the dynamically generated negative samples. Thus the remaining color distortion and loss of detail can be removed. We also extend our method as an image enhancer for low-light face recognition and low-light semantic segmentation. Experiments demonstrate that our approach obtains more attractive results than state-of-the-art methods quantitatively and qualitatively. Gehui Li, Jinyuan Liu 0001, Long Ma 0002, Zhiying Jiang, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 4 |
| 2023 | WaterFlow: Heuristic Normalizing Flow for Underwater Image Enhancement and BeyondabstractUnderwater images suffer from light refraction and absorption, which impairs visibility and interferes the subsequent applications. Existing underwater image enhancement methods mainly focus on image quality improvement, ignoring the effect on practice. To balance the visual quality and application, we propose a heuristic normalizing flow for detection-driven underwater image enhancement, dubbed WaterFlow. Specifically, we first develop an invertible mapping to achieve the translation between the degraded image and its clear counterpart. Considering the differentiability and interpretability, we incorporate the heuristic prior into the data-driven mapping procedure, where the ambient light and medium transmission coefficient benefit credible generation. Furthermore, we introduce a detection perception module to transmit the implicit semantic guidance into the enhancement procedure, where the enhanced images hold more detection-favorable features and are able to promote the detection performance. Extensive experiments prove the superiority of our WaterFlow, against state-of-the-art methods quantitatively and qualitatively. Zengxi Zhang, Zhiying Jiang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 2 |
| 2023 | Examining the Usefulness of Customer Reviews for Mobile Applications: The Role of Developer ResponsivenessabstractIn the context of mobile applications (apps), the role of customers has been transformed from mere passive adopters to active co-creators through contribution of user reviews. However, customers might not always possess the required technical expertise to make commercially feasible suggestions. The value of customer reviews also varied due to their unmanageable volume and content irrelevance. In our study, over 189,000 user reviews with over 50 apps would be analyzed using review analysis and multivariate regression analysis to examine the impacts of innovation and improvement led by customers on app performance in terms of app revenues. The developers' lead time in responding to user reviews would be included as a moderator to investigate whether app performance would be enhanced if developers respond faster. This study should represent one of the first few attempts in offering empirical confirmation of the value of co-creation of apps with customers. The authors also present methodological contributions by establishing operationalization and analyses of user reviews. Zhiying Jiang, Vanessa Liu, Miriam Erne |
J. Database Manag. | 1 |
| 2023 | Bilevel modeling investigated generative adversarial framework for image restoration
Zhiying Jiang, Zengxi Zhang, Yiyao Yu, Risheng Liu |
Vis. Comput. | 1 |
| 2023 | Publisher Correction: Bilevel modeling investigated generative adversarial framework for image restoration
Zhiying Jiang, Zengxi Zhang, Yiyao Yu, Risheng Liu |
Vis. Comput. | 1 |
| 2023 | A unified image fusion framework with flexible bilevel paradigm integration
Jinyuan Liu 0001, Zhiying Jiang, Guanyao Wu, Risheng Liu, Xin Fan 0001 |
Vis. Comput. | 2 |
| 2022 | The World Is Our Classroom: Developing a Model for International Virtual Internships - The Global Innovations ProjectabstractIn the aftermath of COVID-19, remote working has become the norm, and graduates now need an even wider range of skills, which traditional classrooms and internships do not always provide. Working in multiple time zones, within global multi-cultural teams, and only ever meeting colleagues through online technology are just some of the challenges, which require a new type of global graduate. Transversal skills including leadership, collaboration, innovation, digital, green, organization and communication skills are critical. The disruption from COVID-19 also presents unprecedented opportunities to develop more inclusive approaches to internships and international experiences, to level the playing field for students with special needs, from underrepresented groups or with caring commitments. In this position paper, we present a new Global Innovation internship model that has the aim of allowing students to complete technology internships and projects by working together virtually on real world challenges, guided by experienced industry and academic mentors. The model is being developed as part of an Erasmus+ funded project, and the partnership includes seven Higher Education Institutions from six different countries around the world. This position paper describes the design and development of a pilot programme of the Global Innovations internship model. Paul Doyle, Brian Keegan, Damian Gordon, Anna Becevel, J. Paul Gibson, Zhiying Jiang, Dympna O'Sullivan |
CSEDU (1) | 6 |
| 2022 | Towards All Weather and Unobstructed Multi-Spectral Image Stitching: Algorithm and BenchmarkabstractImage stitching is a fundamental task that requires multiple images from different viewpoints to generate a wide field-of-viewing~(FOV) scene. Previous methods are developed on RGB images. However, the severe weather and harsh conditions, such as rain, fog, low light, strong light, etc., on visible images may introduce evident interference, leading to the distortion and misalignment of the stitched results. To remedy the deficient imaging of optical sensors, we investigate the complementarity across infrared and visible images to improve the perception of scenes in terms of visual information and viewing ranges. Instead of the cascaded fusion-stitching process, where the inaccuracy accumulation caused by image fusion hinders the stitch performance, especially content loss and ghosting effect, we develop a learnable feature adaptive network to investigate a stitch-oriented feature representation and perform the information complementary at the feature-level. By introducing a pyramidal structure along with the global fast correlation regression, the quadrature attention based correspondence is more responsible for feature alignment, and the estimation of sparse offsets can be realized in a coarse-to-fine manner. Furthermore, we propose the first infrared and visible image based multi-spectral image stitching dataset, covering a more comprehensive range of scenarios and diverse viewing baselines. Extensive experiments on real-world data demonstrate that our method reconstructs the wide FOV images with more credible structure and complementary information against state-of-the-arts. Zhiying Jiang, Zengxi Zhang, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 1 |
| 2022 | Few-Shot Non-Parametric Learning with Deep Latent Variable ModelabstractMost real-world problems that machine learning algorithms are expected to solve face the situation with (1) unknown data distribution; (2) little domain-specific knowledge; and (3) datasets with limited annotation. We propose Non-Parametric learning by Compression with Latent Variables (NPC-LV), a learning framework for any dataset with abundant unlabeled data but very few labeled ones. By only training a generative model in an unsupervised way, the framework utilizes the data distribution to build a compressor. Using a compressor-based distance metric derived from Kolmogorov complexity, together with few labeled data, NPC-LV classifies without further training. We show that NPC-LV outperforms supervised methods on all three datasets on image classification in the low data regime and even outperforms semi-supervised learning methods on CIFAR-10. We demonstrate how and when negative evidence lowerbound (nELBO) can be used as an approximate compressed length for classification. By revealing the correlation between compression rate and classification accuracy, we illustrate that under NPC-LV how the improvement of generative models can enhance downstream classification accuracy. Zhiying Jiang, Yiqin Dai, Ji Xin, Ming Li 0001, Jimmy Lin |
NeurIPS | 1 |
| 2022 | Target Oriented Perceptual Adversarial Fusion Network for Underwater Image EnhancementabstractDue to the refraction and absorption of light by water, underwater images usually suffer from severe degradation, such as color cast, hazy blur, and low visibility, which would degrade the effectiveness of marine applications equipped on autonomous underwater vehicles. To eliminate the degradation of underwater images, we propose a target oriented perceptual adversarial fusion network, dubbed TOPAL. Concretely, we consider the degradation factors of underwater images in terms of turbidity and chromatism. And according to the degradation issues, we first develop a multi-scale dense boosted module to strengthen the visual contrast and a deep aesthetic render module to perform the color correction, respectively. After that, we employ the dual channel-wise attention module and guide the adaptive fusion of latent features, in which both diverse details and credible appearance are integrated. To bridge the gap between synthetic and real-world images, a global-local adversarial mechanism is introduced in the reconstruction. Besides, perceptual information is also embedded into the process to assist the understanding of scenery content. To evaluate the performance of TOPAL, we conduct extensive experiments on several benchmarks and make comparisons among state-of-the-art methods. Quantitative and qualitative results demonstrate that our TOPAL improves the quality of underwater images greatly and achieves superior performance than others. Zhiying Jiang, Zhuoxiao Li, Shuzhou Yang, Xin Fan 0001, Risheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Twin Adversarial Contrastive Learning for Underwater Image Enhancement and BeyondabstractUnderwater images suffer from severe distortion, which degrades the accuracy of object detection performed in an underwater environment. Existing underwater image enhancement algorithms focus on the restoration of contrast and scene reflection. In practice, the enhanced images may not benefit the effectiveness of detection and even lead to a severe performance drop. In this paper, we propose an object-guided twin adversarial contrastive learning based underwater enhancement method to achieve both visual-friendly and task-orientated enhancement. Concretely, we first develop a bilateral constrained closed-loop adversarial enhancement module, which eases the requirement of paired data with the unsupervised manner and preserves more informative features by coupling with the twin inverse mapping. In addition, to confer the restored images with a more realistic appearance, we also adopt the contrastive cues in the training phase. To narrow the gap between visually-oriented and detection-favorable target images, a task-aware feedback module is embedded in the enhancement process, where the coherent gradient information of the detector is incorporated to guide the enhancement towards the detection-pleasing direction. To validate the performance, we allocate a series of prolific detectors into our framework. Extensive experiments demonstrate that the enhanced results of our method show remarkable amelioration in visual quality, the accuracy of different detectors conducted on our enhanced images has been promoted notably. Moreover, we also conduct a study on semantic segmentation to illustrate how object guidance improves high-level tasks. Code and models are available at https://github.com/Jzy2017/TACL. Risheng Liu, Zhiying Jiang, Shuzhou Yang, Xin Fan 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | A Bilevel Integrated Model With Data-Driven Layer Ensemble for Multi-Modality Image FusionabstractImage fusion plays a critical role in a variety of vision and learning applications. Current fusion approaches are designed to characterize source images, focusing on a certain type of fusion task while limited in a wide scenario. Moreover, other fusion strategies (i.e., weighted averaging, choose-max) cannot undertake the challenging fusion tasks, which furthermore leads to undesirable artifacts facilely emerged in their fused results. In this paper, we propose a generic image fusion method with a bilevel optimization paradigm, targeting on multi-modality image fusion tasks. Corresponding alternation optimization is conducted on certain components decoupled from source images. Via adaptive integration weight maps, we are able to get the flexible fusion strategy across multi-modality images. We successfully applied it to three types of image fusion tasks, including infrared and visible, computed tomography and magnetic resonance imaging, and magnetic resonance imaging and single-photon emission computed tomography image fusion. Results highlight the performance and versatility of our approach from both quantitative and qualitative aspects. Risheng Liu, Jinyuan Liu 0001, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo |
IEEE Trans. Image Process. | 3 |
| 2020 | Knowledge-Driven Deep Unrolling for Robust Image Layer SeparationabstractSingle-image layer separation targets to decompose the observed image into two independent components in terms of different application demands. It is known that many vision and multimedia applications can be (re)formulated as a separation problem. Due to the fundamentally ill-posed natural of these separations, existing methods are inclined to investigate model priors on the separated components elaborately. Nevertheless, it is knotty to optimize the cost function with complicated model regularizations. Effectiveness is greatly conceded by the settled iteration mechanism, and the adaption cannot be guaranteed due to the poor data fitting. What is more, for a universal framework, the most taxing point is that one type of visual cue cannot be shared with different tasks. To partly overcome the weaknesses mentioned earlier, we delve into a generic optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. First, we propose a general energy model with implicit priors, which is based on maximum a posterior, and employ the extensively accepted alternating direction method of multiplier to determine our elementary iteration mechanism. By unrolling with one general residual architecture prior and one task-specific prior, we attain a straightforward, flexible, and data-dependent image separation framework successfully. We apply our method to four different tasks, including single-image-rain streak removal, high-dynamic-range tone mapping, low-light image enhancement, and single-image reflection removal. Extensive experiments demonstrate that the proposed method is applicable to multiple tasks and outperforms the state of the arts by a large margin qualitatively and quantitatively. Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | PaperRobot: Incremental Draft Generation of Scientific IdeasabstractWe present a PaperRobot who performs as an automatic research assistant by (1) conducting deep understanding of a large collection of human-written papers in a target domain and constructing comprehensive background knowledge graphs (KGs); (2) creating new ideas by predicting links from the background KGs, by combining graph attention and contextual text attention; (3) incrementally writing some key elements of a new paper based on memory-attention networks: from the input title along with predicted related entities to generate a paper abstract, from the abstract to generate conclusion and future work, and finally from future work to generate a title for a follow-on paper.Turing Tests, where a biomedical domain expert is asked to compare a system output and a human-authored string, show PaperRobot generated abstracts, conclusion and future work sections, and new titles are chosen over human-written ones up to 30%, 24% and 12% of the time, respectively. 1 keeps almost the same across years.In 2012, US scientists estimated that they read, on average, only 264 papers per year (1 out of 5000 available papers), which is, statistically, not different from what they reported in an identical survey last conducted in 2005.PaperRobot automatically reads existing papers to build background knowledge graphs (KGs), in which nodes are entities/concepts and edges are the relations between these entities (Section 2.2). Qingyun Wang 0005, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji 0001, Mohit Bansal, Yi Luan |
ACL (1) | 3 |
| 2019 | Learning Aggregated Transmission Propagation Networks for Haze Removal and BeyondabstractSingle-image dehazing is an important low-level vision task with many applications. Early studies have investigated different kinds of visual priors to address this problem. However, they may fail when their assumptions are not valid on specific images. Recent deep networks also achieve a relatively good performance in this task. But unfortunately, due to the disappreciation of rich physical rules in hazes, a large amount of data are required for their training. More importantly, they may still fail when there exist completely different haze distributions in testing images. By considering the collaborations of these two perspectives, this paper designs a novel residual architecture to aggregate both prior (i.e., domain knowledge) and data (i.e., haze distribution) information to propagate transmissions for scene radiance estimation. We further present a variational energy-based perspective to investigate the intrinsic propagation behavior of our aggregated deep model. In this way, we actually bridge the gap between prior-driven models and data-driven networks and leverage advantages but avoid limitations of previous dehazing approaches. A lightweight learning framework is proposed to train our propagation network. Finally, by introducing a task-aware image separation formulation with a flexible optimization scheme, we extend the proposed model for more challenging vision tasks, such as underwater image enhancement and single-image rain removal. Experiments on both synthetic and real-world images demonstrate the effectiveness and efficiency of the proposed framework. Risheng Liu, Xin Fan 0001, Minjun Hou, Zhiying Jiang, Zhongxuan Luo, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Deep Layer Prior Optimization for Single Image Rain Streaks RemovalabstractVisible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed priors tend to over-smooth the background and leave too many rain streaks since the distribution of rain streaks is complex and disordered. In this work, we exploit a deep layer prior under the maximum a posterior framework to recover the intrinsic rain structure. The optimization of the resulted variational energy can be understood as simultaneously performing rain and image propagations based on data-dependent residual networks and task cues (e.g., total variation regularization), respectively. Experimental results on both synthetic and real test images demonstrate the effectiveness of our approach against both designed priors and fully data-dependent convolutional neural networks. Risheng Liu, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo |
ICASSP | 2 |
| 2018 | Single Image Layer Separation via Deep Admm UnrollingabstractSingle image layer separation aims to divide the observed image into two independent components according to special task requirements and has been widely used in many vision and multimedia applications. Because this task is fundamentally ill-posed, most existing approaches tend to design complex priors on the separated layers. However, the cost function with complex prior regularization is hard to optimize. The performance is also compromised by fixed iteration schemes and less data fitting ability. More importantly, it is also challenging to design a unified framework to separate image layers for different applications. To partially mitigate the above limitations, we develop a flexible optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. Specifically, we first design a general energy model with implicit priors and adopt the widely used alternating direction method of multiplier (ADMM) to establish our basic iteration scheme. By unrolling with residual convolution architectures, we successfully obtain a simple, flexible, and data-dependent image separation method. Extensive experiments on the tasks of rain streak removal and reflection removal validate the effectiveness of our approach. Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo |
ICME | 2 |
| 2018 | Describing a Knowledge BaseabstractWe aim to automatically generate natural language descriptions about an input structured knowledge base (KB).We build our generation framework based on a pointer network which can copy facts from the input KB, and add two attention mechanisms: (i) slot-aware attention to capture the association between a slot type and its corresponding slot value; and (ii) a new table position self-attention to capture the inter-dependencies among related slots.For evaluation, besides standard metrics including BLEU, METEOR, and ROUGE, we propose a KB reconstruction based metric by extracting a KB from the generation output and comparing it with the input KB.We also create a new data set which includes 106,216 pairs of structured KBs and their corresponding natural language descriptions for two distinct entity types.Experiments show that our approach significantly outperforms stateof-the-art methods.The reconstructed KB achieves 68.8% -72.6% F-score. 1 Qingyun Wang 0005, Xiaoman Pan, Lifu Huang, Boliang Zhang, Zhiying Jiang, Heng Ji 0001, Kevin Knight |
INLG | 5 |
| 2007 | An SRP Target Mode to Improve Read Performance of SRP-Based IB-SANs
Zhiying Jiang, Jizhong Han, Xigui Wang, Yonghao Zhou, Xubin He |
ISPA | 1 |