Jiechong Song

dblp:304/1384 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0001-7064-8097ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2024 DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing
abstract
Large-scale Text-to-Image (T2I) diffusion models have revolutionized image generation over the last few years. Although owning diverse and high-quality generation capabilities, translating these abilities to fine-grained Image editing remains challenging. In this paper, we propose DiffEditor to rectify two weaknesses in existing diffusion-based image editing: (1) in complex scenarios, editing results often lack editing accuracy and exhibit unexpected artifacts; (2) lack of flexibility to harmonize editing operations, e.g., imagine new content. In our solution, we introduce image prompts in fine-grained image editing, cooperating with the text prompt to better describe the editing content. To increase the flexibility while maintaining content consistency, we locally combine stochastic differential equation (SDE) into the ordinary differential equation (ODE) sampling. In addition, we incorporate regional score-based gradient guidance and a time travel strategy into the diffusion sampling, further improving the editing quality. Extensive experiments demonstrate that our method can efficiently achieve state-of-the-art performance on various fine-grained image editing tasks, including editing within a single image (e.g., object moving, resizing, and content dragging) and across images (e.g., appearance replacing and object pasting). Our source code is released at https://github.com/MC-E/DragonDiffusion.
Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018
CVPR3
2024 DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models
abstract
Despite the ability of text-to-image (T2I) diffusion models to generate high-quality images, transferring this ability to accurate image editing remains a challenge. In this paper, we propose a novel image editing method, DragonDiffusion, enabling Drag-style manipulation on Diffusion models. Specifically, we treat image editing as the change of feature correspondence in a pre-trained diffusion model. By leveraging feature correspondence, we develop energy functions that align with the editing target, transforming image editing operations into gradient guidance. Based on this guidance approach, we also construct multi-scale guidance that considers both semantic and geometric alignment. Furthermore, we incorporate a visual cross-attention strategy based on a memory bank design to ensure consistency between the edited result and original image. Benefiting from these efficient designs, all content editing and consistency operations come from the feature correspondence without extra model fine-tuning. Extensive experiments demonstrate that our method has promising performance on various image editing tasks, including within a single image (e.g., object moving, resizing, and content dragging) or across images (e.g., appearance replacing and object pasting). Code is available at https://github.com/MC-E/DragonDiffusion.
Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018
ICLR3
2023 Large-Capacity and Flexible Video Steganography via Invertible Neural Network
abstract
Video steganography is the art of unobtrusively concealing secret data in a cover video and then recovering the secret data through a decoding protocol at the receiver end. Although several attempts have been made, most of them are limited to low-capacity and fixed steganography. To rectify these weaknesses, we propose a Large-capacity and Flexible Video Steganography Network (LF-VSN) in this paper. For large-capacity, we present a reversible pipeline to perform multiple videos hiding and recovering through a single invertible neural network (INN). Our method can hide/recover 7 secret videos in/from 1 cover video with promising performance. For flexibility, we propose a key-controllable scheme, enabling different receivers to recover particular secret videos from the same cover video through specific keys. Moreover, we further improve the flexibility by proposing a scalable strategy in multiple videos hiding, which can hide variable numbers of secret videos in a cover video with a single model and a single training session. Extensive experiments demonstrate that with the significant improvement of the video steganography performance, our proposed LF-VSN has high security, large hiding capacity, and flexibility. The source code is available at https://github.com/MC-E/LF-VSN.
Chong Mou, Youmin Xu, Jiechong Song, Chen Zhao 0002, Bernard Ghanem, Jian Zhang 0018
CVPR3
2023 Optimization-Inspired Cross-Attention Transformer for Compressive Sensing
abstract
By integrating certain optimization solvers with deep neural networks, deep unfolding network (DUN) with good interpretability and high performance has attracted growing attention in compressive sensing (CS). However, existing DUNs often improve the visual quality at the price of a large number of parameters and have the problem of feature information loss during iteration. In this paper, we propose an Optimization-inspired Cross-attention Transformer (OCT) module as an iterative process, leading to a lightweight OCT-based Unfolding Framework (OCTUF) for image CS. Specifically, we design a novel Dual Cross Attention (Dual-CA) sub-module, which consists of an Inertia-Supplied Cross Attention (ISCA) block and a Projection-Guided Cross Attention (PGCA) block. ISCA block introduces multi-channel inertia forces and increases the memory effect by a cross attention mechanism between adjacent iterations. And, PGCA block achieves an enhanced information interaction, which introduces the inertia force into the gradient descent step through a cross attention block. Extensive CS experiments manifest that our OCTUF achieves superior performance compared to state-of-the-art methods while training lower complexity. Codes are available at https://github.com/songjiechong/OCTUF.
Jiechong Song, Chong Mou, Shiqi Wang 0001, Siwei Ma 0001, Jian Zhang 0018
CVPR1
2023 Deep Physics-Guided Unrolling Generalization for Compressed Sensing
Bin Chen 0006, Jiechong Song, Jingfen Xie, Jian Zhang 0018
Int. J. Comput. Vis.2
2023 Deep Memory-Augmented Proximal Unrolling Network for Compressive Sensing
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
Int. J. Comput. Vis.1
2023 Dynamic Path-Controllable Deep Unfolding Network for Compressive Sensing
abstract
Deep unfolding network (DUN) that unfolds the optimization algorithm into a deep neural network has achieved great success in compressive sensing (CS) due to its good interpretability and high performance. Each stage in DUN corresponds to one iteration in optimization. At the test time, all the sampling images generally need to be processed by all stages, which comes at a price of computation burden and is also unnecessary for the images whose contents are easier to restore. In this paper, we focus on CS reconstruction and propose a novel Dynamic Path-Controllable Deep Unfolding Network (DPC-DUN). DPC-DUN with our designed path-controllable selector can dynamically select a rapid and appropriate route for each image and is slimmable by regulating different performance-complexity tradeoffs. Extensive experiments show that our DPC-DUN is highly flexible and can provide excellent performance and dynamic adjustment to get a suitable tradeoff, thus addressing the main requirements to become appealing in practice. Codes are available at https://github.com/songjiechong/DPC-DUN.
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
IEEE Trans. Image Process.1
2021 Spatial-Temporal Synergic Prior Driven Unfolding Network for Snapshot Compressive Imaging
abstract
In order to develop a fast and accurate algorithm for snapshot compressive sensing (SCI), we combine the merits of two kinds of existing SCI methods: the interpretability of traditional model-based methods and the speed of learning-based ones. Concretely, we build a novel Spatial-Temporal synErgic Prior driven unfolding network for SCI, dubbed STEP-SCI, which is inspired by Half-Quadratic Splitting (HQS) for optimizing the SCI reconstruction model. To better utilize the correlations among frames, we develop a spatial-temporal synergic prior module for solving the proximal mapping problem, which explicitly exploits the temporal correlation and the spatial correlation in turn. All parameters in STEP-SCI are learned end-to-end. Extensive experiments demonstrate that our STEP-SCI has superior performance than existing state-of-the-art model-based and learning-based methods while maintaining fast inference speed.
Zhuoyuan Wu, Jiechong Song
ICME3
2021 Memory-Augmented Deep Unfolding Network for Compressive Sensing
abstract
Mapping a truncated optimization method into a deep neural network, deep unfolding network (DUN) has attracted growing attention in compressive sensing (CS) due to its good interpretability and high performance. Each stage in DUNs corresponds to one iteration in optimization. By understanding DUNs from the perspective of the human brain's memory processing, we find there exists two issues in existing DUNs. One is the information between every two adjacent stages, which can be regarded as short-term memory, is usually lost seriously. The other is no explicit mechanism to ensure that the previous stages affect the current stage, which means memory is easily forgotten. To solve these issues, in this paper, a novel DUN with persistent memory for CS is proposed, dubbed Memory-Augmented Deep Unfolding Network (MADUN). We design a memory-augmented proximal mapping module (MAPMM) by combining two types of memory augmentation mechanisms, namely High-throughput Short-term Memory (HSM) and Cross-stage Long-term Memory (CLM). HSM is exploited to allow DUNs to transmit multi-channel short-term memory, which greatly reduces information loss between adjacent stages. CLM is utilized to develop the dependency of deep information across cascading stages, which greatly enhances network representation capability. Extensive CS experiments on natural and MR images show that with the strong ability to maintain and balance information our MADUN outperforms existing state-of-the-art methods by a large margin. The source code is available at https://github.com/jianzhangcs/MADUN/.
Jiechong Song, Bin Chen 0006, Jian Zhang 0018
ACM Multimedia1