EDBT 2026 Demo / reviewers in the wild / expert
Zhiliang Wu
dblp:117/3106
· DBLP profile ↗
33ranked-venue papers
10as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 25 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax ConstraintsabstractDual-lens video inpainting aims to simultaneously restore missing or corrupted contents in videos captured by each lens of binocular systems. Although preliminary explorations have been conducted, existing methods still face two key challenges: limited exploitation of long-range reference information and inadequate modeling of inter-lens consistency in non-standard binocular systems. In this paper, we propose a novel dual-lens video inpainting framework named DLVINet, which addresses these challenges with two core components. Firstly, we develop a sparse spatial-temporal transformer (SSTT) that effectively utilizes the information from distant frames to complete the video contents of each lens individually. By employing sparse spatial-temporal attention with a channel selection mechanism, SSTT not only restores missing regions, but also avoids introducing redundant or irrelevant information. Furthermore, SSTT introduces a multi-scale feed-forward network to enrich the multi-scale representation of completed features. Secondly, we design a cross-lens texture transformer (CLTT) to model inter-lens consistency. By interacting with corresponding features between lenses under the guidance of cross-attention, CLTT captures global inter-lens correspondences. Such a design enables effective cross-view information modeling without being constrained by horizontal parallax, which is particularly critical for non-standard binocular systems. Extensive experiments demonstrate the effectiveness of our DLVINet. Zhiliang Wu, Kun Li 0008, Yunqiu Xu, Hehe Fan, Yi Yang 0001 |
AAAI | 1 |
| 2026 | One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query RefinementabstractLarge Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquiries and the structured logic required for machine activation.Existing alignment methods either incur prohibitive O(N ) costs by fine-tuning each model individually or rely on static prompts that fail to resolve query-level structural complexity.In this paper, we propose ReQueR (Reinforcement Query Refinement), a modular framework that treats reasoning elicitation as an inference-time alignment task.We train a specialized Refiner policy via Reinforcement Learning to rewrite raw queries into explicit logical decompositions, treating frozen LLMs as the environment.Rooted in the classical Zone of Proximal Development from educational psychology, we introduce the Adaptive Solver Hierarchy, a curriculum mechanism that stabilizes training by dynamically aligning environmental difficulty with the Refiner's evolving competence.ReQueR yields consistent absolute gains of 1.7%-7.2%across diverse architectures and benchmarks, outperforming strong baselines by 2.1% on average.Crucially, it provides a promising paradigm for one-to-many inference-time reasoning elicitation, enabling a single Refiner trained on a small set of models to effectively unlock reasoning in diverse unseen models. Yixiao Zhou 0001, Dongzhou Cheng, Zhiliang Wu, Yi Yang 0001, Yu Cheng 0001, Hehe Fan |
ACL (1) | 3 |
| 2026 | BeatDance: Generating beat-consistent 3D dance with hierarchical spatial-temporal modeling
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue 0001, Wei Zhou 0021 |
Pattern Recognit. | 8 |
| 2026 | Uncertainty-Guided Spatiotemporal Consistency Fusion Network for Infrared-Visible Video Fusion Under Extremely Low-Light ConditionsabstractInfrared-visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infrared-visible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID. Cheng Zhao 0003, Tianyun Song, Zhiliang Wu, Tianfu Wang 0001, Moncef Gabbouj, Guanghui Yue 0001, Bai Ying Lei, Wei Zhou 0021 |
IEEE Trans. Image Process. | 3 |
| 2025 | Prototypical Calibrating Ambiguous Samples for Micro-Action RecognitionabstractMicro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguity in micro-actions, which arises from the wide category range and subtle visual differences between categories. This oversight hampers the accuracy of micro-action recognition. In this paper, we propose a novel Prototypical Calibrating Ambiguous Network (PCAN) to unleash and mitigate the ambiguity of MAR. Firstly, we employ a hierarchical action-tree to identify the ambiguous sample, categorizing them into distinct sets of ambiguous samples of false negatives and false positives, considering both body- and action-level categories. Secondly, we implement an ambiguous contrastive refinement module to calibrate these ambiguous samples by regulating the distance between ambiguous samples and their corresponding prototypes. This calibration process aims to pull false negative (FN) samples closer to their respective prototypes and push false positive (FP) samples apart from their affiliated prototypes. In addition, we propose a new prototypical diversity amplification loss to strengthen the model's capacity by amplifying the differences between different prototypes. Finally, we propose a prototype-guided rectification to rectify prediction by incorporating the representability of prototypes. Extensive experiments conducted on the benchmark dataset demonstrate the superior performance of our method compared to existing approaches. Kun Li 0008, Dan Guo 0001, Chunxiao Fan 0002, Zhiliang Wu, Hehe Fan, Meng Wang 0001 |
AAAI | 6 |
| 2025 | MMAD: Multi-Label Micro-Action Detection in VideosabstractHuman body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenarios, human micro-actions often temporally co-occur, with multiple micro-actions overlapping in time, such as concurrent head and hand movements. However, current research primarily focuses on recognizing individual micro-actions while overlooking their co-occurring nature. To address this gap, we propose a new task named Multi-label Micro-Action Detection (MMAD), which involves identifying all micro-actions in a given short video, determining their start and end times, and categorizing them. Accomplishing this requires a model capable of accurately capturing both long-term and short-term action relationships to detect multiple overlapping micro-actions. To facilitate the MMAD task, we introduce a new dataset named Multi-label Micro-Action-52 (MMA-52) and propose a baseline method equipped with a dual-path spatial-temporal adapter to address the challenges of subtle visual change in MMAD. We hope that MMA-52 can stimulate research on micro-action analysis in videos and prompt the development of spatio-temporal modeling in human-centric video understanding. The proposed MMA-52 dataset is available at: https://github.com/VUT-HFUT/Micro-Action. Kun Li 0008, Pengyu Liu 0005, Dan Guo 0001, Fei Wang 0067, Zhiliang Wu, Hehe Fan, Meng Wang 0001 |
ICCV | 5 |
| 2025 | BVINet: Unlocking Blind Video Inpainting With Zero AnnotationsabstractVideo inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusing primarily on the "how to inpaint". This reliance necessitates manual annotation of the corrupted regions using binary masks to indicate "whereto inpaint". However, the annotation of these masks is labor-intensive and expensive, limiting the practicality of current methods. In this paper, we expect to relax this assumption by defining a new blind video inpainting setting, enabling the networks to learn the mapping from corrupted video to inpainted result directly, eliminating the need of corrupted region annotations. Specifically, we propose an end-to-end blind video inpainting network (BVINet) to address both "where to inpaint" and "how to inpaint" simultaneously. On the one hand, BVINet can predict the masks of corrupted regions by detecting semantic-discontinuous regions of the frame and utilizing temporal consistency prior of the video. On the other hand, the predicted masks are incorporated into the BVINet, allowing it to capture valid context information from uncorrupted regions to fill in corrupted ones. Besides, we introduce a consistency loss to regularize the training parameters of BVINet. In this way, mask prediction and video completion mutually constrain each other, thereby maximizing the overall performance of the trained model. Furthermore, we customize a dataset consisting of synthetic corrupted videos, real-world corrupted videos, and their corresponding completed videos. This dataset serves as a valuable resource for advancing blind video inpainting research. Extensive experimental results demonstrate the effectiveness and superiority of our method. Zhiliang Wu, Kerui Chen, Kun Li 0008, Hehe Fan, Yi Yang 0001 |
ICCV | 1 |
| 2025 | Prompt-Aware Controllable Shadow RemovalabstractShadow removal aims to restore the image content in shadowed regions. While deep learning-based methods have shown promising results, they still face key challenges: 1) uncontrolled removal of all shadows, or 2) controllable removal but heavily relies on precise shadow region masks. To address these issues, we introduce a novel paradigm: prompt-aware controllable shadow removal. Unlike existing approaches, our paradigm allows for targeted shadow removal from specific subjects based on user prompts (e.g., dots, lines, or subject masks). This approach eliminates the need for shadow annotations and offers flexible, user-controlled shadow removal. Specifically, we propose an end-to-end learnable model, the Prompt-Aware Controllable Shadow Removal Network (PACSRNet). PACSRNet consists of two key modules: a prompt-aware module that generates shadow masks for the specified subject based on the user prompt, and a shadow removal module that uses the shadow prior from the first module to restore the content in the shadowed areas. Additionally, we enhance the shadow removal module by incorporating feature information from the prompt-aware module through a linear operation, providing prompt-guided support for shadow removal. Recognizing that existing shadow removal datasets lack diverse user prompts, we contribute a new dataset specifically designed for prompt-based controllable shadow removal. Extensive experimental results demonstrate the effectiveness and superiority of PACSRNet. Kerui Chen, Zhiliang Wu, Wenjin Hou, Kun Li 0008, Hehe Fan, Yi Yang 0001 |
IJCAI | 2 |
| 2025 | Drafting and Revision: Advancing High-Fidelity Video InpaintingabstractVideo inpainting aims to fill the missing regions in video with spatial-temporally coherent contents. Existing methods usually treat the missing contents as a whole and adopt a hybrid objective containing a reconstruction loss and an adversarial loss to train the model. However, these two kinds of loss focus on contents at different frequencies, simply combining them may cause inter-frequency conflicts, leading the trained model to generate compromised results. Inspired by the common corrupted painting restoration process of “drawing a draft first and then revising the details later”, this paper proposes a Drafting-and-Revision Completion Network (DRCN) for video inpainting. Specifically, we first design a Drafting Network that utilizes the temporal information to complete the low-frequency semantic structure at low resolution. Then, a Revision Network is developed to hallucinate high-frequency details at high resolution by using the output of Drafting Network. In this way, adversarial loss and reconstruction loss can be applied to high-frequency and low-frequency respectively, effectively mitigating inter-frequency conflicts. Furthermore, Revision Network can be stacked in a pyramid manner to generate higher resolution details, which provide a feasible solution for high-resolution video inpainting. Experiments show that DRCN achieves improvements of 7.43% and 12.64% in E_warp and LPIPS, and can handle higher resolution videos on limited GPU memory. Zhiliang Wu, Kun Li 0008, Hehe Fan, Yi Yang 0001 |
IJCAI | 1 |
| 2025 | Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionabstractMicro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in Micro-Action Recognition often overlook the inherent subtle changes in MAs, which limits the accuracy of distinguishing MAs with subtle changes. To address this issue, we present a novel Motion-guided Modulation Network (MMN) that implicitly captures and modulates subtle motion cues to enhance spatial-temporal representation learning. Specifically, we introduce a Motion-guided Skeletal Modulation module (MSM) to inject motion cues at the skeletal level, acting as a control signal to guide spatial representation modeling. In parallel, we design a Motion-guided Temporal Modulation module (MTM) to incorporate motion information at the frame level, facilitating the modeling of holistic motion patterns in micro-actions. Finally, we propose a motion consistency learning strategy to aggregate the motion cues from multi-scale features for micro-action classification. Experimental results on the Micro-Action 52 and iMiGUE datasets demonstrate that MMN achieves state-of-the-art performance in skeleton-based micro-action recognition, underscoring the importance of explicitly modeling subtle motion cues. The code will be available at https://github.com/momiji-bit/MMN Jihao Gu, Kun Li 0008, Fei Wang 0073, Yanyan Wei, Zhiliang Wu, Hehe Fan, Meng Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Cross-Modality Coupled Prompt Learning for Zero-Shot Sketch-Based Image Retrieval
Caichu Luan, Zhiliang Wu |
PRCV (12) | 4 |
| 2025 | Perceptual Transform Fusion of Infrared and Visible ImagesabstractInfrared and visible image fusion aims to generate fused images with rich textures and clear target representations. Existing methods generally assume high-quality input images, thus overlooking issues such as reduced contrast and loss of details in visible images under low-light conditions. The naive enhance-then-fuse strategy cannot perform fuse-oriented image enhancement, which always reaches a sub-optimal result. To address this challenge, we propose a perceptual transform fusion of infrared and visible images, which simultaneously optimizes low-light enhancement and image fusion. Specifically, to improve computational efficiency and optimize key feature representations while suppressing noise interactions caused by lighting variations, we introduce a lightweight adaptive sparse Transformer block (ASTBlock). This model adaptively integrates sparse and dense attention mechanisms to enhance feature representations and employs a feed-forward network to eliminate redundant information, thereby ensuring the quality of image fusion. Subsequently, to retain significant details while reducing the impact of noise introduced by low-light enhancement, we incorporate discrete wavelet transform (DWT) for feature decomposition and fusion, further enhancing the representation capability and feature preservation of fused images. Meanwhile, to tackle the issues of insufficient contrast and hidden details in low-light conditions, we design an illumination perception module and an illumination consistency loss to improve the contrast and clarity of fused images. Experimental results on multiple public benchmark datasets for quality assessment and downstream tasks, e.g., pedestrian detection, demonstrate that our method significantly outperforms the state-of-the-art (SOTA) methods. The code is available at https://github.com/hinmouc/PIVFusion. Dingli Hua, Qingmao Chen, Zhiliang Wu, Yifan Zuo 0001, Wenying Wen, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Text-Guided Semantic Alignment Network With Spatial-Frequency Interaction for Infrared-Visible Image Fusion Under Extreme IlluminationabstractAlthough text-guided infrared-visible image fusion helps improve content understanding under extreme illumination, existing methods usually ignore semantic differences between textual and visual features, resulting in limited improvement. To address this challenge, we propose a Text-Guided Semantic Alignment Network, termed TSANet, for extreme-illumination infrared-visible image fusion. The network follows an encoder-decoder structure, with two image encoders, two text encoders, and one decoder. It uses a Semantic Alignment and Fusion (SAF) block to bridge the two image encoders in each layer. Specifically, the SAF block consists of two parallel Semantic Alignment (SA) modules, corresponding to the infrared and visible modalities, respectively, and a Spatial-Frequency Interaction (SFI) module. The SA module aligns the visual feature from the image encoder with its corresponding textual feature from the text encoder, to guide the network focus on key semantic regions of infrared and visible images. The SFI module aggregates the spatial and frequency information extracted from the modality-aligned features of two SA modules for complementary representation learning. The network progressively complements two image modalities by connecting the SAF blocks from top to down, and finally provides a visually pleasing fusion effect by feeding the output of the last block into the decoder. Recognizing that existing datasets lack illumination diversity, we contribute a new dataset specifically designed for extreme-illumination image fusion. Extensive experiments show the effectiveness and superiority of TSANet over seven state-of-the-art methods. The source code and dataset are available at https://github.com/WentaoLi-CV/TSANet. Guanghui Yue 0001, Cheng Zhao 0003, Zhiliang Wu, Tianwei Zhou, Qiuping Jiang, Runmin Cong |
IEEE Trans. Image Process. | 4 |
| 2025 | SNN-IoT: Efficient Partitioning and Enabling of Deep Spiking Neural Networks in IoT ServicesabstractSpiking Neural Networks (SNNs), due to their inherent biological plausibility and energy-saving characteristics, naturally align with the requirements of IoT services. However, current SNNs require a multi-layer structure to achieve effective applications across various fields. The multi-layer deep SNNs with massive model parameters demand computational resources, rendering them incompatible with resource-constrained IoT devices. To address this problem, in this work, a deep SNN partitioning framework called SNN-IoT is proposed to run complex SNN models on IoT devices. The SNN-IoT first partitions a full deep SNN model into smaller sub-models, leveraging the event-driven sparsity of SNNs and channel-level firing patterns to distribute filters with lower levels of spike activity onto devices with more constrained resources. The SNN model partitioning and deployment is formulated as an optimization problem and is solved using a greedy search assignment mechanism. Furthermore, a channel-wise pruning method exploits the varying degrees of channel activity, effectively reducing each sub-model's size and computational load without compromising performance. Extensive experiments conducted on four non-neuromorphic and two neuromorphic datasets have demonstrated that the SNN-IoT framework not only efficiently partitions deep SNNs and enables their deployment on IoT devices but also significantly reduces the inference latency and energy consumption for IoT services. The experiment uses 9 Raspberry Pi-4B as the IoT devices, and results show that SNN-IoT may reduce the average latency and energy consumption by about 60.7% and 49.9%, respectively, while maintaining the inference accuracy. Xin Du 0002, Wentao Tong, Linshan Jiang, Di Yu 0001, Zhiliang Wu, Qiang Duan 0002, Shuiguang Deng |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingabstractVideo inpainting aims to fill in the missing regions of the video frames with plausible content. Benefiting from the outstanding long-range modeling capacity, the transformer-based models have achieved unprecedented performance regarding inpainting quality. Essentially, coherent contents from all the frames along both spatial and temporal dimensions are concerned by a patch-wise attention module, and then the missing contents are generated based on the attention-weighted summation. In this way, attention retrieval accuracy has become the main bottleneck to improve the video inpainting performance, where the factors affecting attention calculation should be explored to maximize the advantages of transformer. Towards this end, in this paper, we theoretically certificate that noise is the culprit that entangles the process of attention calculation. Meanwhile, we propose a novel wavelet transformer network with noise robustness for video inpainting, named WaveFormer. Unlike existing transformer-based methods that utilize the whole embeddings to calculate the attention, our WaveFormer first separates the noise existing in the embedding into high-frequency components by introducing the Discrete Wavelet Transform (DWT), and then adopts clean low-frequency components to calculate the attention. In this way, the impact of noise on attention computation can be greatly mitigated and the missing content regarding different frequencies can be generated by sharing the calculated attention. Extensive experiments validate the superior performance of our method over state-of-the-art baselines both qualitatively and quantitatively. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu, Yan Yan 0002 |
AAAI | 1 |
| 2024 | Text-Video Completion Networks With Motion Compensation And Attention AggregationabstractThe purpose of video inpainting is to fill a specified area with reasonable content. However, in the case of multiple targets and complex textures, current methods struggle to distinguish between feature information of the targets, leading to confusing or fuzzy inpainting results. In this paper, we design a new text-video completion network based on a motion compensation and temporal attention feature aggregation. Our network utilizes information from reference frames and target frames to complete the damaged region of the target frame. We first employ motion compensation to align the features of reference frames, and then use the temporal attention module to aggregate these features, resulting in accurate and reasonable content. To evaluate the effectiveness of our method, we introduce a new text video dataset with multiple text objects and complex textures, presenting a novel and challenging task for inpainting research. Through quantitative and qualitative comparison experiments, we demonstrate that our model outperforms existing baseline models in scenarios with multiple objects and complex textures. Zhiliang Wu, Hanyu Xuan, Yan Yan 0002 |
ICASSP | 2 |
| 2024 | Robust Audio-Visual Contrastive Learning for Proposal-Based Self-Supervised Sound Source Localization in VideosabstractBy observing a scene and listening to corresponding audio cues, humans can easily recognize where the sound is. To achieve such cross-modal perception on machines, existing methods take advantage of the maps obtained by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for prospective practical applications, we argue that these map-based methods only offer a coarse-grained and indirect description of the sound source. Additionally, these methods utilize a single audio-visual tuple at a time during self-supervised learning, causing the model to lose the crucial chance to reason about the data distribution of large-scale audio-visual samples. Although the introduction of Audio-Visual Contrastive Learning (AVCL) can effectively alleviate this issue, the contrastive set constructed by randomly sampling is based on the assumption that the audio and visual segments from all other videos are not semantically related. Since the resulting contrastive set contains a large number of faulty negatives, we believe that this assumption is rough. In this paper, we advocate a novel proposal-based solution that directly localizes the semantic object-level sound source, without any manual annotations. The Global Response Map (GRM) is incorporated as an unsupervised spatial constraint to filter those instances corresponding to a large number of sound-unrelated regions. As a result, our proposal-based Sound Source Localization (SSL) can be cast into a simpler Multiple Instance Learning (MIL) problem. To overcome the limitation of random sampling in AVCL, we propose a novel Active Contrastive Set Mining (ACSM) to mine the contrastive sets with informative and diverse negatives for robust AVCL. Our approaches achieve state-of-the-art (SOTA) performance when compared to several baselines on multiple SSL datasets with diverse scenarios. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Bo Jiang 0002, Lei Luo 0001, Xavier Alameda-Pineda, Yan Yan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Deep Stereo Video InpaintingabstractStereo video inpainting aims to fill the missing regions on the left and right views of the stereo video with plausible content simultaneously. Compared with the single video inpainting that has achieved promising results using deep convolutional neural networks, inpainting the missing regions of stereo video has not been thoroughly explored. In essence, apart from the spatial and temporal consistency that single video inpainting needs to achieve, another key challenge for stereo video inpainting is to maintain the stereo consistency between left and right views and hence alleviate the 3D fatigue for viewers. In this paper, we propose a novel deep stereo video inpainting network named SVINet, which is the first attempt for stereo video inpainting task utilizing deep convolutional neural networks. SVINet first utilizes a self-supervised flow-guided deformable temporal alignment module to align the features on the left and right view branches, respectively. Then, the aligned features are fed into a shared adaptive feature aggregation module to generate missing contents of their respective branches. Finally, the parallax attention module (PAM) that uses the cross-view information to consider the significant stereo correlation is introduced to fuse the completed features of left and right views. Furthermore, we develop a stereo consistency loss to regularize the trained parameters, so that our model is able to yield high-quality stereo video inpainting results with better stereo consistency. Experimental results demonstrate that our SVINet outperforms state-of-the-art single video inpainting models. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
CVPR | 1 |
| 2023 | Semi-Supervised Video Inpainting with Cycle Consistency ConstraintsabstractDeep learning-based video inpainting has yielded promising results and gained increasing attention from re-searchers. Generally, these methods assume that the cor-rupted region masks of each frame are known and easily ob-tained. However, the annotation of these masks are labor-intensive and expensive, which limits the practical application of current methods. Therefore, we expect to relax this assumption by defining a new semi-supervised inpainting setting, making the networks have the ability of completing the corrupted regions of the whole video using the anno-tated mask of only one frame. Specifically, in this work, we propose an end-to-end trainable framework consisting of completion network and mask prediction network, which are designed to generate corrupted contents of the current frame using the known mask and decide the regions to be filled of the next frame, respectively. Besides, we introduce a cycle consistency loss to regularize the training parameters of these two networks. In this way, the completion network and the mask prediction network can constrain each other, and hence the overall performance of the trained model can be maximized. Furthermore, due to the natural existence of prior knowledge (e.g., corrupted contents and clear bor-ders), current video inpainting datasets are not suitable in the context of semi-supervised video inpainting. Thus, we create a new dataset by simulating the corrupted video of real-world scenarios. Extensive experimental results are reported to demonstrate the superiority of our model in the video inpainting task. Remarkably, although our model is trained in a semi-supervised manner, it can achieve compa-rable performance as fully-supervised methods. Zhiliang Wu, Hanyu Xuan, Changchang Sun, Weili Guan, Yan Yan 0002 |
CVPR | 1 |
| 2023 | Flow-Guided Deformable Alignment Network with Self-Supervision for Video InpaintingabstractVideo inpainting aims to utilize plausible contents to fill missing regions in the video. State-of-the-art video inpainting methods typically generate the missing contents of the target frame (current frame) by aggregating the temporal information of reference frames (neighboring frames) aligned using deformable convolution. However, these deformable convolution alignment networks often suffer from offset overflow during training, resulting in unsatisfactory alignment, thereby obtaining compromised inpainting performance. In this paper, we propose a self-supervised Flow-Guided Deformable Alignment (FGDA) network for aligning reference frames at the feature level. FGDA computes the residual of the optical flow as the offsets. This design can effectively reduce the burden of offsets learning, thereby avoiding offset overflow. Furthermore, a gradient-weighted reconstruction loss for supervised completed frame reconstruction is designed, which can use the gradients in all directions of the video frame to emphasize the difficultly reconstructed texture regions, so that detail textures get more attention during training. Experiments show that FGDA-based video inpainting model trained with gradient-weighted reconstruction loss outperforms the state-of-the-art by a significant margin in terms of PSNR and SSIM with relative improvements of 6.2% and 2.1%, respectively. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
ICASSP | 1 |
| 2023 | PFTA-Net: Progressive Feature Alignment and Temporal Attention Fusion Networks for Video InpaintingabstractThe goal of video inpainting is to fill in missing regions with reasonable and coherent content in a video sequence. Due to the motion of cameras and objects, the reference frame and the target frame are not aligned, and the useful information of the reference frame cannot be well utilized. Therefore, temporal alignment plays an important role in video inpainting. Some studies have attempted to divide the remote alignment into multiple sub-alignments and process them step by step, but error accumulation is inevitable. In this paper, we present a novel progressive feature alignment and temporal attention fusion network, namely PFTA-Net. Specifically, we design a progressive feature alignment module, which employs sub-alignments with a progressive refinement scheme, resulting in more accurate motion compensation. After alignment, we propose a temporal attention fusion module, which computes temporal attention weights for each aligned reference frame feature, resulting in modulated features for reconstructing the target frame. Our extensive evaluations, including both quantitative and qualitative assessments, demonstrate the better performance and efficacy of our video inpainting network. Zhiliang Wu, Yan Yan 0002 |
ICIP | 2 |
| 2023 | Semantic-Guided Completion Network for Video Inpainting in Complex Urban Scene
Hanyu Xuan, Zhiliang Wu |
PRCV (11) | 3 |
| 2023 | Divide-and-conquer model based on wavelet domain for multi-focus image fusion
Zhiliang Wu, Hanyu Xuan, Xia Yuan, Chunxia Zhao |
Signal Process. Image Commun. | 1 |
| 2023 | Divide-and-Conquer Completion Network for Video InpaintingabstractVideo inpainting aims to utilize plausible contents to complete missing regions in the video. For different components, the reconstruction targets of missing regions are different,e.g.,smoothness preserving for flat regions, sharpening for edges and textures. Typically, existing methods treat the missing regions as a whole and holistically train the model by optimizing homogenous pixel-wise losses (e.g.,MSE). In this way, the trained models will be easily dominated and determined by flat regions, failing to infer realistic details (edges and textures) that are difficult to reconstruct but necessary for practical applications. In this paper, we propose a divide-and-conquer completion network for video inpainting. In particular, our network first uses discrete wavelet transform to decompose the deep features into low-frequency components containing structural information (flat regions) and high-frequency components involving detailed texture information. Thereafter, we feed these components into different branches and adopt the temporal attention feature aggregation module to generate missing contents, separately. It hence can realize flexible supervision utilizing the intermediate supervision learning strategy for each component, which has not been noticed and explored by current state-of-the-art video inpainting methods. Furthermore, we adopt a gradient-weighted reconstruction loss to supervise the completed frame reconstruction process, which can use the gradients in all directions of the video frame to emphasize the difficultly reconstructed textures regions, making the model pay more attention to the complex detailed textures. Extensive experiments validate the superior performance of our divide-and-conquer model over state-of-the-art baselines in both quantitative and qualitative evaluations. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Fractional Optimization Model for Infrared and Visible Image Fusion
Zhiliang Wu, Xia Yuan, Chunxia Zhao |
BMVC | 3 |
| 2022 | A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosabstractHumans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for potential practical applications, we argue that these existing map-based approaches only provide a coarse-grained and indirect description of the sound source. In this pa-per, we advocate a novel proposal-based paradigm that can directly perform semantic object-level localization, without any manual annotations. We incorporate the global re-sponse map as an unsupervised spatial constraint to weight the proposals according to how well they cover the esti-mated global shape of the sound source. As a result, our proposal-based sound source localization can be cast into a simpler Multiple Instance Learning (MIL) problem by filtering those instances corresponding to large sound-unrelated regions. Our method achieves state-of-the-art (SOTA) per-formance when compared to several baselines on multiple datasets. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Yan Yan 0002, Xavier Alameda-Pineda |
CVPR | 2 |
| 2022 | Active Contrastive Set Mining for Robust Audio-Visual Instance DiscriminationabstractThe recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual Instance Discrimination (AVID) extends instance discrimination to the audio-visual realm. Existing AVID methods construct the contrastive set by random sampling based on the assumption that the audio and visual clips from all other videos are not semantically related. We argue that this assumption is rough, since the resulting contrastive sets have a large number of faulty negatives. In this paper, we overcome this limitation by proposing a novel Active Contrastive Set Mining (ACSM) that aims to mine the contrastive sets with informative and diverse negatives for robust AVID. Moreover, we also integrate a semantically-aware hard-sample mining strategy into our ACSM. The proposed ACSM is implemented into two most recent state-of-the-art AVID methods and significantly improves their performance. Extensive experiments conducted on both action and sound recognition on multiple datasets show the remarkably improved performance of our method. Hanyu Xuan, Shuo Chen 0002, Zhiliang Wu, Jian Yang 0003, Yan Yan 0002, Xavier Alameda-Pineda |
IJCAI | 4 |
| 2022 | CFNet: Context fusion network for multi-focus imagesabstractAbstract Multi‐focus image fusion aims to generate a clear image by fusing multiple source images. Existing deep learning‐based fusion methods often neglect the context information resulting in the loss of detail information. To address this issue, a context fusion network to merge multi‐focus images, namely CFNet, is proposed. Specifically, a context fusion module is proposed to make full use of low‐level pixels and high‐level semantic features. Particularly, the pyramid fusion mechanism and cross‐scale transfer strategy are adopted to ensure the visual and semantic consistency of the fused image. Meanwhile, to extract salient features more effectively, a spatial attention mechanism is introduced to enhance these features. Further, the pyramid loss is used to progressively refine the fused features at each scale. Experimental results show that the proposed method is superior to some existing methods in both qualitative and quantitative evaluation. Zhiliang Wu, Xia Yuan, Chunxia Zhao |
IET Image Process. | 2 |
| 2021 | DAPC-Net: Deformable Alignment and Pyramid Context Completion Networks for Video InpaintingabstractVideo inpainting aims to fill missing regions with plausible content in a video sequence. Deep learning-based video inpainting methods have made promising progress over the past few years. However, these methods tend to generate degraded completion content, such as missing textural details. To address this issue, we propose a novel Deformable Alignment and Pyramid-context Completion Network for video inpainting (DAPC-Net), which takes advantage of temporal redundancy information among video sequence. Specifically, we construct a deformable convolution alignment network (DANet) for aligning reference frame at the feature level. After alignment, we further devise a pyramid-context completion network (PCNet) to complete missing regions of the target frame. Particularly, the pyramid completion mechanism and cross-scale transference strategy are used to ensure the visual and semantic coherence of the completed target frame. Experimental results show that the proposed method not only achieves better quantitative and qualitative performance but also improves the inference speed by 35.4%. Zhiliang Wu, Hanyu Xuan, Jian Yang 0003 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Introspective Learning by Distilling Knowledge from Online Self-explanation
Jindong Gu, Zhiliang Wu, Volker Tresp |
ACCV (4) | 2 |
| 2017 | FPGA-Based N×N adaptive channel for array lidar imagerabstractAt present, APD (Avalanche Photo Diode) arrays LIDAR (Light Detection and Ranging) has been broadly accepted as an important means to obtain 3D (Three Dimensional) data. A new method of a scalable adaptive N × N channel communication system is introduced in the paper, which is aimed at improving the transmission rate of APD array LIDAR data. This paper presents the research on multi-channel communication system based on FPGA (Field Programmable Gate Arrays) configuration, which can be realized multi-channel parallel data independent receiving and transmitting. The multi-channel transmission system not only improves the data bandwidth, but also improve the load capacity of the system. The feasibility of multi-channel communication is verified sufficiently in the paper. Guoqing Zhou 0001, Pengyun Chen, Xiang Zhou 0002, Lieping Zhang, Guoqing Gao, Yajun Fan, Zhiliang Wu, Jingjin Huang |
IGARSS | 8 |
| 2017 | Carbon sink estimation of surface carbonate karstification in global karst areaabstractThe lithosphere is the most stored reservoir of carbon on Earth. The study of the effects of rock weathering and in the atmosphere provides an important basis for accurate prediction of CO2in the atmosphere, which is essential for predicting global climate change. One of the main factors that has been neglected in the global carbon cycle study of carbonate weathering carbon sinks is the rate problem. In this paper, some published data were used to study the influence factors of karst carbon sink and the establishment of corrosion velocity model. Based on the data of 18 corrosion test sites in China, a corrosion rate model was fitted, and then the data predicted by the model were compared with the data provided paper of Liu Rui, which are close to the measured data. The validated model can provide a research method for the carbonate certification rate in different regions, and solve the problems that cannot be tested because of time and space constraints. At last, net sink of CO2in the atmosphere is calculated as 4.77×1014g·a−1(or 130 million tC·a−1) in the global surface carbonate certification, which is close to the datameasured by Liu Zaihua and Ichikuni. Guoqing Zhou 0001, Guoqing Gao, Zhiliang Wu, Yajun Fan, Pengyun Chen, Jingjin Huang |
IGARSS | 4 |
| 2012 | Vessel Steering Control Using Generalized Ellipsoidal Basis Function Based Fuzzy Neural Networks
Ning Wang 0002, Zhiliang Wu, Chidong Qiu, Tieshan Li 0001 |
ISNN (2) | 2 |