Yueyi Zhang 0001

dblp:124/7067-1 · DBLP profile ↗
← Back
90ranked-venue papers
5as first author
74since 2021 · last 2026
0000-0003-0788-8826ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 75 · 4 first-author · 60 since 2021Artificial intelligence and machine learning · 49 · 1 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021
YearPublicationVenuePosition
2026 Seeing the Unseen: Zooming in the Dark with Event Cameras
abstract
This paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these challenges, we present RetinexEVSR, the first event-driven LVSR framework that leverages high-contrast event signals and Retinex-inspired priors to enhance video quality under low-light scenarios. Unlike previous approaches that directly fuse degraded signals, RetinexEVSR introduces a novel bidirectional cross-modal fusion strategy to extract and integrate meaningful cues from noisy event data and degraded RGB frames. Specifically, an illumination-guided event enhancement module is designed to progressively refine event features using illumination maps derived from the Retinex model, thereby suppressing low-light artifacts while preserving high-contrast details. Furthermore, we propose an event-guided reflectance enhancement module that utilizes the enhanced event features to dynamically recover reflectance details via a multi-scale fusion mechanism. Experimental results show that our RetinexEVSR achieves state-of-the-art performance on three datasets. Notably, on the SDSD benchmark, our method can get up to 2.95 dB gain while reducing runtime by 65% compared to prior event-based methods.
Dachun Kai, Zeyu Xiao 0002, Huyue Zhu, Jiaxiao Wang, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI5
2026 EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
abstract
Event-based vision has drawn increasing attention owing to its distinctive properties, including ultra-high temporal resolution and extreme dynamic range. Recent works have introduced it to video super-resolution (VSR) to enhance flow estimation and temporal alignment. In contrast, this paper shifts the focus of event signals from motion refinement to texture enhancement in VSR. We propose EvTexture++, the first event-driven framework dedicated to texture enhancement in VSR. It leverages high-frequency spatiotemporal details from events to improve texture recovery. EvTexture++ incorporates a customized texture enhancement branch, along with an iterative texture enhancement module that progressively exploits high-temporal-resolution event information for texture restoration. This enables gradual refinement of texture regions across iterations, yielding more accurate and detailed high-resolution outputs. Besides intra-frame texture recovery, large motions could degrade inter-frame temporal consistency, particularly in texture regions, leading to texture flickering. To mitigate this, we further exploit the continuous-time motion cues of events to enhance temporal consistency, introducing a temporal texture alignment module that estimates event-guided texture-aware flow for precise inter-frame texture alignment. Moreover, EvTexture++ is designed as a plug-and-play tool to flexibly boost the performance of existing VSR models. Experiments on five datasets demonstrate that EvTexture++ achieves state-of-the-art performance. When integrated into recent VSR models, it yields significant improvements, with gains of up to 1.55 dB in PSNR on the texture-rich Vid4 dataset.
Dachun Kai, Jiayao Lu, Yueyi Zhang 0001, Xiaoyan Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 E2SL: Efficient Depth Sensing from Event-Based Structured Light
abstract
Structured light (SL) is a popular approach for 3D reconstruction. Most SL techniques rely on frame-based cameras and are often not robust in high-speed dynamic scenes. Recently, event cameras have sparked growing interest in high-speed SL imaging, due to their high temporal resolution. The event-based SL enjoys the high-speed data acquisition, however, most existing methods tend to pursue the reconstruction accuracy but sacrificing the computational efficiency, limiting the applicability in real-world scenarios. To this end, we propose E2SL, an Efficient deep network tailored for monocular Event-based SL. Specifically, E2SL comprises three key components: binary embedding lookup table (BE-LUT), spatial context enhancement (SCE), and geometric-prior regression (GPR). Given the input event frame, BE-LUT, which is precomputed and stored, first retrieves the features efficiently. Then, SCE extends the receptive field of the features and captures the spatial context. Finally, GPR conducts the geometric-prior-based tree classification for fast and robust depth estimation. To support training and evaluation, we contribute an event-based SL simulator, which generates a large-scale and diverse synthetic dataset. Besides, we develop an event-based SL prototype and collect a dataset with accurate ground truth for real-world evaluation. Extensive experiments demonstrate that our method achieves state-of-the-art accuracy while maintaining a per-frame reconstruction time of 7.7 ms, meeting the demands of high-speed depth sensing. The code and dataset are available on the project page https://dongxin000.github.io/E2SL/.
Jiacheng Fu, Wenming Weng, Yueyi Zhang 0001, Bingyao Huang, Zhiwei Xiong
IEEE Trans. Vis. Comput. Graph.5
2025 Event-Enhanced Blurry Video Super-Resolution
abstract
In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net.
Dachun Kai, Yueyi Zhang 0001, Jin Wang 0023, Zeyu Xiao 0002, Zhiwei Xiong, Xiaoyan Sun 0001
AAAI2
2025 Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach
abstract
Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events, resulting in complex training strategies and increased computational costs. To address these challenges, we propose an efficient hybrid framework for image semantic segmentation, comprising a Spiking Neural Network branch for events and an Artificial Neural Network branch for frames. Specifically, we introduce three specialized modules to facilitate the interaction between these two branches: the Adaptive Temporal Weighting (ATW) Injector, the Event-Driven Sparse (EDS) Injector, and the Channel Selection Fusion (CSF) module. The ATW Injector dynamically integrates temporal features from event data into frame features, enhancing segmentation accuracy by leveraging critical dynamic temporal information. The EDS Injector effectively combines sparse event data with rich frame features, ensuring precise temporal and spatial information alignment. The CSF module selectively merges these features to optimize segmentation performance. Experimental results demonstrate that our framework not only achieves state-of-the-art accuracy across the DDD17-Seg, DSEC-Semantic, and M3ED-Semantic datasets but also significantly reduces energy consumption, achieving a 65% reduction on the DSEC-Semantic dataset.
Hebei Li, Yansong Peng, Jiahui Yuan, Peixi Wu, Jin Wang 0023, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI6
2025 Spiking Point Transformer for Point Cloud Classification
abstract
Spiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud remains underexplored. To this end, we present Spiking Point Transformer (SPT), the first transformer-based SNN framework for point cloud classification. Specifically, we first design Queue-Driven Sampling Direct Encoding for point cloud to reduce computational costs while retaining the most effective support points at each time step. We introduce the Hybrid Dynamics Integrate-and-Fire Neuron (HD-IF), designed to simulate selective neuron activation and reduce over-reliance on specific artificial neurons. SPT attains state-of-the-art results on three benchmark datasets that span both real-world and synthetic datasets in the SNN domain. Meanwhile, the theoretical energy consumption of SPT is at least 6.4x less than its ANN counterpart.
Peixi Wu, Bosong Chai, Hebei Li, Menghua Zheng, Yansong Peng, Xuan Nie, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI8
2025 Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model
abstract
Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of available modalities for brain tumor segmentation. LS3M excels at efficiently modeling long-range dependencies based on the Mamba design, while incorporating differentiable permutation matrices that reorder input sequences based on modality-specific characteristics. This dynamic reordering ensures that critical spatial inductive biases and long-range semantic correlations inherent in 3D brain MRI are preserved, which is crucial for imcomplete multi-modal brain tumor segmentation. Once the input sequences are reordered using the generated permutation matrix, the Series State Space Model (S3M) block models the relationships between them, capturing both local and long-range dependencies. This enables effective representation of intra-modal and inter-modal relationships, significantly improving segmentation accuracy. Extensive experiments on the BraTS2018 and BraTS2020 datasets demonstrate that LS3M outperforms existing methods, offering a robust solution for brain tumor segmentation, particularly in scenarios with missing modalities.
Zheyu Zhang 0002, Yayuan Lu, Feipeng Ma, Yueyi Zhang 0001, Huanjing Yue, Xiaoyan Sun 0001
CVPR4
2025 S2D-LFE: Sparse-to-Dense Light Field Event Generation
abstract
In this paper, we present S2D-LFE, an innovative approach for sparse-to-dense light field event generation. For the first time to our knowledge, S2D-LFE enables controllable novel view synthesis only from sparse-view light field event (LFE) data, and addresses three critical challenges for the LFE generation task: simplicity, controllability, and consistency. The simplicity aspect eliminates the dependency on frame-based modality, which often suffers from motion blur and low frame-rate limitations. The controllability aspect enables precise view synthesis under sparse LFE conditions with view-related constraints. The consistency aspect ensures both cross-view and temporal coherence in the generated results. To realize S2D-LFE, we develop a novel diffusion-based generation network with two key components. First, we design an LFE-customized variational auto-encoder that effectively compresses and reconstructs LFE by integrating cross-view information. Second, we design an LFE-aware injection adaptor to extract comprehensive geometric and texture priors. Furthermore, we construct a large-scale synthetic LFE dataset containing 162 one-minute sequences using simulator, and capture a real-world testset using our custom-built sparse LFE acquisition system, covering diverse indoor and outdoor scenes. Extensive experiments demonstrate that S2D-LFE successfully generates up to 9 × 9 dense LFE from 2 × 2 sparse inputs and outperforms existing methods on both synthetic and real-world data. The datasets and code are available at https://github.com/Yutong2022/S2D-LFE.
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
CVPR3
2025 Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
abstract
Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce $\textbf{VLADBench}$, a challenging and fine-grained dataset featuring close-form QAs that progress from static foundational knowledge and elements to advanced reasoning for dynamic on-road situations. The elaborate $\textbf{VLADBench}$ spans 5 key domains: Traffic Knowledge Understanding, General Element Recognition, Traffic Graph Generation, Target Attribute Comprehension, and Ego Decision-Making and Planning. These domains are further broken down into 11 secondary aspects and 29 tertiary tasks for a granular evaluation. A thorough assessment of general and domain-specific (DS) VLMs on this benchmark reveals both their strengths and critical limitations in AD contexts. To further exploit the cognitive and reasoning interactions among the 5 domains for AD understanding, we start from a small-scale VLM and train the DS models on individual domain datasets (collected from 1.4M DS QAs across public sources). The experimental results demonstrate that the proposed benchmark provides a crucial step toward a more comprehensive assessment of VLMs in AD, paving the way for the development of more cognitively sophisticated and reasoning-capable AD systems.
Zhenyu Lin, Jiangtong Zhu, Dechang Zhu, Yueyi Zhang 0001, Zhiwei Xiong, Xinhai Zhao
ICCV7
2025 GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud Sequences
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
ICCV3
2025 Generalizable Non-Line-of-Sight Imaging with Learnable Physical Priors
abstract
Non-line-of-sight (NLOS) imaging, recovering the hidden volume from indirect reflections, has attracted increasing attention due to its potential applications. Despite promising results, existing NLOS reconstruction approaches are constrained by the reliance on empirical physical priors, e.g., single fixed path compensation. Moreover, these approaches still possess limited generalization ability, particularly when dealing with scenes at a low signal-to-noise ratio (SNR). To overcome the above problems, we introduce a novel learning-based solution, comprising two key designs: Learnable Path Compensation (LPC) and Adaptive Phasor Field (APF). The LPC applies tailored path compensation coefficients to adapt to different objects in the scene, effectively reducing light wave attenuation, especially in distant regions. Meanwhile, the APF learns the precise Gaussian window of the illumination function for the phasor field, dynamically selecting the relevant spectrum band of the transient measurement. Experimental validations demonstrate that our proposed approach, only trained on synthetic data, exhibits the capability to seamlessly generalize across various real-world datasets captured by different imaging systems and characterized by low SNRs.
Shida Sun, Yueyi Zhang 0001, Zhiwei Xiong
ICCV3
2025 Event-Boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction
Wenming Weng, Yueyi Zhang 0001, Ruikang Xu, Zhiwei Xiong
ICCV3
2025 D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement
abstract
We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and models: https://github.com/Peterande/D-FINE.
Yansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0001
ICLR4
2025 Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects
abstract
Diffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our knowledge, this is a pioneering work enabling users to "create anything anywhere".
Hebei Li, Yansong Peng, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICME5
2025 Event-based HDR Structured Light
abstract
Event-based structured light (SL) systems have attracted increasing attention for their potential in high-performance 3D measurement. Despite the inherent HDR capability of event cameras, reflective and absorptive surfaces still cause event cluttering and absence, which produce overexposed and underexposed regions that degrade the reconstruction quality. In this work, we present the first HDR 3D measurement framework specifically designed for event-based SL systems. First, we introduce a multi-contrast HDR coding strategy that facilitates imaging of areas with different reflectance. Second, to alleviate inter-frame interference caused by overexposed and underexposed areas, we propose a universal confidence-driven stereo matching strategy. Specifically, we estimate a confidence map as the fusion weight for features via an energy-guided confidence estimation. Further, we propose the confidence propagation volume, an innovative cost volume that offers both effective suppression of inter-frame interference and strong representation capability. Third, we contribute an event-based SL simulator and propose the first event-based HDR SL dataset. We also collect a real-world benchmarking dataset with ground truth. We validate the effectiveness of our method with the proposed confidence-driven strategy on both synthetic and real-world datasets. Experimental results demonstrate that our proposed HDR framework enables accurate 3D measurement even under extreme conditions.
Jiacheng Fu, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
NeurIPS5
2025 Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition
Yucheng Zhan, Yijun Pan, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001
Neurocomputing4
2025 Graph relation distillation for efficient biomedical instance segmentation
Xiaoyu Liu 0006, Yueyi Zhang 0001, Zhiwei Xiong, Wei Huang 0036, Bo Hu 0014, Xiaoyan Sun 0001, Feng Wu 0001
Pattern Recognit.2
2025 Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action Recognition
abstract
Fine-grained action recognition typically faces challenges with lower inter-class variances and higher intra-class variances. Supervised contrastive learning is inherently suitable for this task, as it can decrease intra-class feature distances while increasing inter-class ones. However, directly applying it into fine-grained action recognition encounters two main problems. The first problem stems from the heavy training cost associated with supervised contrastive learning, which requires numerous training epochs, each involving double augmentation views per instance. To address this issue, we propose the late-stage supervised contrastive learning (late-SC) strategy, which effectively reduces the number of training epochs needed for the contrastive learning process. The second problem is that supervised contrastive loss does not explicitly consider the semantic distances between fine-grained actions when adjusting representation distances. This results in less reasonable and efficient adjustments to the representation space. To overcome this limitation, we introduce the semantic-aware temperature adaptation (STA) mechanism, enhancing the suitability of the supervised contrastive loss for fine-grained action recognition. We conduct experiments on several benchmark datasets for fine-grained action recognition, including Epic-Kitchens-55/100, SomethingSomething-V1, and Diving48-V2. The results demonstrate that our proposed method (referred to as LSC-STA) consistently enhances performance across various base feature extractors, without introducing additional inference overhead and incurring only a marginal increase in training expenses.
Yijun Pan, Yueyi Zhang 0001, Zilei Wang, Xiaoyan Sun 0001, Feng Wu 0005
IEEE Trans. Circuits Syst. Video Technol.3
2025 BVSR-EvD: Blurry Video Space-Time Super-Resolution With Events via Diffusion Models
abstract
Video restoration from low-resolution and low-frame-rate blurry sources remains challenging due to insufficient data priors. In this paper, we propose BVSR-EvD, leveraging event cameras and diffusion models to boost blurry video space-time super-resolution. Specifically, we identify three distinct data priors from event-video dual modalities: motion prior from events, content prior from videos, and physical prior from their integration, contributing to temporal stability, content preservation, and detail enhancement respectively. To effectively utilize these data priors, BVSR-EvD creates the Trident Diffusion Model (Trident-DM), which decomposes each denoising step into trident decoupling and adaptive self-composition stages. The former employs single-modal and dual-modal meta-networks to extract the three unique data priors, while the latter dynamically integrates them through learned prior-aware weight maps. BVSR-EvD achieves up to $\times 8$ spatial super-resolution and $\times 64$ temporal super-resolution from blurry videos, surpassing existing methods on public video datasets.
Wenming Weng, Yueyi Zhang 0001, Zeyu Xiao 0002, Zhiwei Xiong
IEEE Trans. Image Process.2
2025 EGVD: Event-Guided Video Deraining
abstract
Recent research has explored leveraging event cameras, known for their prowess in capturing scenes with nonuniform motion, for video deraining, leading to performance improvements. However, the existing event-based method still faces the challenge that the complex spatiotemporal distribution disrupts temporal information fusion and complicates feature separation. This article proposes a novel end-to-end learning framework for video deraining that effectively extracts the rich dynamic information provided by the event stream. Our framework incorporates two key modules: an event-aware motion detection (EAMD) module that adaptively aggregates multiframe motion information using event-driven masks and a pyramidal adaptive selection module that separates background and rain layers by leveraging contextual priors from both event and conventional camera data. To facilitate efficient training, we introduce a real-world dataset of synchronized rainy videos and event streams. Extensive evaluations on both synthetic and real-world datasets demonstrate the superiority of our proposed method compared to state-of-the-art approaches. The code is available at https://github.com/booker-max/EGVD.
Yueyi Zhang 0001, Jin Wang 0023, Wenming Weng, Xiaoyan Sun 0001, Zhiwei Xiong
IEEE Trans. Neural Networks Learn. Syst.1
2024 TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities
abstract
Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes for missing modalities, inadvertently introducing feature bias and redundant computations. To address these issues, we present the Token Merging transFormer (TMFormer) for robust brain tumor segmentation with missing modalities. TMFormer tackles these challenges by extracting and merging accessible modalities into more compact token sequences. The architecture comprises two core components: the Uni-modal Token Merging Block (UMB) and the Multi-modal Token Merging Block (MMB). The UMB enhances individual modality representation by adaptively consolidating spatially redundant tokens within and outside tumor-related regions, thereby refining token sequences for augmented representational capacity. Meanwhile, the MMB mitigates multi-modal feature fusion bias, exclusively leveraging tokens from present modalities and merging them into a unified multi-modal representation to accommodate varying modality combinations. Extensive experimental results on the BraTS 2018 and 2020 datasets demonstrate the superiority and efficacy of TMFormer compared to state-of-the-art methods when dealing with missing modalities.
Zheyu Zhang 0002, Yueyi Zhang 0001, Huanjing Yue, Aiping Liu, Yunwei Ou, Xiaoyan Sun 0001
AAAI3
2024 Image Captioning with Multi-Context Synthetic Data
abstract
Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This potential can be harnessed to create synthetic image-text pairs for training captioning models. Synthetic data can improve cost and time efficiency in data collection, allow for customization to specific domains, bootstrap generalization capability for zero-shot performance, and circumvent privacy concerns associated with real-world data. However, existing methods struggle to attain satisfactory performance solely through synthetic data. We identify the issue as generated images from simple descriptions mostly capture a solitary perspective with limited context, failing to align with the intricate scenes prevalent in real-world imagery. To tackle this, we present an innovative pipeline that introduces multi-context data generation. Beginning with an initial text corpus, our approach employs a large language model to extract multiple sentences portraying the same scene from diverse viewpoints. These sentences are then condensed into a single sentence with multiple contexts. Subsequently, we generate intricate images using the condensed captions through diffusion models. Our model is exclusively trained on synthetic image-text pairs crafted through this process. The effectiveness of our pipeline is validated through experimental results in both the in-domain and cross-domain settings, where it achieves state-of-the-art performance on well-known datasets such as MSCOCO, Flickr30k, and NoCaps.
Feipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI4
2024 Multi-modal Diffusion Network with Controllable Variability for Medical Image Segmentation
abstract
In diffusion-based medical segmentation models, stochastic sampling is commonly used to generate multiple masks. However, the inherent variability in diffusion models can lead to significant biases in some masks, resulting in the fused mask deviating from the true mask. In this study, we propose a novel multi-modal diffusion segmentation network (MMDSN) with controllable variability, specifically designed to address the issue of variability in diffusion models. MMDSN achieves multi-modal conditional control through medical text annotations, thereby enhancing consistency of visual semantic representation and establishing a correspondence between vision and language for diffusion models. Additionally, MMDSN constrains the uncertainty distributions of multiple timesteps within the latent Gaussian space, controlling the variability at each denoising timestep. Extensive experiments on the Qata-Covid19 and MosMed datasets demonstrate that our proposed method surpasses existing state-of-the-art diffusion networks, producing a high-quality, controllable segmentation map with just a single reverse diffusion step and one sampling.
Zheyu Zhang 0002, Yueyi Zhang 0001, Jing Zhang 0165, Yunwei Ou, Xiaoyan Sun 0001
BIBM3
2024 Event-Assisted Low-Light Video Object Segmentation
abstract
In the realm of video object segmentation (VOS), the challenge of operating under low-light conditions persists, resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras, characterized by their high dynamic range and ability to capture motion information of objects, offer promise in enhancing object visibility and aiding VOS methods under such low-light conditions. This paper introduces a pioneering framework tai-lored for low-light VOS, leveraging event camera data to elevate segmentation accuracy. Our approach hinges on two pivotal components: the Adaptive Cross-Modal Fusion (ACMF) module, aimed at extracting pertinent features while fusing image and event modalities to mitigate noise interference, and the Event-Guided Memory Matching (EGMM) module, designed to rectify the issue of in-accurate matching prevalent in low-light settings. Additionally, we present the creation of a synthetic LLE-DAVIS dataset and the curation of a real-world LLE-vas dataset, encompassing frames and events. Experimental evaluations corroborate the efficacy of our method across both datasets, affirming its effectiveness in low-light scenarios. The datasets are available at https://github.com/HebeiFast/EventLowLightVOS.
Hebei Li, Jin Wang 0023, Jiahui Yuan, Wenming Weng, Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001
CVPR7
2024 Cross-dimension Affinity Distillation for 3D EM Neuron Segmentation
abstract
Accurate 3D neuron segmentation from electron mi-croscopy (EM) volumes is crucial for neuroscience re-search. However, the complex neuron morphology often leads to over-merge and over-segmentation results. Recent advancements utilize 3D CNNs to predict a 3D affinity map with improved accuracy but suffer from two challenges: high computational cost and limited input size, especially for practical deployment for large-scale EM volumes. To address these challenges, we propose a novel method to leverage lightweight 2D CNNs for efficient neuron segmen-tation. Our method employs a 2D Y-shape network to generate two embedding maps from adjacent 2D sections, which are then converted into an affinity map by measuring their embedding distance. While the 2D network better captures pixel dependencies inside sections with larger in-put sizes, it overlooks inter-section dependencies. To over-come this, we introduce a cross-dimension affinity distillation (CAD) strategy that transfers inter-section dependency knowledge from a 3D teacher network to the 2D student network by ensuring consistency between their output affin-ity maps. Additionally, we design a feature grafting in-teraction (FGI) module to enhance knowledge transfer by grafting embedding maps from the 2D student onto those from the 3D teacher. Extensive experiments on multiple EM neuron segmentation datasets, including a newly built one by ourselves, demonstrate that our method achieves supe-rior performance over state-of-the-art methods with only 1/20 inference latency. We release our code and dataset at https://github.com/liuxyll03/CAD.
Xiaoyu Liu 0006, Yinda Chen, Yueyi Zhang 0001, Te Shi 0003, Ruobing Zhang, Xuejin Chen, Zhiwei Xiong
CVPR4
2024 Scene Adaptive Sparse Transformer for Event-based Object Detection
abstract
While recent Transformer-based approaches have shown impressive performances on event-based object detection tasks, their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. However, they display inade-quate sparsity and adaptability when applied to event-based object detection, since these approaches cannot balance the fine granularity of token-level sparsification and the efficiency of window-based Transformers, leading to re-duced performance and efficiency. Furthermore, they lack scene-specific sparsity optimization, resulting in information loss and a lower recall rate. To overcome these limi-tations, we propose the Scene Adaptive Sparse Transformer (SAST). SAST enables window-token co-sparsification, sig-nificantly enhancing fault tolerance and reducing compu-tational overhead. Leveraging the innovative scoring and selection modules, along with the Masked Sparse Window Self-Attention, SAST showcases remarkable scene-aware adaptability: It focuses only on important objects and dy-namically optimizes sparsity level according to scene complexity, maintaining a remarkable balance between performance and computational cost. The evaluation results show that SAST outperforms all other dense and sparse networks in both performance and efficiency on two large-scale event-based object detection datasets (1 Mpx and Genl). Code: https://github.com/Peterande/SAST.
Yansong Peng, Hebei Li, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005
CVPR3
2024 Anatomical Consistency Distillation and Inconsistency Synthesis for Brain Tumor Segmentation with Missing Modalities
abstract
Multi-modal Magnetic Resonance Imaging (MRI) is imperative for accurate brain tumor segmentation, offering indispensable complementary information. Nonetheless, the absence of modalities poses significant challenges in achieving precise segmentation. Recognizing the shared anatomical structures between mono-modal and multi-modal representations, it is noteworthy that mono-modal images typically exhibit limited features in specific regions and tissues. In response to this, we present Anatomical Consistency Distillation and Inconsistency Synthesis (ACDIS), a novel framework designed to transfer anatomical structures from multi-modal to mono-modal representations and synthesize modality-specific features. ACDIS consists of two main components: Anatomical Consistency Distillation (ACD) and Modality Feature Synthesis Block (MFSB). ACD incorporates the Anatomical Feature Enhancement Block (AFEB), meticulously mining anatomical information. Simultaneously, Anatomical Consistency ConsTraints (ACCT) are employed to facilitate the consistent knowledge transfer, i.e., the richness of information and the similarity in anatomical structure, ensuring precise alignment of structural features across mono-modality and multi-modality. Complementarily, MFSB produces modality-specific features to rectify anatomical inconsistencies, thereby compensating for missing information in the segmented features. Through validation on the BraTS2018 and BraTS2020 datasets, ACDIS substantiates its efficacy in the segmentation of brain tumors with missing MRI modalities.
Zheyu Zhang 0002, Xinzhao Liu, Yueyi Zhang 0001, Huanjing Yue, Yunwei Ou, Xiaoyan Sun 0001
ECAI4
2024 Exploiting Dual-Correlation for Multi-frame Time-of-Flight Denoising
Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong
ECCV (22)2
2024 Event-Adapted Video Super-Resolution
Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Zhengjun Zha, Xiaoyan Sun 0001, Zhiwei Xiong
ECCV (42)3
2024 High-Resolution and Few-Shot View Synthesis from Asymmetric Dual-Lens Inputs
Ruikang Xu, Mingde Yao, Yueyi Zhang 0001, Zhiwei Xiong
ECCV (3)4
2024 Event-Based Head Pose Estimation: Benchmark and Method
Jiahui Yuan, Hebei Li, Yansong Peng, Jin Wang 0023, Yuheng Jiang, Yueyi Zhang 0001, Xiaoyan Sun 0001
ECCV (15)6
2024 Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation
abstract
The Segment Anything Model (SAM) has demonstrated outstanding zero-shot performance in image segmentation through efficient point and box prompts. In this paper, we propose a SAM-based Semantic-enhanced Point-Box joint prompting (SAM-SPB) framework for Video Object Segmentation (VOS). SAM-SPB leverages the local structure information and the global semantic cues of interest objects, leading to strong and robust segmentation. To be specific, the local structure information of the objects is maintained by a point tracking branch, and the semantic consistency of the objects across frames are propagated through our proposed semantic-aware memory-based box tracking branch. Compared with previous SAM-based point-centric video segmentation method, we highlight the importance of point-box joint prompting for video object segmentation. The state-of-the-art experimental results on popular VOS benchmarks in the zero-shot setting demonstrate the strong zero-shot ability of the proposed method.
Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICIP3
2024 Joint Flow Estimation from Point Clouds and Event Streams
abstract
Understanding scene dynamics relies heavily on optical flow and scene flow. Most existing flow estimation methods use low-rate RGB images and point clouds, and match the frames geometrically. However, this approach faces challenges in real-world scenes with intricate motion, occlusion, and noise. To tackle this problem, we combine point clouds with events, which introduce dynamic inter-frame information. We propose a bi-stream neural network that jointly estimates optical flow and scene flow. The event branch extracts dynamic information and estimates optical flow, while the point branch captures scene structure and estimate scene flow. A Spatio-temporal Fusion Block is introduced to fuse the complementary information from points and events. Additionally, we adopt a result-level fusion strategy for direct refinement between the flow predictions of the two branches. We evaluate our model on the real-world datasets DSEC and MVSEC. The experimental results demonstrate superior performance compared to existing methods.
Yueyi Zhang 0001, Shida Sun, Zhiwei Xiong
ICME2
2024 ESTME: Event-driven Spatio-temporal Motion Enhancement for Micro-Expression Recognition
abstract
The inherently rapid and subtle changes in micro-expressions pose significant challenges for micro-expression recognition (MER). Previous methods, typically relying on frame aggregation or optical flow, struggle to accurately capture subtle changes because of low frame rate. In this paper, we propose an Event-driven Spatio-temporal Motion Enhancement Network, which incorporates event signals captured by an event camera, to assist MER. Specifically, we introduce an Event-Enhanced Motion Extractor module to exploit event signals’ high temporal resolution property, enhancing subtle motion details. We also propose an Event-Guided Attention module to focus on subtle changes in specific areas, capturing more precise spatial features of micro-expressions. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method on MER, showcasing its strong ability to capture subtle motion changes.
Peilin Xiao, Yueyi Zhang 0001, Dachun Kai, Yansong Peng, Zheyu Zhang 0002, Xiaoyan Sun 0001
ICME2
2024 EvTexture: Event-driven Texture Enhancement for Video Super-Resolution
abstract
Event-based vision has drawn increasing attention due to its unique characteristics, such as high temporal resolution and high dynamic range. It has been used in video super-resolution (VSR) recently to enhance the flow estimation and temporal alignment. Rather than for motion learning, we propose in this paper the first VSR method that utilizes event signals for texture enhancement. Our method, called EvTexture, leverages high-frequency details of events to better recover texture regions in VSR. In our EvTexture, a new texture enhancement branch is presented. We further introduce an iterative texture enhancement module to progressively explore the high-temporal-resolution event information for texture restoration. This allows for gradual refinement of texture regions across multiple iterations, leading to more accurate and rich high-resolution details. Experimental results show that our EvTexture achieves state-of-the-art performance on four datasets. For the Vid4 dataset with rich textures, our method can get up to 4.67dB gain compared with recent event-based methods. Code: https://github.com/DachunKai/EvTexture.
Dachun Kai, Jiayao Lu, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICML3
2024 GRACE: GRadient-based Active Learning with Curriculum Enhancement for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) aims to predict sentiment from text, audio, and visual data of videos. Existing works focus on designing fusion strategies or decoupling mechanisms, which suffer from low data utilization and a heavy reliance on large amounts of labeled data. However, acquiring large-scale annotations for multimodal sentiment analysis is extremely labor-intensive and costly. To address this challenge, we propose GRACE, a GRadient-based Active learning method with Curriculum Enhancement, designed for MSA under a multi-task learning framework. Our approach achieves annotation reduction by strategically selecting valuable samples from the unlabeled data pool while maintaining high-performance levels. Specifically, we introduce informativeness and representativeness criteria, calculated from gradient magnitudes and sample distances, to quantify the active value of unlabeled samples. Additionally, an easiness criterion is incorporated to avoid outliers, considering the relationship between modality consistency and sample difficulty. During the learning process, we dynamically balance sample difficulty and active value, guided by the curriculum learning principle. This strategy prioritizes easier, modality-aligned samples for stable initial training, then gradually increases the difficulty by incorporating more challenging samples with modality conflicts. Extensive experiments demonstrate the effectiveness of our approach on both multimodal sentiment regression and classification benchmarks.
Wenqing Ye, Yueyi Zhang 0001, Xiaoyan Sun 0001
ACM Multimedia3
2024 Asymmetric Event-Guided Video Super-Resolution
abstract
Event cameras are novel bio-inspired cameras that record asynchronous events with high temporal resolution and dynamic range. Leveraging the auxiliary temporal information recorded by event cameras holds great promise for the task of video super-resolution (VSR). However, existing event-guided VSR methods assume that the event and RGB cameras are strictly calibrated (e.g., pixel-level sensor designs in DAVIS 240/346). This assumption proves limiting in emerging high-resolution devices, such as dual-lens smartphones and unmanned aerial vehicles, where such precise calibration is typically unavailable. To unlock more event-guided application scenarios, we perform the task of asymmetric event-guided VSR for the first time, and we propose an Asymmetric Event-guided VSR Network (AsEVSRN) for this new task. AsEVSRN incorporates two specialized designs for leveraging the asymmetric event stream in VSR. Firstly, the content hallucination module dynamically enhances event and RGB information by exploiting their complementary nature, thereby adaptively boosting representational capacity. Secondly, the event-enhanced bidirectional recurrent cells align and propagate temporal features fused with features from content-hallucinated frames. Within the bidirectional recurrent cells, event-enhanced flow is employed to simultaneously utilize and fuse temporal information at both the feature and pixel levels. Comprehensive experimental results affirm that our method consistently generates superior quantitative and qualitative results.
Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong
ACM Multimedia3
2024 Toward Dynamic Non-Line-of-Sight Imaging with Mamba Enforced Temporal Consistency
abstract
Dynamic reconstruction in confocal non-line-of-sight imaging encounters great challenges since the dense raster-scanning manner limits the practical frame rate. A fewer pioneer works reconstruct high-resolution volumes from the under-scanning transient measurements but overlook temporal consistency among transient frames. To fully exploit multi-frame information, we propose the first spatial-temporal Mamba (ST-Mamba) based method tailored for dynamic reconstruction of transient videos. Our method capitalizes on neighbouring transient frames to aggregate the target 3D hidden volume. Specifically, the interleaved features extracted from the input transient frames are fed to the proposed ST-Mamba blocks, which leverage the time-resolving causality in transient measurement. The cross ST-Mamba blocks are then devised to integrate the adjacent transient features. The target high-resolution transient frame is subsequently recovered by the transient spreading module. After transient fusion and recovery, a physical-based network is employed to reconstruct the hidden volume. To tackle the substantial noise inherent in transient videos, we propose a wave-based loss function to impose constraints within the phasor field. Besides, we introduce a new dataset, comprising synthetic videos for training and real-world videos for evaluation. Extensive experiments showcase the superior performance of our method on both synthetic data and real world data captured by different imaging setups. The code and data are available at https://github.com/Depth2World/Dynamic_NLOS.
Shida Sun, Juntian Ye, Yueyi Zhang 0001, Feihu Xu, Zhiwei Xiong
NeurIPS5
2024 Visual Perception by Large Language Model's Weights
abstract
Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstrate promising results on various vision-language tasks but are limited by the high computational effort due to the extended input sequence resulting from the involvement of visual tokens. In this paper, instead of input space alignment, we propose a novel parameter space alignment paradigm that represents visual information as model weights. For each input image, we use a vision encoder to extract visual features, convert features into perceptual weights, and merge the perceptual weights with LLM's weights. In this way, the input of LLM does not require visual tokens, which reduces the length of the input sequence and greatly improves efficiency. Following this paradigm, we propose VLoRA with the perceptual weights generator. The perceptual weights generator is designed to convert visual features to perceptual weights with low-rank property, exhibiting a form similar to LoRA. The experimental results show that our VLoRA achieves comparable performance on various benchmarks for MLLMs, while significantly reducing the computational costs for both training and inference. Code and models are released at \url{https://github.com/FeipengMa6/VLoRA}.
Feipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang, Fengyun Rao, Shilin Yan, Yueyi Zhang 0001, Siying Wu, Zheng Shou 0001, Xiaoyan Sun 0001
NeurIPS7
2024 Depth from Asymmetric Frame-Event Stereo: A Divide-and-Conquer Approach
abstract
Event cameras asynchronously measure brightness changes in a scene without motion blur or saturation, while frame cameras capture images with dense intensity and fine details at a fixed rate. The exclusive advantages of the two modalities make depth estimation from Stereo Asymmetric Frame-Event (SAFE) systems appealing. However, due to the inevitable information absence of one modality in certain challenging regions, existing stereo matching methods lose efficacy for asymmetric inputs from SAFE systems. In this paper, we propose a divide-and-conquer approach that decomposes depth estimation from SAFE systems into three sub-tasks, i.e., frame-event stereo matching, frame-based Structure-from-Motion (SfM), and event-based SfM. In this way, the above challenging regions are addressed by monocular SfM, which estimates robust depth with two views belonging to the same functioning modality. Moreover, we propose a dual sampling strategy to construct cost volumes with identical spatial locations and depth hypotheses for different sub-tasks, which enables sub-task fusion at the cost volume level. To tackle the occlusion issue raised by the sampling strategy, we further introduce a temporal fusion scheme to utilize long-term sequential inputs with multi-view information. Experimental results validate the superior performance of our method over existing solutions.
Xihao Chen, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
WACV3
2024 Deep multi-threshold spiking-UNet for image processing
Hebei Li, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001
Neurocomputing2
2024 Event-Based Stereo Depth Estimation by Temporal-Spatial Context Learning
abstract
Event cameras represent a cutting-edge sensor technology, recording asynchronous pixel-level intensity changes with high temporal resolution and a wide dynamic range. These attributes make event-based stereo depth estimation particularly robust for scenarios characterized by rapid changes and challenging lighting conditions. However, previous learning-based approaches for event-based stereo have often overlooked exploiting the temporal context information within the scene, resulting in suboptimal depth estimations. In this paper, we introduce a novel learning-based network for event-based stereo that incorporates two innovative modules: the Event-based Temporal Aggregation Module (E-TAM) and the Temporal-guided Spatial Context Learning Module (T-SCLM). The E-TAM is designed to capture temporal context information among temporal features extracted from the entire event stream, further the T-SCLM exploits the temporal context information to provide guidance for spatial context learning. Subsequently, these merged features are input into the stereo matching network, ultimately yielding the final disparity map. Experimental evaluations conducted on two real-world datasets affirm the superiority of our method when compared to state-of-the-art approaches.
Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005
IEEE Signal Process. Lett.2
2024 Learned Rate-Distortion Cost Prediction for Ultrafast Screen Content Intra Coding
abstract
As online collaborations become more prevalent, screen content has become increasingly important in real-time video communications. To reduce communication costs, the H.265/HEVC standard introduced the Screen Content Coding (SCC) extension, which achieves significant bits savings but comes with a higher encoding complexity. There is a need for ultrafast SCC encoding to meet the demands of real-time applications. Our key idea is to predict the rate-distortion (RD) cost of each possible coding unit under each possible mode, rather than performing actual coding to obtain the RD cost. Specifically, we construct neural networks to predict RD costs for intra prediction, palette, and normal intra block copy (IBC) modes. For IBC merge mode, we conduct motion compensation trials and use a linear regression network for prediction. Using the predicted RD costs, we create a partition-mode map set that determines not only block partitioning but also optimal modes, significantly reducing encoding complexity. Our experimental results demonstrate that our method achieves a more than 90% reduction in encoding time with an average 9.4% BD-rate increase compared to the HEVC-SCC reference software in the all-intra configuration.
Yanchen Zuo, Changsheng Gao, Dong Liu 0002, Li Li 0040, Yueyi Zhang 0001, Xiaoyan Sun 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 WASPSYN: A Challenge for Domain Adaptive Synapse Detection in Microwasp Brain Connectomes
abstract
The size of image volumes in connectomics studies now reaches terabyte and often petabyte scales with a great diversity of appearance due to different sample preparation procedures. However, manual annotation of neuronal structures (e.g., synapses) in these huge image volumes is time-consuming, leading to limited labeled training data often smaller than 0.001% of the large-scale image volumes in application. Methods that can utilize in-domain labeled data and generalize to out-of-domain unlabeled data are in urgent need. Although many domain adaptation approaches are proposed to address such issues in the natural image domain, few of them have been evaluated on connectomics data due to a lack of domain adaptation benchmarks. Therefore, to enable developments of domain adaptive synapse detection methods for large-scale connectomics applications, we annotated 14 image volumes from a biologically diverse set of Megaphragma viggianii brain regions originating from three different whole-brain datasets and organized the WASPSYN challenge at ISBI 2023. The annotations include coordinates of pre-synapses and post-synapses in the 3D space, together with their one-to-many connectivity information. This paper describes the dataset, the tasks, the proposed baseline, the evaluation method, and the results of the challenge. Limitations of the challenge and the impact on neuroscience research are also discussed. The challenge is and will continue to be available at https://codalab.lisn.upsaclay.fr/competitions/9169. Successful algorithms that emerge from our challenge may potentially revolutionize real-world connectomics research and further the cause that aims to unravel the complexity of brain structure and function.
Yicong Li 0002, Wanhua Li 0001, Qi Chen 0014, Wei Huang 0036, Yuda Zou, Kazunori Shinomiya, Pat Gunn, Nishika Gupta, Alexey Polilov, Yongchao Xu, Yueyi Zhang 0001, Zhiwei Xiong, Hanspeter Pfister, Donglai Wei 0001, Jingpeng Wu
IEEE Trans. Medical Imaging12
2023 Better and Faster: Adaptive Event Conversion for Event-Based Object Detection
abstract
Event cameras are a kind of bio-inspired imaging sensor, which asynchronously collect sparse event streams with many advantages. In this paper, we focus on building better and faster event-based object detectors. To this end, we first propose a computationally efficient event representation Hyper Histogram, which adequately preserves both the polarity and temporal information of events. Then we devise an Adaptive Event Conversion module, which converts events into Hyper Histograms according to event density via an adaptive queue. Moreover, we introduce a novel event-based augmentation method Shadow Mosaic, which significantly improves the event sample diversity and enhances the generalization ability of detection models. We equip our proposed modules on three representative object detection models: YOLOv5, Deformable-DETR, and RetinaNet. Experimental results on three event-based detection datasets (1Mpx, Gen1, and MVSEC-NIGHTL21) demonstrate that our proposed approach outperforms other state-of-the-art methods by a large margin, while achieving a much faster running speed (< 14 ms and < 4 ms for 50 ms event data on the 1Mpx and Gen1 datasets).
Yansong Peng, Yueyi Zhang 0001, Peilin Xiao, Xiaoyan Sun 0001, Feng Wu 0001
AAAI2
2023 Depth Estimation from Indoor Panoramas with Neural Scene Representation
abstract
Depth estimation from indoor panoramas is challenging due to the equirectangular distortions of panoramas and inaccurate matching. In this paper, we propose a practical framework to improve the accuracy and efficiency of depth estimation from multi-view indoor panoramic images with the Neural Radiance Field technology. Specifically, we develop two networks to implicitly learn the Signed Distance Function for depth measurements and the radiance field from panoramas. We also introduce a novel spherical position embedding scheme to achieve high accuracy. For better convergence, we propose an initialization method for the network weights based on the Manhattan World Assumption. Furthermore, we devise a geometric consistency loss, leveraging the surface normal, to further refine the depth estimation. The experimental results demonstrate that our proposed method outperforms state-of-the-art works by a large margin in both quantitative and qualitative evaluations. Our source code is available at https://github.com/WJ-Chang-42/IndoorPanoDepth.
Wenjie Chang, Yueyi Zhang 0001, Zhiwei Xiong
CVPR2
2023 Progressive Spatio-temporal Alignment for Efficient Event-based Motion Estimation
abstract
In this paper, we propose an efficient event-based motion estimation framework for various motion models. Different from previous works, we design a progressive event-to-map alignment scheme and utilize the spatio-temporal correlations to align events. In detail, we progressively align sampled events in an event batch to the time-surface map and obtain the updated motion model by minimizing a novel time-surface loss. In addition, a dynamic batch size strategy is applied to adaptively adjust the batch size so that all events in the batch are consistent with the current motion model. Our framework has three advantages: a) the progressive scheme refines motion parameters iteratively, achieving accurate motion estimation; b) within one iteration, only a small portion of events are involved in optimization, which greatly reduces the total runtime; c) the dynamic batch size strategy ensures that the constant velocity assumption always holds. We conduct comprehensive experiments to evaluate our framework on challenging high-speed scenes with three motion models: rotational, homography, and 6-DOF models. Experimental results demonstrate that our framework achieves state-of-the-art estimation accuracy and efficiency. The code is available at https://github.com/huangxueyan/PEME.
Xueyan Huang, Yueyi Zhang 0001, Zhiwei Xiong
CVPR2
2023 NLOST: Non-Line-of-Sight Imaging with Transformer
abstract
Time-resolved non-line-of-sight (NLOS) imaging is based on the multi-bounce indirect reflections from the hidden objects for 3D sensing. Reconstruction from NLOS measurements remains challenging especially for complicated scenes. To boost the performance, we present NLOST, the first transformer-based neural network for NLOS reconstruction. Specifically, after extracting the shallow features with the assistance of physics-based priors, we design two spatial-temporal self attention encoders to explore both local and global correlations within 3D NLOS data by splitting or downsampling the features into different scales, respectively. Then, we design a spatial-temporal cross attention decoder to integrate local and global features in the token space of transformer, resulting in deep features with high representation capabilities. Finally, deep and shallow features are fused to reconstruct the 3D volume of hidden scenes. Extensive experimental results demonstrate the superior performance of the proposed method over existing solutions on both synthetic data and real-world data captured by different NLOS imaging systems.
Jiayong Peng, Juntian Ye, Yueyi Zhang 0001, Feihu Xu, Zhiwei Xiong
CVPR4
2023 A Soma Segmentation Benchmark in Full Adult Fly Brain
abstract
Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural functions. However, due to the absence of EM datasets specifically annotated for somas, existing deep learning-based neuron reconstruction methods cannot directly provide accurate soma distribution and morphology. Moreover, full brain neuron reconstruction remains extremely time-consuming due to the unprecedentedly large size of EM data. In this paper, we develop an efficient soma reconstruction method for obtaining accurate soma distribution and morphology information in a full adult fly brain. To this end, we first make a high-resolution EM dataset with fine-grained 3D manual annotations on somas. Relying on this dataset, we propose an efficient, two-stage deep learning algorithm for predicting accurate locations and boundaries of 3D soma instances. Further, we deploy a parallelized, high-throughput data processing pipeline for executing the above algorithm on the full brain. Finally, we provide quantitative and qualitative benchmark comparisons on the testset to validate the superiority of the proposed method, as well as preliminary statistics of the reconstructed somas in the full adult fly brain from the biological perspective. We release our code and dataset at https://github.com/liuxy1103/EMADS.
Xiaoyu Liu 0006, Bo Hu 0014, Mingxing Li 0003, Wei Huang 0036, Yueyi Zhang 0001, Zhiwei Xiong
CVPR5
2023 Event-based Blurry Frame Interpolation under Blind Exposure
abstract
Restoring sharp high frame-rate videos from low frame-rate blurry videos is a challenging problem. Existing blurry frame interpolation methods assume a predefined and known exposure time, which suffer from severe performance drop when applied to videos captured in the wild. In this paper, we study the problem of blurry frame interpolation under blind exposure with the assistance of an event camera. The high temporal resolution of the event camera is beneficial to obtain the exposure prior that is lost during the imaging process. Besides, sharp frames can be restored using event streams and blurry frames relying on the mutual constraint among them. Therefore, we first propose an exposure estimation strategy guided by event streams to estimate the lost exposure prior, transforming the blind exposure problem well-posed. Second, we propose to model the mutual constraint with a temporal-exposure control strategy through iterative residual learning. Our blurry frame interpolation method achieves a distinct performance boost over existing methods on both synthetic and self-collected real- world datasets under blind exposure.
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
CVPR2
2023 Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation
abstract
Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision. In this paper, we propose a sparsely supervised biomedical instance segmentation framework via cross-representation affinity consistency regularization. Specifically, we adopt two individual networks to enforce the perturbation consistency between an explicit affinity map and an implicit affinity map to capture both feature-level instance discrimination and pixel-level instance boundary structure. We then select the highly confident region of each affinity map as the pseudo label to supervise the other one for affinity consistency learning. To obtain the highly confident region, we propose a pseudo-label noise filtering scheme by integrating two entropy-based decision strategies. Extensive experiments on four biomedical datasets with sparse instance annotations show the state-of-the-art performance of our proposed framework. For the first time, we demonstrate the superiority of sparse instance-level supervision on 3D volumetric datasets, compared to common semi-supervision under the same annotation cost. Code is available at https://github.com/liuxy1103/CRAC.
Xiaoyu Liu 0006, Wei Huang 0036, Zhiwei Xiong, Shenglong Zhou 0002, Yueyi Zhang 0001, Xuejin Chen, Zhengjun Zha, Feng Wu 0001
ICCV5
2023 GET: Group Event Transformer for Event-Based Vision
abstract
Event cameras are a type of novel neuromorphic sensor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity. To address this issue, we propose a novel Group-based vision Transformer backbone for Event-based vision, called Group Event Transformer (GET), which decouples temporal-polarity information from spatial information throughout the feature extraction process. Specifically, we first propose a new event representation for GET, named Group Token, which groups asynchronous events based on their timestamps and polarities. Then, GET applies the Event Dual Self-Attention block, and Group Token Aggregation module to facilitate effective feature communication and integration in both the spatial and temporal-polarity domains. After that, GET can be integrated with different downstream tasks by connecting it with various heads. We evaluate our method on four event-based classification datasets (Cifar10-DVS, N-MNIST, N-CARS, and DVS128Gesture) and two event-based object detection datasets (1Mpx and Gen1), and the results demonstrate that GET outperforms other state-of-the-art methods. The code is available at https://github.com/Peterande/GET-Group-Event-Transformer.
Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001
ICCV2
2023 Unsupervised Video Deraining with An Event Camera
abstract
Current unsupervised video deraining methods are inefficient in modeling the intricate spatio-temporal properties of rain, which leads to unsatisfactory results. In this paper, we propose a novel approach by integrating a bio-inspired event camera into the unsupervised video deraining pipeline, which enables us to capture high temporal resolution information and model complex rain characteristics. Specifically, we first design an end-to-end learning-based network consisting of two modules, the asymmetric separation module and the cross-modal fusion module. The two modules are responsible for segregating the features of the rain-background layer, and for positive enhancement and negative suppression from a cross-modal perspective, respectively. Second, to regularize the network training, we elaborately design a cross-modal contrastive learning method that leverages the complementary information from event cameras, exploring the mutual exclusion and similarity of rain-background layers in different domains. This encourages the deraining network to focus on the distinctive characteristics of each layer and learn a more discriminative representation. Moreover, we construct the first real-world dataset comprising rainy videos and events using a hybrid imaging system. Extensive experiments demonstrate the superior performance of our method on both synthetic and real-world datasets.
Jin Wang 0023, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
ICCV3
2023 Video Super-Resolution Via Event-Driven Temporal Alignment
abstract
Video super-resolution aims to restore low-resolution videos into their high-resolution counterparts. Existing methods typically rely on optical flow, which assumes linear motion and is sensitive to rapid lighting changes, to capture inter-frame information. Event cameras can asynchronously output high temporal resolution event streams, which can reflect nonlinear motion and are robust to lighting changes. Inspired by these characteristics, we propose an Event-driven Bidirectional Video Super-Resolution (EBVSR) framework. Firstly, we propose an event-assisted temporal alignment module that utilizes events to generate nonlinear motion to align adjacent frames, complementing flow-based methods. Secondly, we build an event-based frame synthesis module that enhances the network’s robustness to lighting changes through a bidirectional cross-modal fusion design. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method. The code is available at https://github.com/DachunKai/EBVSR.
Dachun Kai, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICIP2
2023 Attention-Guided Contrastive Masked Image Modeling for Transformer-Based Self-Supervised Learning
abstract
Self-supervised learning with vision transformer (ViT) has gained much attention recently. Most existing methods rely on either contrastive learning or masked image modeling. The former is suitable for global feature extraction but underperforms in fine-grained tasks. The later explores the internal structure of images but ignores the high information sparsity and unbalanced information distribution. In this paper, we propose a new approach called Attention-guided Contrastive Masked Image Modeling (ACoMIM), which integrates the merits of both paradigms and leverages the attention mechanism of ViT for effective representation. Specifically, it has two pretext tasks, predicting the features of masked regions guided by attention and comparing the global features of masked and unmasked images. We show that these two pretext tasks complement each other and improve our method’s performance. The experiments demonstrate that our model transfers well to various downstream tasks such as classification and object detection. Code is available at https://github.com/yczhan/ACoMIM.
Yucheng Zhan, Chong Luo 0001, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICIP4
2023 Multimodal Sentiment Analysis with Preferential Fusion and Distance-aware Contrastive Learning
abstract
Recent efforts on multimodal sentiment analysis (MSA) leverage data from multiple modalities, among which the text modality is heavily relied on. However, the text modality often contains false correlations between text tokens and sentiment labels, leading to errors in sentiment analysis. To address this issue, we propose a new framework, PriSA, which incorporates the preferential fusion and distance-aware contrastive learning. Specifically, we first propose a preferential inter-modal fusion method, which utilizes the text modality to guide the calculation of the inter-modal correlations. Then the resulting inter-modal features are further used to calculate mixed-modal correlations through our proposed distance-aware contrastive learning, which leverages the distance information of the sentiment labels. At last, we identify the sentiment information based on both the mixed-modal correlations and the discriminative intra-modal features extracted from the visual and audio modalities via a self-attention module. Experimental results show that our proposed PriSA achieves the state-of-the-art performance on four datasets, including MOSEI, MOSI, SIMS, and UR-FUNNY. The code is available at https://github.com/FeipengMa6/PriSA.
Feipeng Ma, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICME2
2023 EoFormer: Edge-Oriented Transformer for Brain Tumor Segmentation
Dong She, Yueyi Zhang 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001
MICCAI (4)2
2023 Deep Non-line-of-sight Imaging from Under-scanning Measurements
abstract
Active confocal non-line-of-sight (NLOS) imaging has successfully enabled seeing around corners relying on high-quality transient measurements. However, acquiring spatial-dense transient measurement is time-consuming, raising the question of how to reconstruct satisfactory results from under-scanning measurements (USM). The existing solutions, involving the traditional algorithms, however, are hindered by unsatisfactory results or long computing times. To this end, we propose the first deep-learning-based approach to NLOS imaging from USM. Our proposed end-to-end network is composed of two main components: the transient recovery network (TRN) and the volume reconstruction network (VRN). Specifically, TRN takes the under-scanning measurements as input, utilizes a multiple kernel feature extraction module and a multiple feature fusion module, and outputs sufficient-scanning measurements at the high-spatial resolution. Afterwards, VRN incorporates the linear physics prior of the light-path transport model and reconstructs the hidden volume representation. Besides, we introduce regularized constraints that enhance the perception of more local details while suppressing smoothing effects. The proposed method achieves superior performance on both synthetic data and public real-world data, as demonstrated by extensive experimental results with different under-scanning grids. Moreover, the proposed method delivers impressive robustness at an extremely low scanning grid (i.e., 8$\times$8) and offers high-speed inference (i.e., 50 times faster than the existing iterative solution).
Yueyi Zhang 0001, Juntian Ye, Feihu Xu, Zhiwei Xiong
NeurIPS2
2023 QISampling: An Effective Sampling Strategy for Event-Based Sign Language Recognition
abstract
Event cameras are innovative neuromorphic sensors that detect changes in brightness and generate a stream of events, thus providing high temporal resolution and dynamic range advantages. With its ability to perceive motion information, event data is well-suited for sign language recognition (SLR). Existing methods of event-based SLR rely on a uniform sampling strategy, which may result in redundant and indiscriminate information being captured when segments are randomly sampled. In this letter, we propose an effective sampling strategy called Quantity-inspired Sampling (QISampling) that takes advantage of the quantity feature of event distribution to sample key segments containing more discriminative and remarkable motion features from a sign. These segments are then converted into frame-like representations, which serve as input to subsequent Deep Neural Networks (DNNs). In addition, we introduce a synthetic event-based sign language dataset N-WLASL to the community. We apply our strategy to DNNs and conduct experiments on this synthetic dataset and another real-world dataset. The results demonstrate that our QISampling strategy could accelerate training and make event data provide superior performance for the SLR task.
Zhongfu Ye, Jin Wang 0023, Yueyi Zhang 0001
IEEE Signal Process. Lett.4
2022 Degradation-agnostic Correspondence from Resolution-asymmetric Stereo
abstract
In this paper, we study the problem of stereo matching from a pair of images with different resolutions, e.g., those acquired with a tele-wide camera system. Due to the difficulty of obtaining ground-truth disparity labels in diverse real-world systems, we start from an unsupervised learning perspective. However, resolution asymmetry caused by unknown degradations between two views hinders the effectiveness of the generally assumed photometric consistency. To overcome this challenge, we propose to impose the consistency between two views in a feature space instead of the image space, named feature-metric consistency. Interestingly, we find that, although a stereo matching network trained with the photometric loss is not optimal, its feature extractor can produce degradation-agnostic and matching-specific features. These features can then be utilized to formulate a feature-metric loss to avoid the photometric inconsistency. Moreover, we introduce a self-boosting strategy to optimize the feature extractor progressively, which further strengthens the feature-metric consistency. Experiments on both simulated datasets with various degradations and a self-collected real-world dataset validate the superior performance of the proposed method over existing solutions.
Xihao Chen, Zhiwei Xiong, Zhen Cheng 0002, Jiayong Peng, Yueyi Zhang 0001, Zhengjun Zha
CVPR5
2022 Exploiting Rigidity Constraints for LiDAR Scene Flow Estimation
abstract
Previous LiDAR scene flow estimation methods, especially recurrent neural networks, usually suffer from structure distortion in challenging cases, such as sparse reflection and motion occlusions. In this paper, we propose a novel optimization method based on a recurrent neural network to predict LiDAR scene flow in a weakly supervised manner. Specifically, our neural recurrent network exploits direct rigidity constraints to preserve the geometric structure of the warped source scene during an iterative alignment procedure. An error awarded optimization strategy is proposed to update the LiDAR scene flow by minimizing the point measurement error instead of reconstructing the cost volume multiple times. Trained on two autonomous driving datasets, our network outperforms recent state-of-the-art networks on lidarKITTI by a large margin. The code and models will be available at https://github.com/gtdong-ustc/LiDARSceneFlow.
Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong
CVPR2
2022 Boosting Event Stream Super-Resolution with a Recurrent Neural Network
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
ECCV (6)2
2022 Biological Instance Segmentation with a Superpixel-Guided Graph
abstract
Recent advanced proposal-free instance segmentation methods have made significant progress in biological images. However, existing methods are vulnerable to local imaging artifacts and similar object appearances, resulting in over-merge and over-segmentation. To reduce these two kinds of errors, we propose a new biological instance segmentation framework based on a superpixel-guided graph, which consists of two stages, i.e., superpixel-guided graph construction and superpixel agglomeration. Specifically, the first stage generates enough superpixels as graph nodes to avoid over-merge, and extracts node and edge features to construct an initialized graph. The second stage agglomerates superpixels into instances based on the relationship of graph nodes predicted by a graph neural network (GNN). To solve over-segmentation and prevent introducing additional over-merge, we specially design two loss functions to supervise the GNN, i.e., a repulsion-attraction (RA) loss to better distinguish the relationship of nodes in the feature space, and a maximin agglomeration score (MAS) loss to pay more attention to crucial edge classification. Extensive experiments on three representative biological datasets demonstrate the superiority of our method over existing state-of-the-art methods. Code is available at https://github.com/liuxy1103/BISSG.
Xiaoyu Liu 0006, Wei Huang 0036, Yueyi Zhang 0001, Zhiwei Xiong
IJCAI3
2022 Domain Adaptive Mitochondria Segmentation via Enforcing Inter-Section Consistency
Wei Huang 0036, Xiaoyu Liu 0006, Zhen Cheng 0002, Yueyi Zhang 0001, Zhiwei Xiong
MICCAI (4)4
2022 Efficient Biomedical Instance Segmentation via Knowledge Distillation
Xiaoyu Liu 0006, Bo Hu 0014, Wei Huang 0036, Yueyi Zhang 0001, Zhiwei Xiong
MICCAI (4)4
2022 RPPformer-Flow: Relative Position Guided Point Transformer for Scene Flow Estimation
abstract
Estimating scene flow for point clouds is one of the key problems in 3D scene understanding and autonomous driving. Recently the point transformer architecture has become a popular and successful solution for 3D computer vision tasks, e.g., point cloud object detection and completion, but its application to scene flow estimation is rarely explored. In this work, we provide a full transformer based solution for scene flow estimation. We first introduce a novel relative position guided point attention mechanism. Then to relax the memory consumption in practice, we provide an efficient implementation of our proposed point attention layer via matrix factorization and nearest neighbor sampling. Finally, we build a pyramid transformer, named RPPformer-Flow, to estimate the scene flow between two consecutive point clouds in a coarse-to-fine manner. We evaluate our RPPformer-Flow on the FlyingThings3D and KITTI Scene Flow 2015 benchmarks. Experimental results show that our method outperforms previous state-of-the-art methods with large margins.
Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong
ACM Multimedia3
2022 Semi-Supervised Neuron Segmentation via Reinforced Consistency Learning
abstract
Emerging deep learning-based methods have enabled great progress in automatic neuron segmentation from Electron Microscopy (EM) volumes. However, the success of existing methods is heavily reliant upon a large number of annotations that are often expensive and time-consuming to collect due to dense distributions and complex structures of neurons. If the required quantity of manual annotations for learning cannot be reached, these methods turn out to be fragile. To address this issue, in this article, we propose a two-stage, semi-supervised learning method for neuron segmentation to fully extract useful information from unlabeled data. First, we devise a proxy task to enable network pre-training by reconstructing original volumes from their perturbed counterparts. This pre-training strategy implicitly extracts meaningful information on neuron structures from unlabeled data to facilitate the next stage of learning. Second, we regularize the supervised learning process with the pixel-level prediction consistencies between unlabeled samples and their perturbed counterparts. This improves the generalizability of the learned model to adapt diverse data distributions in EM volumes, especially when the number of labels is limited. Extensive experiments on representative EM datasets demonstrate the superior performance of our reinforced consistency learning compared to supervised learning, i.e., up to 400% gain on the VOI metric with only a few available labels. This is on par with a model trained on ten times the amount of labeled data in a supervised manner. Code is available at https://github.com/weih527/SSNS-Net.
Wei Huang 0036, Chang Chen 0004, Zhiwei Xiong, Yueyi Zhang 0001, Xuejin Chen, Xiaoyan Sun 0001, Feng Wu 0001
IEEE Trans. Medical Imaging4
2021 Training Spiking Neural Networks with Accumulated Spiking Flow
abstract
The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demonstrate the potential of SNNs in supervised learning by introducing approximated backpropagation (BP) methods. To deal with the non-differentiable spike function in SNNs, these BP methods utilize information from the spatio-temporal domain to adjust the model parameters. With the increasing of time window and network size, the computational complexity of spatio-temporal backpropagation augments dramatically. In this paper, we propose a new backpropagation method for SNNs based on the accumulated spiking flow (ASF), i.e. ASF-BP. In the proposed ASF-BP method, updating parameters does not rely on the spike train of spiking neurons but leverage accumulated inputs and outputs of spiking neurons over the time window, which reduces the BP complexity significantly. We further present an adaptive linear estimation model to approach the dynamic characteristics of spiking neurons statistically. Experimental results demonstrate that with our proposed ASF-BP method, light-weight convolutional SNNs achieve superior performances compared with other spike-based BP methods on both non-neuromorphic (MNIST, CIFAR10) and neuromorphic (CIFAR10-DVS) datasets. The code is available at https://github.com/neural-lab/ASF-BP.
Hao Wu 0042, Yueyi Zhang 0001, Wenming Weng, Yongting Zhang, Zhiwei Xiong, Zhengjun Zha, Xiaoyan Sun 0001, Feng Wu 0001
AAAI2
2021 Transformer-based Monocular Depth Estimation with Attention Supervision
Wenjie Chang, Yueyi Zhang 0001, Zhiwei Xiong
BMVC2
2021 Event-based Video Reconstruction Using Transformer
abstract
Event cameras, which output events by detecting spatio- temporal brightness changes, bring a novel paradigm to image sensors with high dynamic range and low latency. Previous works have achieved impressive performances on event-based video reconstruction by introducing convolutional neural networks (CNNs). However, intrinsic locality of convolutional operations is not capable of modeling long-range dependency, which is crucial to many vision tasks. In this paper, we present a hybrid CNN- Transformer network for event-based video reconstruction (ET-Net), which merits the fine local information from CNN and global contexts from Transformer In addition, we further propose a Token Pyramid Aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token-space. Experimental results demonstrate that our proposed method achieves superior performance over state-of-the-art methods on multiple real-world event datasets. The code is available at https://github.com/WarranWeng/ET-Net.
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong
ICCV2
2021 Asymmetric Stereo Color Transfer
abstract
Dual-camera systems containing a color camera and a monochrome camera are widely equipped on smartphones. The color camera captures chrominance information while the monochrome camera captures fine details, which causes asymmetry across spectral and spatial dimensions. In these imaging systems, the chrominance information of low-resolution (LR) color images and the spatial information of high-resolution (HR) monochrome images are highly complementary. In this paper, we propose an elaborate convolutional neural network to recover HR color images by transferring color information from LR color images to HR monochrome images. The network contains a novel feature extraction module named U-ASPP and an asymmetric parallax attention module (APAM). Our network achieves state-of-the-art performance on the Flickr1024 stereo dataset with high efficiency. Moreover, the effectiveness of our trained network is validated in real-world asymmetric image pairs captured by a smartphone, which demonstrates that our method has high generalization capability in real-world imaging systems.
Jiayong Peng, Yueyi Zhang 0001, Shan Liu 0001, Xiaoyan Sun 0001, Zhiwei Xiong
ICME3
2021 Learning Neuron Stitching for Connectomics
Xiaoyu Liu 0006, Yueyi Zhang 0001, Zhiwei Xiong, Chang Chen 0004, Wei Huang 0036, Xuejin Chen, Feng Wu 0001
MICCAI (8)2
2021 Stereo Video Super-Resolution via Exploiting View-Temporal Correlations
abstract
Stereo Video Super-Resolution (StereoVSR) aims to generate high-resolution video steams from two low-resolution videos under stereo settings. Existing video super-resolution and stereo image super-resolution techniques can be extended to tackle the StereoVSR task, yet they cannot make full use of the multi-view and temporal information to achieve satisfactory performance. In this paper, we propose a novel Stereo Video Super-Resolution Network (SVSRNet) to fulfill the StereoVSR task via exploiting view-temporal correlations. First, we devise a view-temporal attention module (VTAM) to integrate the information of cross-time-cross-view for constructing high-resolution stereo videos. Second, we propose a spatial-temporal fusion module (STFM), which aggregates the information across time in intra-view to emphasize important features for subsequent restoration. In addition, we design a view-temporal consistency loss function to enforce consistency constraint of superresolved stereo videos. Comprehensive experimental results demonstrate that our method generates superior results.
Ruikang Xu, Zeyu Xiao 0002, Mingde Yao, Yueyi Zhang 0001, Zhiwei Xiong
ACM Multimedia4
2021 Revisiting Flipping Strategy for Learning-based Stereo Depth Estimation
abstract
Deep neural networks (DNNs) have been widely used for stereo depth estimation, which achieve great success in performance. In this paper, we introduce a novel flipping strategy for DNN on the stereo depth estimation task. Specifically, based on a common DNN for stereo matching, we apply the flipping operation for both input stereo images, which are further fed to the original DNN. A flipping loss function is proposed to jointly train the network with the initial loss. We apply our strategy to many representative networks in both supervised and self-supervised manners. Extensive experimental results demonstrate that our proposed strategy improves the performance of these networks.
Yueyi Zhang 0001, Zhiwei Xiong
VCIP2
2020 Spatial Hierarchy Aware Residual Pyramid Network for Time-of-Flight Depth Denoising
Yueyi Zhang 0001, Zhiwei Xiong
ECCV (24)2
2020 Towards Neuron Segmentation from Macaque Brain Images: A Weakly Supervised Approach
Meng Dong, Dong Liu 0002, Zhiwei Xiong, Xuejin Chen, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (5)5
2020 Volumetric End-to-End Optimized Compression for Brain Images
abstract
The amount of volumetric brain image increases rapidly, which requires a vast amount of resources for storage and transmission, so it's urgent to explore an efficient volumetric compression method. Recent years have witnessed the progress of deep learning-based approaches for two-dimensional (2D) natural image compression, but the field of learned volumetric image compression still remains unexplored. In this paper, we propose the first end-to-end learning framework for volumetric image compression by extending the advanced techniques of 2D image compression to volumetric images. Specifically, a convolutional autoencoder is used to compress 3D image cubes, and the non-local attention models are embedded in the convolutional autoencoder to jointly capture local and global correlations. Both hyperprior and autoregressive models are used to perform the conditional probability estimation in entropy coding. To reduce model complexity, we introduce a convolutional long short-term memory network for the autoregressive model based on channel-wise prediction. Experimental results on volumetric mouse brain images show that the proposed method outperforms JPEG2000-3D, HEVC and state-of-the-art 2D methods.
Yueyi Zhang 0001, Dong Liu 0002, Zhiwei Xiong
VCIP2
2020 Neuronal Population Reconstruction From Ultra-Scale Optical Microscopy Images via Progressive Learning
abstract
Reconstruction of neuronal populations from ultra-scale optical microscopy (OM) images is essential to investigate neuronal circuits and brain mechanisms. The noises, low contrast, huge memory requirement, and high computational cost pose significant challenges in the neuronal population reconstruction. Recently, many studies have been conducted to extract neuron signals using deep neural networks (DNNs). However, training such DNNs usually relies on a huge amount of voxel-wise annotations in OM images, which are expensive in terms of both finance and labor. In this paper, we propose a novel framework for dense neuronal population reconstruction from ultra-scale images. To solve the problem of high cost in obtaining manual annotations for training DNNs, we propose a progressive learning scheme for neuronal population reconstruction (PLNPR) which does not require any manual annotations. Our PLNPR scheme consists of a traditional neuron tracing module and a deep segmentation network that mutually complement and progressively promote each other. To reconstruct dense neuronal populations from a terabyte-sized ultra-scale image, we introduce an automatic framework which adaptively traces neurons block by block and fuses fragmented neurites in overlapped regions continuously and smoothly. We build a dataset "VISoR-40" which consists of 40 large-scale OM image blocks from cortical regions of a mouse. Extensive experimental results on our VISoR-40 dataset and the public BigNeuron dataset demonstrate the effectiveness and superiority of our method on neuronal population reconstruction and single neuron reconstruction. Furthermore, we successfully apply our method to reconstruct dense neuronal populations from an ultra-scale mouse brain slice. The proposed adaptive block propagation and fusion strategies greatly improve the completeness of neurites in dense neuronal population reconstruction.
Jie Zhao 0020, Xuejin Chen, Zhiwei Xiong, Dong Liu 0002, Chaoyu Xie, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
IEEE Trans. Medical Imaging7
2019 Instance Segmentation from Volumetric Biomedical Images Without Voxel-Wise Labeling
Meng Dong, Dong Liu 0002, Zhiwei Xiong, Xuejin Chen, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (2)5
2019 Progressive Learning for Neuronal Population Reconstruction from Optical Microscopy Images
Jie Zhao 0020, Xuejin Chen, Zhiwei Xiong, Dong Liu 0002, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (1)6
2019 Fast and Accurate Electron Microscopy Image Registration with 3D Convolution
Shenglong Zhou 0002, Zhiwei Xiong, Chang Chen 0004, Xuejin Chen, Dong Liu 0002, Yueyi Zhang 0001, Zhengjun Zha, Feng Wu 0001
MICCAI (1)6
2017 LF-fusion: Dense and accurate 3D reconstruction from light field images
abstract
Light field (LF) cameras offer the capability of depth estimation in a single shot, which facilitates real-time 3D reconstruction of dynamic scenes. However, the accuracy of depth estimated from LF is still limited. Different from previous methods that generally focus on improving the fidelity of the central view depth, we argue that depth maps obtained at different views contain complementary information. Inspired by the principle of Kinect-fusion, we then propose a novel method for dense and accurate 3D reconstruction from LF images, namely, LF-fusion. Specifically, we use the iterative closest point (ICP) algorithm to register the point clouds generated from different views, and then employ a volumetric integration algorithm based on the truncated signed distance function (TSDF) to reconstruct the final 3D surface. Experiments demonstrate that the proposed method produces superior 3D reconstruction results on two representative LF datasets.
Jiayong Peng, Zhiwei Xiong, Yueyi Zhang 0001, Dong Liu 0002, Feng Wu 0001
VCIP3
2015 Fusion of Time-of-Flight and Phase Shifting for high-resolution and low-latency depth sensing
abstract
Depth sensors based on Time-of-Flight (ToF) and Phase Shifting (PS) have complementary strengths and weaknesses. ToF can provide real-time depth but limited in resolution and sensitive to noise. PS can generate accurate and robust depth with high resolution but requires a number of patterns that leads to high latency. In this paper, we propose a novel fusion framework to take advantages of both ToF and PS. The basic idea is using the coarse depth from ToF to disambiguate the wrapped depth from PS. Specifically, we address two key technical problems: cross-modal calibration and interference-free synchronization between ToF and PS sensors. Experiments demonstrate that the proposed method generates accurate and robust depth with high resolution and low latency, which is beneficial to tremendous applications.
Yueyi Zhang 0001, Zhiwei Xiong, Feng Wu 0001
ICME1
2014 Robust depth sensing with adaptive structured light illumination
Yueyi Zhang 0001, Zhiwei Xiong, Pengyu Cong, Feng Wu 0001
J. Vis. Commun. Image Represent.1
2014 Real-Time Scalable Depth Sensing With Hybrid Structured Light Illumination
abstract
Time multiplexing (TM) and spatial neighborhood (SN) are two mainstream structured light techniques widely used for depth sensing. The former is well known for its high accuracy and the latter for its low delay. In this paper, we explore a new paradigm of scalable depth sensing to integrate the advantages of both the TM and SN methods. Our contribution is twofold. First, we design a set of hybrid structured light patterns composed of phase-shifted fringe and pseudo-random speckle. Under the illumination of the hybrid patterns, depth can be decently reconstructed either from a few consecutive frames with the TM principle for static scenes or from a single frame with the SN principle for dynamic scenes. Second, we propose a scene-adaptive depth sensing framework based on which a global or region-wise optimal depth map can be generated through motion detection. To validate the proposed scalable paradigm, we develop a real-time (20 fps) depth sensing system. Experimental results demonstrate that our method achieves an efficient balance between accuracy and speed during depth sensing that has rarely been exploited before.
Yueyi Zhang 0001, Zhiwei Xiong, Feng Wu 0001
IEEE Trans. Image Process.1
2013 Depth Acquisition from Density Modulated Binary Patterns
abstract
This paper proposes novel density modulated binary patterns for depth acquisition. Similar to Kinect, the illumination patterns do not need a projector for generation and can be emitted by infrared lasers and diffraction gratings. Our key idea is to use the density of light spots in the patterns to carry phase information. Two technical problems are addressed here. First, we propose an algorithm to design the patterns to carry more phase information without compromising the depth reconstruction from a single captured image as with Kinect. Second, since the carried phase is not strictly sinusoidal, the depth reconstructed from the phase contains a systematic error. We further propose a pixel-based phase matching algorithm to reduce the error. Experimental results show that the depth quality can be greatly improved using the phase carried by the density of light spots. Furthermore, our scheme can achieve 20 fps depth reconstruction with GPU assistance.
Zhiwei Xiong, Yueyi Zhang 0001, Feng Wu 0001
CVPR3
2013 Dense single-shot 3D scanning via stereoscopic fringe analysis
abstract
In this paper, we present a novel single-shot method for dense and accurate 3D scanning. Our method takes advantage of two conventional techniques, i.e., stereo and Fourier fringe analysis (FFA). While FFA is competent for high-density and high-precision phase measurement, stereo solves the phase ambiguity caused by the periodicity of the fringe. By jointly using the intensity images and unwrapped phase maps from stereo, the pixel-wise absolute depth can be obtained through a sparse matching process efficiently and reliably. Due to its single-shot property and low complexity, the proposed method facilitates dense and accurate 3D scanning in time-critical applications.
Pengyu Cong, Zhiwei Xiong, Yueyi Zhang 0001, Feng Wu 0001
ICIP3
2013 Accurate 3D reconstruction of dynamic scenes with Fourier transform assisted phase shifting
abstract
Phase shifting is a widely used method for accurate and dense 3D reconstruction. However, at least three images of the same scene are required for each reconstruction, so measurement errors are inevitable in dynamic scenes, even with high-speed hardware. In this paper, we propose a Fourier transform assisted phase shifting method to overcome the motion vulnerability in phase shifting. A new model with motion-related phase shifts is formulated, and the coarse phase measurements obtained by Fourier transform profilemetry are used to estimate the unknown phase shifts. The phase errors caused by motion are greatly reduced in this way. Experimental results show that the proposed method can obtain accurate and dense 3D reconstruction of dynamic scenes, with regard to different kinds of motion.
Pengyu Cong, Yueyi Zhang 0001, Zhiwei Xiong, Feng Wu 0001
VCIP2
2012 Hybrid structured light for scalable depth sensing
abstract
Time multiplexing and spatial neighborhood are two mainstream structured light techniques widely used for 3D shape measurement. In this paper, we explore a way to subtly integrate their advantages for scalable depth sensing. This is realized through a set of elaborate hybrid structured patterns, which consists of three sinusoidal fringe patterns with different initial phases modulated by a pseudo-random speckle signal. For temporally static scenes, a high resolution, high accuracy depth map can be recovered from the latest three frames by the phase-shifting method; for dynamic scenes, a decent depth map can still be recovered from the current single frame by image matching. Since it provides seamless transition between high quality and quick response options, our method validates a new paradigm of scalable depth sensing in practice.
Yueyi Zhang 0001, Zhiwei Xiong, Feng Wu 0001
ICIP1
2012 Depth sensing with focus and exposure adaptation
abstract
Automatic focus and exposure are the key components in digital cameras nowadays, which jointly play an essential role for capturing a high quality image. In this paper, we make an attempt to address these two challenging issues for future depth cameras. Relying on a programmable projector, we establish a structured light system for depth sensing with focus and exposure adaptation. The basic idea is to change current illumination pattern and intensity locally according to the prior depth information. Consequently, object surfaces appearing at different depths in the scene can receive proper illumination respectively. In this way, more flexible and robust depth sensing can be achieved in comparison with fixed illumination, especially at near depth.
Zhiwei Xiong, Yueyi Zhang 0001, Pengyu Cong, Feng Wu 0001
VCIP2